Career Profile
Hi, I’m a Ph.D. student in CSE at Seoul National University, advised by Gunhee Kim at the Vision and Learning Lab. I work on real-time conversational AI — spoken interaction where a model listens, thinks, speaks, and acts at once, rather than treating speech as text with a voice attached. I’m currently a research intern at KRAFTON AI, where I develop speech language models and full-duplex conversational models, Raon-Speech and Raon-SpeechChat.
- Full-Duplex Spoken Dialogue: Agents that listen, speak, and act at once — streaming, low-latency interaction, and the behavioral dynamics of how people actually talk.
- Omnimodal Understanding: Language models that perceive speech, audio, and vision natively, including the paralinguistic cues that a text transcript throws away.
- Situated Interaction: Grounding conversation in the surrounding scene — egocentric perception, multi-party settings, and resolving what an utterance refers to.
Education
Advisor: Gunhee Kim
Graduated with Summa Cum Laude
Experiences
Projects
A.X K2 Raon-Speech-21B-A3B
- Bilingual (Korean/English) speech language model built during my research internship at KRAFTON AI. A 21B-A3B mixture-of-experts backbone coupled with an AuT speech encoder and a Mimi-style neural audio codec, supporting speech recognition, speech generation, voice conditioning, and paralinguistic understanding.
Publications
SitCom: Scaling Egocentric Multi-Party Spoken Dialogue for Situated Communication Assistance
Under Review, 2026
StreamAlign: Streaming Text-Aligned Speech Tokenization
Under Review, 2026
DExTER: Can Omnimodal Language Models Resolve Audio-Visual Deixis?
Under Review, 2026