Career Profile

Hi, I’m a Ph.D. student in CSE at Seoul National University, advised by Gunhee Kim at the Vision and Learning Lab. I work on real-time conversational AI — spoken interaction where a model listens, thinks, speaks, and acts at once, rather than treating speech as text with a voice attached. I’m currently a research intern at KRAFTON AI, where I develop speech language models and full-duplex conversational models, Raon-Speech and Raon-SpeechChat.

  • Full-Duplex Spoken Dialogue: Agents that listen, speak, and act at once — streaming, low-latency interaction, and the behavioral dynamics of how people actually talk.
  • Omnimodal Understanding: Language models that perceive speech, audio, and vision natively, including the paralinguistic cues that a text transcript throws away.
  • Situated Interaction: Grounding conversation in the surrounding scene — egocentric perception, multi-party settings, and resolving what an utterance refers to.

Education

M.S./Ph.D. in Computer Science and Engineering

2022 - Now
Seoul National University

Advisor: Gunhee Kim

B.S. in Electrical and Computer Engineering

2015 - 2022
Seoul National University

Graduated with Summa Cum Laude

Experiences

Research Intern

2025.11 - Present
KRAFTON AI

Projects

A.X K2 Raon-Speech-21B-A3B - Bilingual (Korean/English) speech language model built during my research internship at KRAFTON AI. A 21B-A3B mixture-of-experts backbone coupled with an AuT speech encoder and a Mimi-style neural audio codec, supporting speech recognition, speech generation, voice conditioning, and paralinguistic understanding.

Publications

SitCom: Scaling Egocentric Multi-Party Spoken Dialogue for Situated Communication Assistance
Heeseung Yun*, Sehun Lee*, Sang Hoon Woo, Yoonji Nam, Sung-Feng Huang, Chao-Han Huck Yang, Gunhee Kim
Under Review, 2026
StreamAlign: Streaming Text-Aligned Speech Tokenization
Kang-wook Kim, Jinyoung Park, Jinsoo Kim, Sehun Lee, Sang Hoon Woo, Gunhee Kim
Under Review, 2026
DExTER: Can Omnimodal Language Models Resolve Audio-Visual Deixis?
Sehun Lee, Yoonji Nam, Sang Hoon Woo, Gunhee Kim
Under Review, 2026
Raon-Speech / Raon-SpeechChat
KRAFTON AI
In Technical Report, 2026
Think, Verbalize, then Speak: Bridging Complex Thoughts and Comprehensible Speech
Sehun Lee*, Sang Hoon Woo*, Kang-wook Kim, Gunhee Kim
In EMNLP, 2025
Behavior-SD: Behaviorally Aware Spoken Dialogue Generation with Large Language Models
Sehun Lee*, Kang-wook Kim*, Gunhee Kim
In NAACL, 2025 (Oral)
🏆 SAC Award Winner 🏆 in Speech Processing and Spoken Language Understanding
Meta-Learning Approach for Joint Multimodal Signals with Multimodal Iterative Adaptation
Sehun Lee*, Wonkwang Lee*, Gunhee Kim
In TMLR, 2024
Panoramic Vision Transformer for Saliency Detection on 360º Videos
Heeseung Yun, Sehun Lee, Gunhee Kim
In ECCV, 2022