Jaeyeon Kim

prof_pic.jpg

Hi everyone! I’m Jaeyeon Kim, a PhD student at the Language Technologies Institute, Carnegie Mellon University, co-advised by Professor Carlos Busso and Professor Shinji Watanabe. I also spent summer 2026 as a research intern at Microsoft Research, mentored by Dr. Hannes Gamper. Before joining CMU, I was a research intern at the Vision and Learning Lab at Seoul National University, advised by Professor Gunhee Kim.

I aim to build AI that perceives and reasons about the world through multiple modalities, bridging language and vision with auditory information. Specifically, I am interested in (1) advancing audio understanding and reasoning in large audio and omni-modal language models and (2) improving the comprehension and generation of audio(-visual) scenes, including spatial audio.

selected publications

  1. llt.png
    Long Listening Thoughts: Eliciting Open Auditory Reasoning with Deliberative Perception and Cognitive Refinement
    Jaeyeon Kim , Chao-Han Huck Yang, Luoyi Zhang, Chan-Jan Hsu, Fernando Ruiloba Portilla , and 3 more authors
    In Findings of EMNLP, 2026
  2. wow_bench.png
    WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations
    Jaeyeon Kim , Heeseung Yun, Sang Hoon Woo, Chao-Han Huck Yang, and Gunhee Kim
    In Findings of ACL, 2026
  3. scene_aware_v2sa.png
    Towards Scene-Aware Video-to-Spatial Audio Generation
    Jaeyeon Kim* , Heeseung Yun*, and Gunhee Kim
    International Journal of Computer Vision (IJCV), 2026
  4. visage.png
    ViSAGe: Video-to-Spatial Audio Generation
    Jaeyeon Kim , Heeseung Yun, and Gunhee Kim
    In ICLR, 2025
  5. enclap.png
    EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
    Jaeyeon Kim , Jaeyoon Jung, Jinjoo Lee, and Sang Hoon Woo
    In ICASSP, 2024