Jaeyeon Kim

prof_pic.jpg

Hi everyone! I’m Jaeyeon Kim, a PhD student at the Language Technologies Institute, Carnegie Mellon University, co-advised by Professor Carlos Busso and Professor Shinji Watanabe. I also spent summer 2026 as a research intern at Microsoft Research, mentored by Dr. Hannes Gamper. Before joining CMU, I was a research intern at the Vision and Learning Lab at Seoul National University, advised by Professor Gunhee Kim.

I aim to build AI that perceives and reasons about the world through multiple modalities, bridging language and vision with auditory information. Specifically, I am interested in (1) advancing audio understanding and reasoning in large audio and omni-modal language models and (2) improving the comprehension and generation of audio(-visual) scenes, including spatial audio.

selected publications

  1. wow_bench.png
    WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations
    Jaeyeon Kim , Heeseung Yun, Sang Hoon Woo, Chao-Han Huck Yang, and Gunhee Kim
    In Findings of ACL, 2026
  2. scene_aware_v2sa.png
    Towards Scene-Aware Video-to-Spatial Audio Generation
    Jaeyeon Kim* , Heeseung Yun*, and Gunhee Kim
    International Journal of Computer Vision (IJCV), 2026
  3. visage.png
    ViSAGe: Video-to-Spatial Audio Generation
    Jaeyeon Kim , Heeseung Yun, and Gunhee Kim
    In ICLR, 2025
  4. enclap.png
    EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
    Jaeyeon Kim , Jaeyoon Jung, Jinjoo Lee, and Sang Hoon Woo
    In ICASSP, 2024