I am Mohan Shi, a second-year Ph.D. student in Electrical and Computer Engineering at University of California, Los Angeles (UCLA). I received my master degree at the University of Science and Technology of China (USTC) and worked under the guidance of Prof. Li-Rong Dai for three years. My research interests span a variety of domains in the world of speech processing:

  • Automatic Speech Recognition
  • Speech-centric Large Language Models
  • Child/Low-resource Speech Processing
  • Speech Tokenization
  • Cocktail Party Problems

I am seeking research internship opportunities for the Summer of 2027. Please reach out if you have any leads.

Experience


Microsoft Research, Redmond, USA

Microsoft Research, Redmond, USA

  • Research Intern, CoreAI Speech Team, Jun 2025 – Sep 2025
  • With Jinyu Li, Xiong Xiao, Ruchao Fan, Shaoshi Ling
  • Topic: In-context learning and emergent capabilities of Speech-LLMs; joint ASR and speaker diarization

Tencent AI Lab, Bellevue, USA (remote)

Alibaba Group, Hangzhou, China

  • Research Intern, Speech Team (now the Qwen-Audio team), Jul 2022 – May 2023
  • With Shiliang Zhang, Zhihao Du
  • Topic: ASR; voice activity detection

Selected Publications


  1. Train Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio, ICASSP 2026 [pdf]
    Mohan Shi, Xiong Xiao, Ruchao Fan, Shaoshi Ling, Jinyu Li

  2. STACodec: Semantic Token Assignment for Balancing Acoustic Fidelity and Semantic Information in Audio Codecs, ICASSP 2026 [pdf] [code]
    Kaiyuan Zhang*, Mohan Shi*, Eray Eren, Natarajan Balaji Shankar, Zilai Wang, Abeer Alwan

  3. Advancing Multi-talker ASR Performance with Large Language Models, SLT 2024 [pdf]
    Mohan Shi, Zengrui Jin, Yaoxun Xu, Yong Xu, Shi-Xiong Zhang, Kun Wei, Yiwen Shao, Chunlei Zhang, Dong Yu

  4. LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization, Interspeech 2024 (Oral) [pdf]
    Zengrui Jin*, Yifan Yang*, Mohan Shi*, Wei Kang, Xiaoyu Yang, Zengwei Yao, Fangjun Kuang, Liyong Guo, Lingwei Meng, Long Lin, Yong Xu, Shi-Xiong Zhang, Daniel Povey

  5. CASA-ASR: Context-Aware Speaker-Attributed ASR, Interspeech 2023 [pdf]
    Mohan Shi, Zhihao Du, Qian Chen, Fan Yu, Yangze Li, Shiliang Zhang, Jie Zhang, Li-Rong Dai

  6. Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction, Interspeech 2023 (Oral) [pdf]
    Mohan Shi, Yuchun Shu, Lingyun Zuo, Qian Chen, Shiliang Zhang, Jie Zhang, Li-Rong Dai

Education


University of California, Los Angeles

  • Ph.D. student, Electrical and Computer Engineering, Sep 2024 – Present

University of Science and Technology of China

  • Master of Engineering, Electronic Engineering and Information Science, Sep 2021 – Jun 2024

Dalian University of Technology

  • Bachelor of Engineering, Electronic Information Engineering, Sep 2017 – Jun 2021