I am Mohan Shi, a third-year Ph.D. student in Electrical and Computer Engineering at University of California, Los Angeles (UCLA). I received my master degree at the University of Science and Technology of China (USTC). My research interests span a variety of domains in the world of speech processing:
- Automatic Speech Recognition
- Speech-centric Large Language Models
- Child/Low-resource Speech Processing
- Cocktail Party Problems
I am seeking research internship opportunities in 2027. Please reach out if you have any leads.
Experience
Microsoft Research, Redmond, USA
- Research Intern, CoreAI Speech Team, Jun 2026 – Sep 2026
- With Jinyu Li, Ruchao Fan, Keqi Deng, Sunit Sivasankaran
- Topic: Native encoder-free Speech-LLMs pretraining
Microsoft Research, Redmond, USA
- Research Intern, CoreAI Speech Team, Jun 2025 – Sep 2025
- With Jinyu Li, Xiong Xiao, Ruchao Fan, Shaoshi Ling
- Topic: In-context learning and emergent capabilities of Speech-LLMs; joint ASR and speaker diarization
Tencent AI Lab, Bellevue, USA (remote)
- Research Intern, Seattle Speech Lab, Sep 2023 – Aug 2024
- With Yong Xu, Shi-Xiong (Austin) Zhang, Dong Yu
- Topic: Speech-LLMs; multi-talker ASR
Alibaba Group, Hangzhou, China
- Research Intern, Speech Team (now the Qwen-Audio team), Jul 2022 – May 2023
- With Shiliang Zhang, Zhihao Du
- Topic: ASR; voice activity detection
Selected Publications
Teaching LLMs to Hear Who Spoke What: Metadata-Supervised Pretraining for Encoder-Free Speech-LLMs, Preprint [link]
Mohan Shi, Ruchao Fan, Sunit Sivasankaran, Keqi Deng, Jinyu LiEncoder Awakening via Adapters: Effective Domain-Adaptive Fine-tuning of Speech-LLMs, SLT 2026 [link]
Mohan Shi, Zilai Wang, Natarajan Balaji Shankar, Kaiyuan Zhang, Eray Eren, Abeer AlwanTrain Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio, ICASSP 2026 [link]
Mohan Shi, Xiong Xiao, Ruchao Fan, Shaoshi Ling, Jinyu LiSTACodec: Semantic Token Assignment for Balancing Acoustic Fidelity and Semantic Information in Audio Codecs, ICASSP 2026 [link]
Kaiyuan Zhang*, Mohan Shi*, Eray Eren, Natarajan Balaji Shankar, Zilai Wang, Abeer AlwanAdvancing Multi-talker ASR Performance with Large Language Models, SLT 2024 [link]
Mohan Shi, Zengrui Jin, Yaoxun Xu, Yong Xu, Shi-Xiong Zhang, Kun Wei, Yiwen Shao, Chunlei Zhang, Dong YuLibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization, Interspeech 2024 (Oral) [link]
Zengrui Jin*, Yifan Yang*, Mohan Shi*, Wei Kang, Xiaoyu Yang, Zengwei Yao, Fangjun Kuang, Liyong Guo, Lingwei Meng, Long Lin, Yong Xu, Shi-Xiong Zhang, Daniel PoveySemantic VAD: Low-Latency Voice Activity Detection for Speech Interaction, Interspeech 2023 (Oral) [link]
Mohan Shi, Yuchun Shu, Lingyun Zuo, Qian Chen, Shiliang Zhang, Jie Zhang, Li-Rong DaiCASA-ASR: Context-Aware Speaker-Attributed ASR, Interspeech 2023 [link]
Mohan Shi, Zhihao Du, Qian Chen, Fan Yu, Yangze Li, Shiliang Zhang, Jie Zhang, Li-Rong Dai
Education
University of California, Los Angeles
- Ph.D. student, Electrical and Computer Engineering, Sep 2024 – Present
University of Science and Technology of China
- Master of Engineering, Electronic Engineering and Information Science, Sep 2021 – Jun 2024
Dalian University of Technology
- Bachelor of Engineering, Electronic Information Engineering, Sep 2017 – Jun 2021
