See Google Scholar for a complete listing of papers and patents.
Papers
PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs
Artem Dementyev, Wazeer Zulfikar, Sinan Hersek, Pascal Getreuer, Anurag Kumar, Vivek Kumar
International Conference on Machine Learning (ICML), 2026
[paper / arXiv] [pdf]SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization
Artem Dementyev, Wazeer Zulfikar, Sinan Hersek, Vivek Kumar
arXiv preprint, 2025
[paper / arXiv] [pdf]SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMs
Lijun Yu, Yong Cheng, Zhiruo Wang, Vivek Kumar, Wolfgang Macherey, Yanping Huang, David A. Ross, Irfan Essa, Yonatan Bisk, Ming-Hsuan Yang, Kevin P. Murphy, Alexander G. Hauptmann, Lu Jiang
Neural Information Processing Systems (NeurIPS), 2023
[paper / NeurIPS] [pdf]Voice conversion with conditional SampleRNN
Vivek Kumar et al.
Interspeech, 2018
[paper / arXiv] [pdf]Transform-domain decorrelation in Dolby Digital Plus
Vivek Kumar et al.
IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2014
[IEEE Xplore] [pdf]Pseudo-Reliable Code Development
Vivek Kumar
Embedded Design, 2000
[article] [pdf]
Talks & Presentations
I speak about Audio AI, deep learning for signal processing, and multimodal understanding and generation. My talks cover practical applications of AI in audio, the intersection of machine learning and traditional signal processing, and how foundation models are reshaping how we work with sound, speech, and music.
Audio AI — Challenges, Breakthroughs & Applications
PyTorch DevCon, 2019
[video / YouTube] [event & slides info]When Deep Learning Takes Over Signal Processing
Machine Learning Summit, San Francisco, 2018
[slides / SlideShare] [presentation info]Artificial Intelligence in Audio – Applications, Advancements And Trends
Dolby Soho, 477 Broadway, New York, 2019
[video / YouTube] [event info]Future of Audio
IEEE MIPR, Santa Clara, 2019
[slides & event info]BISH Bash Meetup: Audio AI
Hosted by Dolby Laboratories, San Francisco, 2019
[event info]Challenges of Doing Deep Learning with Audio
RE·WORK Applied AI Summit, San Francisco, 2018
[schedule & pdf]Learning Deep Learning
Deep Learning Meetup, 2017
[video / YouTube]
Patents
- Systems and methods for adapting human speaker embeddings in speech synthesis, US Patent 11,929,058, 2024
- Speech style transfer, US Patent 11,538,455, 2022
- Audio capture for aerial devices, US Patent 10,979,613, 2021
- Low bit rate parametric encoding and transport of haptic-tactile signals, US Patent application WO2017024001A1, 2017
- Adaptive quantization, US Patent application WO2017132366A1, 2017
- Time-varying filters for generating decorrelation signals, US Patent application 2014126684, 2014
- Signal decorrelation in an audio processing system, US Patent application US201443877, 2014
- Bit error concealment for audio coding system, US Patent US8301440B2, 2012
- Real time monitoring & control for audio device, US Patent US7778829B2, 2010
- Sampling rate mismatch solution, US Patent US7778373B2, 2010
- Bit error management methods for wireless audio communication channel, US Patent US8578247B2, 2013
- Method and apparatus for sharing a bluetooth module with two computing devices, US Patent 7,263,331, 2007