My research focuses on multimodal foundation models, spatial acoustic reasoning, and generative audio architectures.
For complete citation indexes, visit my Google Scholar Profile and Google Research Profile.
Foundation Models & Publications
2026
- PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs
Artem Dementyev, Wazeer Zulfikar, Sinan Hersek, Pascal Getreuer, Anurag Kumar, Vivek Kumar
International Conference on Machine Learning (ICML), 2026
2025
- SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization
Artem Dementyev, Wazeer Zulfikar, Sinan Hersek, Vivek Kumar
arXiv Preprint, 2025
2023
- SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMs
Lijun Yu, Yong Cheng, Zhiruo Wang, Vivek Kumar, Wolfgang Macherey, Yanping Huang, David A. Ross, Irfan Essa, Yonatan Bisk, Ming-Hsuan Yang, Kevin P. Murphy, Alexander G. Hauptmann, Lu Jiang
Neural Information Processing Systems (NeurIPS), 2023
Featured Keynote Talks
Selected Patents
- AI-Based Visual Content Collage Generation, US Patent Application 2025/0265751 A1, Google LLC, 2025
- Systems and methods for adapting human speaker embeddings in speech synthesis, US Patent 11,929,058, 2024
- Speech style transfer, US Patent 11,538,455, 2022
- Audio capture for aerial devices, US Patent 10,979,613, 2021