My research focuses on multimodal foundation models, spatial acoustic reasoning, and generative audio architectures.
For complete citation indexes, visit my Google Scholar Profile and Google Research Profile.
Foundation Models & Publications
2026
- PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs
Artem Dementyev, Wazeer Zulfikar, Sinan Hersek, Pascal Getreuer, Anurag Kumar, Vivek Kumar
International Conference on Machine Learning (ICML), 2026
[Paper / arXiv] [PDF]
BibTeX
@inproceedings{dementyev2026phasecoder,
title={PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs},
author={Dementyev, Artem and Zulfikar, Wazeer and Hersek, Sinan and Getreuer, Pascal and Kumar, Anurag and Kumar, Vivek},
booktitle={International Conference on Machine Learning (ICML)},
year={2026},
url={https://arxiv.org/abs/2601.21124}
}
2025
- SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization
Artem Dementyev, Wazeer Zulfikar, Sinan Hersek, Vivek Kumar
arXiv Preprint, 2025
[Paper / arXiv] [PDF]
BibTeX
@article{dementyev2025speechcompass,
title={SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization},
author={Dementyev, Artem and Zulfikar, Wazeer and Hersek, Sinan and Kumar, Vivek},
journal={arXiv preprint arXiv:2502.08848},
year={2025},
url={https://arxiv.org/abs/2502.08848}
}
2023
- SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMs
Lijun Yu, Yong Cheng, Zhiruo Wang, Vivek Kumar, Wolfgang Macherey, Yanping Huang, David A. Ross, Irfan Essa, Yonatan Bisk, Ming-Hsuan Yang, Kevin P. Murphy, Alexander G. Hauptmann, Lu Jiang
Neural Information Processing Systems (NeurIPS), 2023
[Paper / NeurIPS] [PDF]
BibTeX
@inproceedings{yu2023spae,
title={SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMs},
author={Yu, Lijun and Cheng, Yong and Wang, Zhiruo and Kumar, Vivek and Macherey, Wolfgang and Huang, Yanping and Ross, David A and Essa, Irfan and Bisk, Yonatan and Yang, Ming-Hsuan and Murphy, Kevin P and Hauptmann, Alexander G and Jiang, Lu},
booktitle={Neural Information Processing Systems (NeurIPS)},
year={2023}
}
Keynote Talks & Presentations
Audio AI: Challenges, Breakthroughs & Applications
PyTorch DevCon, 2019
[Video / YouTube] [Event & Slides]When Deep Learning Takes Over Signal Processing
Machine Learning Summit, San Francisco, 2018
[Slides / SlideShare]Artificial Intelligence in Audio: Applications, Advancements and Trends
Dolby Soho, New York, 2019
[Video / YouTube]
Selected Patents
- AI-Based Visual Content Collage Generation, US Patent Application 2025/0265751 A1, Google LLC, 2025
- Systems and methods for adapting human speaker embeddings in speech synthesis, US Patent 11,929,058, 2024
- Speech style transfer, US Patent 11,538,455, 2022
- Audio capture for aerial devices, US Patent 10,979,613, 2021