See Google Scholar for a complete listing of papers and patents.


Papers

  • PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs
    Artem Dementyev, Wazeer Zulfikar, Sinan Hersek, Pascal Getreuer, Anurag Kumar, Vivek Kumar
    International Conference on Machine Learning (ICML), 2026
    [paper / arXiv] [pdf]

  • SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization
    Artem Dementyev, Wazeer Zulfikar, Sinan Hersek, Vivek Kumar
    arXiv preprint, 2025
    [paper / arXiv] [pdf]

  • SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMs
    Lijun Yu, Yong Cheng, Zhiruo Wang, Vivek Kumar, Wolfgang Macherey, Yanping Huang, David A. Ross, Irfan Essa, Yonatan Bisk, Ming-Hsuan Yang, Kevin P. Murphy, Alexander G. Hauptmann, Lu Jiang
    Neural Information Processing Systems (NeurIPS), 2023
    [paper / NeurIPS] [pdf]

  • Voice conversion with conditional SampleRNN
    Vivek Kumar et al.
    Interspeech, 2018
    [paper / arXiv] [pdf]

  • Transform-domain decorrelation in Dolby Digital Plus
    Vivek Kumar et al.
    IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2014
    [IEEE Xplore] [pdf]

  • Pseudo-Reliable Code Development
    Vivek Kumar
    Embedded Design, 2000
    [article] [pdf]


Talks & Presentations

I speak about Audio AI, deep learning for signal processing, and multimodal understanding and generation. My talks cover practical applications of AI in audio, the intersection of machine learning and traditional signal processing, and how foundation models are reshaping how we work with sound, speech, and music.


Patents