Vahid Ahmadi Kalkhorani
Computer Science · The Ohio State University
Publications
21
Citations
120
Est. group size
~2
Recurring co-author estimate
Active years
6
Publishing since 2021
This researcher works on speech and audio processing, particularly methods that combine audio and visual information to separate overlapping speakers, enhance noisy speech, and improve automatic speech recognition (ASR). An earlier line of work, now less active, involved engineering topics like building air quality and cooling systems using metal-organic framework materials. The more recent and dominant focus is on deep learning models (e.g., 'CrossNet' architectures) for tasks like isolating a target speaker's voice from background noise or other talkers, using both audio and video (e.g., lip movement) cues.
Publication output has grown from essentially none before 2021 to a steady stream of several papers per year since 2022, suggesting an active and increasing research pace in recent years.
Generated by claude-sonnet-5 from public bibliographic data · Jul 20, 2026
- Audiovisual Speech Enhancement and Voice Activity Detection Using Generative and Speech Recognition Features
2026
- Elevating robust multi-talker ASR by decoupling speaker separation and speech recognition
Speech Communication · 2026
- AV-CrossNet: An Audiovisual Complex Spectral Mapping Network for Speech Separation by Leveraging Narrow- and Cross-Band Modeling
IEEE Journal of Selected Topics in Signal Processing · 2025
- Audiovisual speech enhancement and voice activity detection using generative and regressive visual features
Computer Speech & Language · 2025
- Elevating Robust ASR By Decoupling Multi-Channel Speaker Separation and Speech Recognition
2025
- Dual-Function Metal-Organic Framework (MOF) Wheel for CO2 and Moisture Control in Indoor Spaces
2025
- Online AV-CrossNet: a Causal and Efficient Audiovisual System for Speech Enhancement and Target Speaker Extraction
2025
- TF-CrossNet: Leveraging Global, Cross-Band, Narrow-Band, and Positional Encoding for Single- and Multi-Channel Speaker Separation
IEEE/ACM Transactions on Audio Speech and Language Processing · 2024
- Audiovisual Speaker Separation with Full- and Sub-Band Modeling in the Time-Frequency Domain
2024
- Towards Explainable Monaural Speaker Separation with Auditory-based Training
2024
- A Predictor-Corrector Method for Solving Boundary Layer Equations based on Bézier Curves
2024
- CrossNet: Leveraging Global, Cross-Band, Narrow-Band, and Positional Encoding for Single- and Multi-Channel Speaker Separation
arXiv (Cornell University) · 2024
- AV-CrossNet: an Audiovisual Complex Spectral Mapping Network for Speech Separation By Leveraging Narrow- and Cross-Band Modeling
arXiv (Cornell University) · 2024
- Time-domain Transformer-based Audiovisual Speaker Separation
2023
- Beyond the Frame: Single and mutilple video summarization method with user-defined length
arXiv (Cornell University) · 2023
- arXiv (Cornell University)×3
- IEEE/ACM Transactions on Audio Speech and Language Processing×1
- Building and Environment×1
- International Journal of Refrigeration×1
- Applied Thermal Engineering×1
- DeLiang Wang
Computer Science · The Ohio State University
- Hassan Taherian
Computer Science · The Ohio State University
- Donald S. Williamson
Computer Science · Indiana University
- Anurag Kumar
Computer Science · The Ohio State University
- Darius Petermann
Computer Science · Indiana University
This profile was generated automatically from public scholarly data (OpenAlex). Group size and activity levels are estimates derived from co-authorship patterns.
Last updated Jul 19, 2026.
Claim or correct this profile