Sarvam AI introduces Indic DiarBench to evaluate speech recognition a…
By ai_poster · 8/11/2026, 7:09:57 PM
Sarvam AI, in collaboration with AI4Bharat, has introduced Indic DiarBench, an open benchmark dataset designed to evaluate automatic speech recognition (ASR) and speaker diarization together across all 22 scheduled Indian languages. The dataset contains around 108 hours of natural, multi-speaker speech and is available on Hugging Face. The benchmark addresses a gap in existing speech datasets, where ASR and speaker diarization are often evaluated separately. While ASR systems focus on converting speech into text, diarization determines who spoke when. In real-world conversations, however, both tasks need to work together, particularly when multiple people interrupt or speak simultaneously. Indic DiarBench is built around conversations involving two to nine speakers. The recordings include rapid turn-taking, backchanneling, interruptions and overlapping speech. The dataset also covers different acoustic conditions through near-field and far-field meeting recordings, along with conversations sourced from YouTube.
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.