Multilingual, multi-speaker speech dataset covering Indian languages and dialects, built for ASR and speech research.
IndicVoices is a large multilingual speech dataset developed by AI4Bharat, capturing spontaneous and read speech from a wide range of speakers across Indian languages and dialects. The dataset is designed to reflect natural conversational speech patterns, regional accents, and diverse recording conditions, addressing a major gap in speech resources for Indian languages. Recordings are paired with transcriptions to support supervised speech model training.
Indicvoices Is Used To Train And Evaluate Automatic Speech Recognition (Asr) And Related Speech Technologies For Indian Languages. It Supports Research On Accent And Dialect Robustness, Multilingual Acoustic Modeling, And Inclusive Voice Technology For Underrepresented Languages. The Dataset Plays A Key Role In Building Open Speech Models Tailored To India's Linguistic Diversity.
Attribution 4.0 International (CC BY- 4.0)
© 2026 - Copyright AIKosh. All rights reserved.