Pan-India multilingual, multi-domain speech corpus collected across districts, capturing regional accents and dialects.
Vaani is a pan-India speech data collection initiative led by the Indian Institute of Science (IISc) and ARTPARK, aimed at capturing multilingual and multi-domain speech from volunteers across hundreds of districts. The project systematically records speech samples that reflect regional accents, dialects, and everyday topics, building one of the most geographically diverse speech corpora for Indian languages. Data collection is conducted in partnership with local communities to ensure broad linguistic representation.
Vaani Is Used To Develop And Evaluate Speech And Language Technologies That Are Robust To India's Regional And Dialectal Diversity. It Supports Building Inclusive Voice Assistants, Asr Systems, And Speech-based Applications Tailored To Underserved Regions And Languages. The Dataset Is Also Used In Research On Dialectal Variation, Accent Robustness, And Equitable Ai For Multilingual Populations.
Attribution 4.0 International (CC BY- 4.0)
© 2026 - Copyright AIKosh. All rights reserved.