Indian Flag
Government Of India
A-
A
A+
ORGANISATION
Vaani (IISc / ARTPARK)

Vaani (IISc / ARTPARK)

Pan-India multilingual, multi-domain speech corpus collected across districts, capturing regional accents and dialects.

About Dataset

Vaani is a pan-India speech data collection initiative led by the Indian Institute of Science (IISc) and ARTPARK, aimed at capturing multilingual and multi-domain speech from volunteers across hundreds of districts. The project systematically records speech samples that reflect regional accents, dialects, and everyday topics, building one of the most geographically diverse speech corpora for Indian languages. Data collection is conducted in partnership with local communities to ensure broad linguistic representation.

Purpose of Dataset

Vaani Is Used To Develop And Evaluate Speech And Language Technologies That Are Robust To India's Regional And Dialectal Diversity. It Supports Building Inclusive Voice Assistants, Asr Systems, And Speech-based Applications Tailored To Underserved Regions And Languages. The Dataset Is Also Used In Research On Dialectal Variation, Accent Robustness, And Equitable Ai For Multilingual Populations.

Activity Overview Activity Overview

  • Downloads0
  • Redirect 1
  • File Size 0
  • Views 9

Tags Tags

  • ASR
  • Indic Languages
  • regional dialects
  • Speech Data

License Control License Control

Attribution 4.0 International (CC BY- 4.0)