Labelled ASR corpus of over 6,400 hours mined from All India Radio news bulletins across 12 Indian languages, funded by Bhashini and MeitY.
Shrutilipi is a large labelled speech corpus created by AI4Bharat by mining and aligning over 6,400 hours of audio from All India Radio news bulletins across 12 Indian languages. The project was supported by the Bhashini mission and the Ministry of Electronics and Information Technology (MeitY), reflecting a national effort to build large-scale, high-quality speech resources for Indian languages. The dataset provides transcribed audio suitable for training automatic speech recognition systems at scale.
Shrutilipi Is Used To Train And Benchmark Automatic Speech Recognition Systems For Indian Languages, Particularly In News And Formal Broadcast Contexts. It Supports The Broader Bhashini Initiative To Build Inclusive Language Technology Infrastructure For India And Is Used By Researchers To Study Large-scale Weakly Supervised Asr Data Mining Techniques.
Attribution 4.0 International (CC BY- 4.0)
© 2026 - Copyright AIKosh. All rights reserved.