
This dataset contains spontaneous speech samples in Dehwali Bhili language, along with translations in English and Marathi, and metadata such as duration and recording information, to support research in language documentation, language learning, and speech recognition.
Bhili is an indigenous Indo-Aryan language spoken by over 10 million members of the Bhil tribal community across western and central India. Through Project Astitva, Karya collaborated with 400+ native Bhili speakers in partnership with the district administration of Nandurbar, Maharashtra, to create high-quality digital language resources for Dehwali Bhili. These efforts enabled the development of India’s first tribal language model in Bhili, which has subsequently been integrated into MahaVISTAAR AI with support from the Maharashtra state government and Bhashini. The Dehwali Bhili Spontaneous Speech Dataset and Language Translation is a collection of speech samples in the Dehwali Bhili language, which is a dialect spoken in certain regions of India. The dataset includes translations of these speech samples in English and Marathi, as well as metadata such as duration and recording information. The dataset is designed to support research in language documentation, language learning, and speech recognition, and can be used to improve the accuracy of speech recognition systems, develop more effective language translation models, and gain insights into the linguistic and cultural characteristics of the Dehwali Bhili language. The dataset can also be used to develop applications such as language learning software, speech-to-text systems, and voice assistants. The dataset is particularly useful for researchers and developers working in natural language processing, speech recognition, and language documentation. The dataset can be used to analyze the structure and syntax of the Dehwali Bhili language, as well as to develop and evaluate language learning tools and resources, such as language learning apps and language exchange platforms. The dataset can also be used to develop speech recognition systems, such as virtual assistants and voice recognition software, and to improve the accuracy of language translation models. The dataset can be used to support research in language documentation, language learning, and speech recognition, and can be used to develop and evaluate language learning tools and resources, such as language learning apps and language exchange platforms, and to improve the accuracy of speech recognition systems and language translation models.
To Support Research In Language Documentation, Language Learning, And Speech Recognition. Use Cases: 1. Language Documentation: The Dataset Can Be Used To Analyze The Structure And Syntax Of The Dehwali Bhili Language, As Well As To Develop And Evaluate Language Documentation Tools And Resources. 2. Language Learning: The Dataset Can Be Used To Develop And Evaluate Language Learning Tools And Resources, Such As Language Learning Apps And Language Exchange Platforms. 3. Speech Recognition: The Dataset Can Be Used To Develop And Evaluate Speech Recognition Systems, Such As Virtual Assistants And Voice Recognition Software. 4. Natural Language Processing: Improve The Accuracy Of Speech Recognition Systems And Language Translation Models.
Attribution 4.0 International (CC BY- 4.0)
13043 files
4.34 MB
© 2026 - Copyright AIKosh. All rights reserved.