
This dataset contains audio speech data in the Dehwali Bhili language, along with corresponding sentences and audio file metadata.
Bhili is an indigenous Indo-Aryan language spoken by over 10 million members of the Bhil tribal community across western and central India. Through Project Astitva, Karya collaborated with 400+ native Bhili speakers in partnership with the district administration of Nandurbar, Maharashtra, to create high-quality digital language resources for Dehwali Bhili. These efforts enabled the development of India’s first tribal language model in Bhili, which has subsequently been integrated into MahaVISTAAR AI with support from the Maharashtra state government and Bhashini. This dataset is comprised of audio speech data in the Dehwali Bhili language, collected for the purpose of language research and development. The dataset contains speech samples from various individuals, each accompanied by a corresponding sentence and audio file metadata. The dataset is primarily intended for use in language modeling, speech recognition, and natural language processing applications. It has the potential to contribute to the development of more accurate and efficient language processing systems, particularly for the Dehwali Bhili language. The dataset can be utilized by researchers, developers, and organizations working in the fields of natural language processing, speech recognition, and language research. The dataset provides valuable resources for language researchers and enthusiasts, enabling them to analyze the characteristics of the Dehwali Bhili language and its applications in various fields. It can be used to train and evaluate natural language processing models, improving their accuracy and efficiency in processing Dehwali Bhili language data. The dataset can also be employed to develop and fine-tune speech recognition systems for the Dehwali Bhili language, enhancing their ability to recognize and transcribe spoken language. Additionally, the dataset can be utilized to build and refine language models for the Dehwali Bhili language, enabling more accurate and context-aware language generation. Furthermore, the dataset can be used to document and preserve the Dehwali Bhili language, providing valuable resources for language researchers and enthusiasts.
To Facilitate Language Research And Development, Particularly In The Dehwali Bhili Language. Use Cases: 1. Natural Language Processing: This Dataset Can Be Used To Train And Evaluate Natural Language Processing Models, Improving Their Accuracy And Efficiency In Processing Dehwali Bhili Language Data. 2. Speech Recognition: The Dataset Can Be Employed To Develop And Fine-tune Speech Recognition Systems For The Dehwali Bhili Language, Enhancing Their Ability To Recognize And Transcribe Spoken Language. 3. Language Modeling: This Dataset Can Be Utilized To Build And Refine Language Models For The Dehwali Bhili Language, Enabling More Accurate And Context-aware Language Generation. 4. Language Documentation: The Dataset Can Be Used To Document And Preserve The Dehwali Bhili Language, Providing Valuable Resources For Language Researchers And Enthusiasts.
Attribution 4.0 International (CC BY- 4.0)
54563 files
23.49 MB
© 2026 - Copyright AIKosh. All rights reserved.