
This dataset contains transcriptions of studio recordings of Dehwali Bhili from Nandurbar, along with corresponding audio files and duration in seconds, providing a collection of transcribed audio recordings for research and model training.
Bhili is an indigenous Indo-Aryan language spoken by over 10 million members of the Bhil tribal community across western and central India. Through Project Astitva, Karya collaborated with 400+ native Bhili speakers in partnership with the district administration of Nandurbar, Maharashtra, to create high-quality digital language resources for Dehwali Bhili. These efforts enabled the development of India’s first tribal language model in Bhili, which has subsequently been integrated into MahaVISTAAR AI with support from the Maharashtra state government and Bhashini.This dataset contains transcriptions of studio recordings from Nandurbar in Dehwali Bhili covering agriculture, forestry, and rural development topics. It is intended for researchers, analysts, educators, and developers working in agriculture, linguistics, and natural language processing. The dataset can support research on agricultural practices and farmer perspectives, language documentation, educational initiatives, and the development of Marathi language technologies. Potential applications include language modeling, sentiment analysis, topic modeling, speech and text processing, and studies of the linguistic and cultural context of the Nandurbar region.
To Provide A Collection Of Transcribed Audio Recordings That Support Research, Education, And Technology Development In Agriculture And Marathi Language Processing. Use Cases:1. Agricultural Research: Analyze The Transcriptions To Identify Trends And Patterns In Agricultural Practices. 2. Language Technology Development: Develop And Evaluate Marathi Language Models, Speech Technologies, And Other Nlp Applications. 3. Education: Support Marathi Language Learning And Provide Educational Content On Agricultural Practices. 4. Natural Language Processing: Enable Tasks Such As Language Modeling, Text Classification, Information Extraction, Sentiment Analysis, And Language Understanding.
Attribution 4.0 International (CC BY- 4.0)
7102 files
2.06 MB
© 2026 - Copyright AIKosh. All rights reserved.