Indian Flag
Government Of India
A-
A
A+
Vaani-KanniyaKumari-Dialect-Transcripts

Vaani-KanniyaKumari-Dialect-Transcripts

This dataset contains audio transcripts in the Tamil language from the Kanniya Kumari district, with speaker information, success metrics, and linguistic features.

About Dataset

This dataset is a collection of audio transcripts in the Tamil language, gathered from the Kanniya Kumari district. The transcripts are paired with speaker information, including speaker IDs and districts. The dataset also includes success metrics, indicating the accuracy of the speech-to-text model used to transcribe the audio. The dataset provides a resource for linguistic research and analysis of Vaani Kanniya Kumari dialect, supporting the development of speech recognition technology, language models, and language teaching materials. The dataset can be used to study the phonetics, phonology, and syntax of the dialect, as well as to develop language-based applications, such as language translation systems, speech-to-text systems, and text-to-speech systems. The dataset is particularly useful for researchers and developers working in the field of natural language processing, speech recognition, and language teaching.

Purpose of Dataset

To Support Research And Development Of Speech Recognition Technology And Linguistic Analysis Of Vaani Kanniya Kumari Dialect. Use Cases: 1. Speech Recognition: Improve Speech Recognition Models For Indian Languages. 2. Natural Language Processing: Develop Language Models And Speech Recognition Systems Using The Dataset. 3. Linguistics: Analyze The Phonetics, Phonology, And Syntax Of Vaani Kanniya Kumari Dialect. 4. Language Teaching: Use The Dataset To Develop Language Teaching Materials And Resources For Students Learning Vaani Kanniya Kumari Dialect.

Activity Overview Activity Overview

  • Downloads1
  • Downloads 2
  • File Size 44.50 MB
  • Views 29

Tags Tags

  • Linguistics
  • Tourism
  • Vaani_KanniyaKumari_Dialect
  • Tamil_language
  • Speech_recognition
  • Natural_language_processing
  • Indian_languages
  • AIKosh_Submission
  • Audio_transcripts
  • Vaani Kanniya Kumari dialect
  • Language_models
  • Tamil_Southern_Dialect
  • vachana-audio-intelligence-v2

License Control License Control

Attribution-Non-Commercial 4.0 International (CC BY-NC 4.0)

No Record(s) Found

Select a file to preview its contents.

Data Quality Score BetaData Quality Score Beta

Version Control Version Control

FolderVersion 1(44.50 MB)
  • S Arumuga Perumal·5 day(s) ago
    • text/csv
      Vaani_KanniyaKumari_Dialect - Cleaned_AIKosh_Submission.csv

Related Datasets Related Datasets

Updated 1 year(s) ago
VAANI: Multi-modal, Multi-lingual Dataset
VAANI: Multi-modal, Multi-lingual Dataset
Information-
VAANI is a multi-modal, multi-lingual dataset designed to represent the rich linguistic diversity of India. It currently includes data from two phases—Phase 1 (80 districts) and Phase 2 (40 districts)—spanning a total of ~21,500 hours of spontaneous, image-prompted speech collected from more than 110K speakers across 120 districts, describing 210K images in 86 languages. From this, 835 hours of transcribed audio data is available, distributed nearly evenly across all 120 districts.
Bengali
Gujarati
Kannada
Nepali
Punjabi
Telugu
Urdu
Sindhi
English
Tamil
Low-Resource Languages
Marathi
Malayalam
Odia
speech transcription
spontaneous speech
Sangtam
Ao
Halbi
Malvi
Sadri
Chhattisgarhi
Bhili
Malvani
multilingual corpus
Sumi
Bagheli
Khorth
Nyishi
multimodal dataset
Shekhawati
Bagri
Mewati
Meitei
86 Indian languages
Surgujia
audio-visual dataset
Garo
MagadhiMagahi
TTS training
Wancho
Awadhi
Galo
Oriya
speech + image + text
Tenyidie
regional dialects
Rajbanshi
Nagamese
manual transcription
Thethi
LLM speech integration
KhorthKhotta
Konkani
quality evaluated
speaker identification
diverse demographics
telemedicine AI applications
Marwadi
speaker diversity
Tulu
Assamese
Harauti
Rajasthani
language identification
120+ districts
Tagin
Marwari
Gondi
Bajjika
22 Indian states
Surjapuri
dialect diversity
Kokborok
image-prompted data
Santali
speech enhancement
Khariboli
Rengma
geo-centric data collection
Hajong
Hindi
ASR training
Wagdi
Bhatri
Dorli
multi-modal language resources
reallife recording environments
Mewari
linguistic diversity
Nimadi
Khandeshi
data for conversational AI
benchmarking dataset
Lotha
Kurmali
Bhojpuri
NissiDafla
Angika
Lambani
Magahi
Jaipuri
Bihari
Chakhesang
Magadhi
Angami
Chakma
Duruwa
Bearybashe
Khortha
Kumaoni
Kurukh
Garhwali
Maithili
Agariya
Bundeli
large-scale speech corpus
  • See Upvoters1
  • Downloads168
  • File Size0
  • Views1,946

INDIAN INSTITUTE OF SCIENCE (IISC), BANGALORE