Indian Flag
Government Of India
A-
A
A+
ORGANISATION
Dravida Dataset

Dravida Dataset

Dravida dataset, a multilingual dataset for core Natural Language Understanding (NLU) tasks in four major Dravidian languages: Kannada, Malayalam, Tamil, Telugu. The dataset supports the following tasks: Morphological Analysis (MA), POS Tagging (POS), Named Entity Recognition (NER), Dependency Parsing (DEP), Coreference Resolution (CR)

About Dataset

The Dravida Dataset is a multilingual annotated corpus developed to advance Natural Language Understanding (NLU) for four major Dravidian languages: Kannada, Malayalam, Tamil, and Telugu. The dataset provides task-specific annotations for Morphological Analysis (MA), Part-of-Speech (POS) Tagging, Named Entity Recognition (NER), Dependency Parsing (DEP), and Coreference Resolution (CR), enabling research on morphologically rich and low-resource languages. The corpus combines manually annotated and carefully curated multilingual data, with synthetic data augmentation used for selected tasks where annotated resources are limited. // @inproceedings{p-m-etal-2025-family, title = "Family helps one another: {D}ravidian {NLP} suite for Natural Language Understanding", author = "P M, Abhinav and Dasari, Priyanka and Vuppala, Nagaraju and Krishnamurthy, Parameswari", editor = "Inui, Kentaro and Sakti, Sakriani and Wang, Haofen and Wong, Derek F. and Bhattacharyya, Pushpak and Banerjee, Biplab and Ekbal, Asif and Chakraborty, Tanmoy and Singh, Dhirendra Pratap", booktitle = "Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics", month = dec, year = "2025", address = "Mumbai, India", publisher = "The Asian Federation of Natural Language Processing and The Association for Computational Linguistics", url = "https://aclanthology.org/2025.findings-ijcnlp.120/", pages = "1926--1941", ISBN = "979-8-89176-303-6", abstract = "Developing robust Natural Language Understanding (NLU) for morphologically rich Dravidian languages like Kannada, Malayalam, Tamil, and Telugu presents significant challenges due to their agglutinative nature and syntactic complexity. In this work, we present the Dravidian NLP Suite tackling five core tasks: Morphological Analysis (MA), POS Tagging (POS), Named Entity Recognition (NER), Dependency Parsing (DEP), and Coreference Resolution (CR), trained for monolingual models and multilingual models. To facilitate this, we present the Dravida dataset, meticulously annotated multilingual corpus for these tasks across all four languages. Our experiments demonstrate that a multilingual model, which utilizes shared linguistic features and cross-lingual patterns inherent to the Dravidian family, consistently outperforms its monolingual counterparts across all tasks. These findings suggest that multilingual learning is an effective approach for enhancing Natural Language Understanding (NLU) capabilities, particularly for languages belonging to the same family. To the best of our knowledge, this is the first work to jointly address all these core tasks on the Dravidian languages." } // This dataset was identified and facilitated for onboarding as part of the Dataset Onboarding Support Team (DOST) initiative led by CivicDataLab (CDL), partnering with the Gates Foundation in collaboration with BHASHINI. CivicDataLab provided support for dataset discovery, validation, metadata preparation and onboarding facilitation. All dataset ownership and intellectual property rights remain with the original author(s).

Purpose of Dataset

The Purpose Of This Dataset Is To Support Natural Language Understanding Research For Dravidian Languages Through Multilingual Annotated Resources Covering Morphology, Syntax, Semantics, And Discourse. It Enables The Development And Evaluation Of Multilingual Language Models, Linguistic Analysis Tools, Information Extraction Systems, Dependency Parsers, Named Entity Recognizers, And Other Downstream Nlp Applications For Kannada, Malayalam, Tamil, And Telugu.

Activity Overview Activity Overview

  • Downloads0
  • Redirect 1
  • File Size 0
  • Views 11

Tags Tags

  • Kannada
  • Telugu
  • Computational Linguistics
  • Named Entity Recognition
  • cross-lingual
  • natural language processing (NLP)
  • multilingual corpus
  • indic-nlp
  • Malayalam language
  • Language Resource Development
  • Multilingual Applications
  • Tamil language
  • Linguistic annotation
  • Part-of-speech tagging
  • Morphological annotation
  • Universal Dependencies
  • Corpus linguistics
  • Natural Language Understanding

License Control License Control

Attribution 4.0 International (CC BY- 4.0)