Indian Flag
Government Of India
A-
A
A+
ORGANISATION
AI4Bharat IndicCorp v2

AI4Bharat IndicCorp v2

Large-scale monolingual text corpus spanning 23 Indian languages, built for pretraining Indic language models.

About Dataset

AI4Bharat IndicCorp v2 is a large monolingual text corpus covering 23 major Indian languages, curated by the AI4Bharat initiative at IIT Madras. The corpus aggregates cleaned web text, news articles, and other publicly available sources to provide substantial per-language text volume, addressing the scarcity of large-scale digital text for many Indic languages. It is deduplicated and filtered for quality, offering one of the most comprehensive open text resources for Indian language NLP.

Purpose of Dataset

Indiccorp V2 Is Primarily Used To Pretrain And Fine-tune Language Models For Indian Languages, Including Masked Language Models And Generative Llms. It Supports Research In Low-resource Nlp, Multilingual Representation Learning, And Cross-lingual Transfer For Indic Scripts. The Corpus Underlies Several Indic Language Model Releases And Is Widely Used To Benchmark Tokenization, Embedding Quality, And Downstream Task Performance Across Indian Languages.

Activity Overview Activity Overview

  • Downloads0
  • Redirect 2
  • File Size 0
  • Views 15

Tags Tags

  • Indic Languages
  • low-resource-NLP
  • pretraining
  • Text Corpus

License Control License Control

CC0 1.0 Public Domain