Large-scale monolingual text corpus spanning 23 Indian languages, built for pretraining Indic language models.
AI4Bharat IndicCorp v2 is a large monolingual text corpus covering 23 major Indian languages, curated by the AI4Bharat initiative at IIT Madras. The corpus aggregates cleaned web text, news articles, and other publicly available sources to provide substantial per-language text volume, addressing the scarcity of large-scale digital text for many Indic languages. It is deduplicated and filtered for quality, offering one of the most comprehensive open text resources for Indian language NLP.
Indiccorp V2 Is Primarily Used To Pretrain And Fine-tune Language Models For Indian Languages, Including Masked Language Models And Generative Llms. It Supports Research In Low-resource Nlp, Multilingual Representation Learning, And Cross-lingual Transfer For Indic Scripts. The Corpus Underlies Several Indic Language Model Releases And Is Widely Used To Benchmark Tokenization, Embedding Quality, And Downstream Task Performance Across Indian Languages.
CC0 1.0 Public Domain
© 2026 - Copyright AIKosh. All rights reserved.