Largest publicly available parallel corpus for Indic languages, covering 11 languages paired with English for machine translation.
Samanantar is a large parallel text corpus developed by AI4Bharat that pairs sentences in 11 Indian languages with their English translations. Compiled from a combination of existing parallel resources, web-mined data, and back-translation techniques, it is designed to be the largest publicly available parallel corpus for Indic languages at the time of release. The dataset spans diverse domains including news, government publications, and general web content.
Samanantar Is Used Primarily To Train And Evaluate Machine Translation Systems Between English And Indian Languages, As Well As Translation Among Indic Languages Themselves. It Supports Research In Low-resource Machine Translation, Multilingual Nlp, And Cross-lingual Transfer Learning. The Dataset Has Been Foundational For Building Open-source Indic Translation Models And Benchmarking Translation Quality Across The Indian Language Landscape.
CC0 1.0 Public Domain
© 2026 - Copyright AIKosh. All rights reserved.