Indian Flag
Government Of India
A-
A
A+
ORGANISATION
Samanantar

Samanantar

Largest publicly available parallel corpus for Indic languages, covering 11 languages paired with English for machine translation.

About Dataset

Samanantar is a large parallel text corpus developed by AI4Bharat that pairs sentences in 11 Indian languages with their English translations. Compiled from a combination of existing parallel resources, web-mined data, and back-translation techniques, it is designed to be the largest publicly available parallel corpus for Indic languages at the time of release. The dataset spans diverse domains including news, government publications, and general web content.

Purpose of Dataset

Samanantar Is Used Primarily To Train And Evaluate Machine Translation Systems Between English And Indian Languages, As Well As Translation Among Indic Languages Themselves. It Supports Research In Low-resource Machine Translation, Multilingual Nlp, And Cross-lingual Transfer Learning. The Dataset Has Been Foundational For Building Open-source Indic Translation Models And Benchmarking Translation Quality Across The Indian Language Landscape.

Activity Overview Activity Overview

  • Downloads0
  • Redirect 2
  • File Size 0
  • Views 10

Tags Tags

  • Multilingual
  • Machine Translation
  • Parallel Corpus
  • Indic Languages

License Control License Control

CC0 1.0 Public Domain