Indian Flag
Government Of India
A-
A
A+
ORGANISATION

Indic Trans2

AI4Bharat's Indic-Trans-v2 is a multilingual Transformer (~1.1BM) NMT model trained on Samanantar v2 dataset which is the largest publicly available parallel corpora collection for languages of India at the time of writing (23 March 2023). We currently release two models - Indic to English and English to Indic and support all the 22 scheduled languages of India.

  • See Upvoters1
  • Downloads94
  • File Size214.60 KB
  • Views2,727

About Model

Bhashini - IndicTrans2 is the first open-source transformer-based multilingual NMT model that supports high-quality translations across all the 22 scheduled Indic languages — including multiple scripts for low-resouce languages like Kashmiri, Manipuri and Sindhi. It adopts script unification wherever feasible to leverage transfer learning by lexical sharing between languages. Overall, the model supports five scripts Perso-Arabic (Kashmiri, Sindhi, Urdu), Ol Chiki (Santali), Meitei (Manipuri), Latin (English), and Devanagari (used for all the remaining languages).

We open-souce all our training dataset (BPCC), back-translation data (BPCC-BT), final IndicTrans2 models, evaluation benchmarks (IN22, which includes IN22-Gen and IN22-Conv) and training and inference scripts for easier use and adoption within the research community. We hope that this will foster even more research in low-resource Indic languages, leading to further improvements in the quality of low-resource translation through contributions from the research community.

This code repository contains instructions for downloading the artifacts associated with IndicTrans2, as well as the code for training/fine-tuning the multilingual NMT models.

For more details about the use of model, refer to github: https://github.com/AI4Bharat/IndicTrans2/tree/main

Indic Trans2

Metadata Metadata

MIT

AI4Bharat

Machine Translation Model

Other

Open

Sector Agnostic

06/07/26 16:12:22

214.60 KB

Tags Tags

  • Machine Translation
  • Computational Linguistics
  • Language Modeling
  • Bilingual Translation
  • Multilingual Translation
  • Machine Translation
  • Regional Languages
  • Indian Languages
  • Indic-TransV2
  • NLP

IndicTrans2-main ( 20 files, 5 directories )


Directory
baseline_eval

5 files

Directory
huggingface_interface

10 files

undefined
.gitignore

2.12 KB

undefined
apply_sentence_piece.sh

1.65 KB

undefined
compute_comet_score.sh

3.45 KB

undefined
compute_metrics_significance.sh

3.14 KB

undefined
compute_metrics.sh

1.30 KB

undefined
eval_rev.sh

2.10 KB

undefined
eval.sh

2.03 KB

undefined
finetune.sh

1.47 KB

This preview shows 10 out of 25 items. Load more

License Control License Control

MIT

Version Control Version Control

FolderVersion 1(214.60 KB)
  • admin·1 year(s) ago
    • chevron_rightFolder
      IndicTrans2-main
      • chevron_rightFolder
        baseline_eval
      • chevron_rightFolder
        huggingface_interface
      • undefined
        .gitignore
      • undefined
        apply_sentence_piece.sh
      • undefined
        compute_comet_score.sh
      • undefined
        compute_metrics_significance.sh
      • undefined
        compute_metrics.sh
      • undefined
        eval_rev.sh
      • undefined
        eval.sh
      • undefined
        finetune.sh
      • more_horiz 15 more

More Models from Digital India BHASHINI Division More Models from Digital India BHASHINI Division

SPRING-INX-DATA2VEC-AQC-GUJARATI
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
ssl
Low-resource languages
SSL_finetunning
Data2vec_aqc
spring_lab
IITM
gujarati
  • See Upvoters0
  • Downloads2
  • File Size3.52 GB
  • Views31
Updated 24 day(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-HINDI
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
Low-resource languages
ssl
IITM
spring_lab
Data2vec_aqc
SSL_finetunning
hindi
  • See Upvoters0
  • Downloads0
  • File Size3.53 GB
  • Views27
Updated 24 day(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-MANIPURI
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
Manipuri
Low Resource Languages
SSL_finetunning
Data2vec_aqc
spring_lab
IITM
ssl
  • See Upvoters0
  • Downloads1
  • File Size3.52 GB
  • Views22
Updated 24 day(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-ASSAMESE
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
SSL_finetunning
Assamese
ssl
IITM
spring_lab
Data2vec_aqc
Low-resource languages
  • See Upvoters0
  • Downloads1
  • File Size3.52 GB
  • Views16
Updated 24 day(s) ago

DIGITAL INDIA BHASHINI DIVISION

IndicXlit
A Transformer-based multilingual transliteration model
NLP
Language Modeling
Multilingual Translation
Machine Translation
Regional Languages
Indian Languages
transliteration
  • See Upvoters0
  • Downloads54
  • File Size3.94 MB
  • Views1,332
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

Indic Trans2
AI4Bharat's Indic-Trans-v2 is a multilingual Transformer (~1.1BM) NMT model trained on Samanantar v2 dataset which is the largest publicly available parallel corpora collection for languages of India at the time of writing (23 March 2023). We currently release two models - Indic to English and English to Indic and support all the 22 scheduled languages of India.
Indian Languages
Computational Linguistics
NLP
Indic-TransV2
Regional Languages
Machine Translation
Multilingual Translation
Bilingual Translation
Language Modeling
Machine Translation
  • See Upvoters1
  • Downloads94
  • File Size214.60 KB
  • Views2,728
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

Bhashini - Fastspeech2 Model using (HS)
Text-to-speech models trained using FastPitch and HiFi-GAN vocoder, separately for each language. Supports both 'female' and 'male' voices.
Transformer
Text to Speech
Language Detection
Multilingual
NLP
Text Processing
  • See Upvoters0
  • Downloads111
  • File Size286.72 MB
  • Views2,095
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

Bhashini - IndicNER
IndicNER is a multilingual Named Entity Recognition model fine-tuned on 11 Indian languages to identify named entities in text
NLP
Foreigners
Multilingual
Transformer
Token Classification
Pytorch
Samanantar
Bert
NER
  • See Upvoters2
  • Downloads176
  • File Size591.28 MB
  • Views2,914
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

Bhashini-AI4Bharat Textual Language Detection v1.0
Detect language from provided text, Currently supports 23 languages (English, Bangla, Manipuri, Bodo, Konkani, Oriya, Nepali, Marathi, Sindhi, Sanskrit, Malayalam, Urdu, Assamese, Telugu, Dogri, Gujarati, Kashmiri, Punjabi, Santali, Maithili, Hindi, Tamil, Kannada)
NLP
Text Processing
Deep Learning
Transformer
Text Language Detection
Multilingual
AI4Bharat
Bhashini
  • See Upvoters5
  • Downloads289
  • File Size3 MB
  • Views5,522
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-SANSKRIT
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text. The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
low-resource-language
SSL_finetunning
Data2vec_aqc
spring_lab
IITM
ssl
Sanskrit
  • See Upvoters0
  • Downloads5
  • File Size3.52 GB
  • Views220
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION