Indian Flag
Government Of India
A-
A
A+
ORGANISATION

IndicXlit

A Transformer-based multilingual transliteration model

  • See Upvoters0
  • Downloads69
  • File Size3.94 MB
  • Views1,584

About Model

Bhashini - IndicXlit is a Transformer-based multilingual transliteration model, trained on Aksharantar dataset which is the largest publicly available parallel transliteration corpora collection for Indic languages at the time of writing (20 May 2022). It is used to convert any roman text written in Indian language (like Hinglish) to the native Indic-script (like Devanagari for Hindi). It supports 21 Indic languages: Assamese, Bangla, Bodo, Gujarati, Hindi, Kannada, Kashmiri, Konkani, Maithili, Malayalam, Manipuri, Marathi, Nepali, Oriya, Panjabi, Sanskrit, Sindhi, Sinhala, Tamil, Telugu, Urdu.

IndicXlit

Metadata Metadata

MIT

AI4Bharat

Machine Translation Model

Other

Open

Sector Agnostic

06/07/26 16:14:08

3.94 MB

Tags Tags

  • Language Modeling
  • Multilingual Translation
  • Machine Translation
  • Regional Languages
  • Indian Languages
  • NLP
  • transliteration

IndicXlit-master ( 3 files, 12 directories )


Directory
ablation_study

2 directories

Directory
app

10 files, 1 directories

Directory
Checker

3 files

Directory
corpus_preprocessing

5 directories

Directory
data_mining

1 files, 2 directories

Directory
Dataset_Format

2 files

Directory
inference

2 directories

Directory
model_training_scripts

1 files, 7 directories

undefined
.gitignore

1.79 KB

undefined
LICENSE

1.04 KB

This preview shows 10 out of 15 items. Load more

License Control License Control

MIT

Version Control Version Control

FolderVersion 1(3.94 MB)
  • admin·1 year(s) ago
    • chevron_rightFolder
      IndicXlit-master
      • chevron_rightFolder
        ablation_study
      • chevron_rightFolder
        app
      • chevron_rightFolder
        Checker
      • chevron_rightFolder
        corpus_preprocessing
      • chevron_rightFolder
        data_mining
      • chevron_rightFolder
        Dataset_Format
      • chevron_rightFolder
        inference
      • chevron_rightFolder
        model_training_scripts
      • undefined
        .gitignore
      • undefined
        LICENSE
      • more_horiz 5 more

More Models from Digital India BHASHINI Division More Models from Digital India BHASHINI Division

BHASHINI IISc Sourashtra VITS TTS Models
VITS text-to-speech models for Sourashtra, a low-resource Indo-Aryan language. There are male and female voices, and each voice accepts text in either Tamil script or Sourashtra script.
sourashtra
Text to Speech
VITS
low-resource
tts
  • See Upvoters0
  • Downloads1
  • File Size0
  • Views40
Updated 5 day(s) ago

DIGITAL INDIA BHASHINI DIVISION

MIZO LANGUAGE OCR
This dataset is designed for training and evaluating Optical Character Recognition (OCR) models capable of extracting text from document images. It contains 14,000 labeled samples consisting of image filenames and their corresponding ground-truth text annotations, along with metadata such as font type, font size, and document category. The dataset includes diverse document styles, including printed books and scanned documents, enabling models to learn robust text recognition under varying visual
Mizoram
Mizo
  • See Upvoters0
  • Downloads1
  • File Size93.73 MB
  • Views26
Updated 20 day(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-GUJARATI
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
SSL_finetunning
Data2vec_aqc
spring_lab
IITM
gujarati
ssl
Low-resource languages
  • See Upvoters0
  • Downloads4
  • File Size3.52 GB
  • Views42
Updated 2 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-HINDI
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
Low-resource languages
ssl
IITM
spring_lab
Data2vec_aqc
SSL_finetunning
hindi
  • See Upvoters0
  • Downloads2
  • File Size3.53 GB
  • Views44
Updated 2 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-MANIPURI
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
Manipuri
Low Resource Languages
SSL_finetunning
Data2vec_aqc
spring_lab
IITM
ssl
  • See Upvoters0
  • Downloads2
  • File Size3.52 GB
  • Views39
Updated 2 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-ASSAMESE
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
ssl
IITM
spring_lab
Data2vec_aqc
SSL_finetunning
Low-resource languages
Assamese
  • See Upvoters0
  • Downloads3
  • File Size3.52 GB
  • Views27
Updated 2 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

IndicXlit
A Transformer-based multilingual transliteration model
NLP
transliteration
Indian Languages
Regional Languages
Machine Translation
Multilingual Translation
Language Modeling
  • See Upvoters0
  • Downloads69
  • File Size3.94 MB
  • Views1,585
Updated 3 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

Indic Trans2
AI4Bharat's Indic-Trans-v2 is a multilingual Transformer (~1.1BM) NMT model trained on Samanantar v2 dataset which is the largest publicly available parallel corpora collection for languages of India at the time of writing (23 March 2023). We currently release two models - Indic to English and English to Indic and support all the 22 scheduled languages of India.
Computational Linguistics
Machine Translation
Language Modeling
Bilingual Translation
Multilingual Translation
Machine Translation
Regional Languages
Indian Languages
Indic-TransV2
NLP
  • See Upvoters1
  • Downloads97
  • File Size214.60 KB
  • Views3,236
Updated 3 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

Bhashini - Fastspeech2 Model using (HS)
Text-to-speech models trained using FastPitch and HiFi-GAN vocoder, separately for each language. Supports both 'female' and 'male' voices.
Multilingual
NLP
Text Processing
Transformer
Text to Speech
Language Detection
  • See Upvoters0
  • Downloads125
  • File Size286.72 MB
  • Views2,472
Updated 3 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

Bhashini - IndicNER
IndicNER is a multilingual Named Entity Recognition model fine-tuned on 11 Indian languages to identify named entities in text
NLP
Transformer
Token Classification
Pytorch
Samanantar
Bert
NER
Multilingual
Foreigners
  • See Upvoters2
  • Downloads205
  • File Size591.28 MB
  • Views3,189
Updated 3 month(s) ago

DIGITAL INDIA BHASHINI DIVISION