Indian Flag
Government Of India
A-
A
A+
ORGANISATION

Bhashini-AI4Bharat Textual Language Detection v1.0

Detect language from provided text, Currently supports 23 languages (English, Bangla, Manipuri, Bodo, Konkani, Oriya, Nepali, Marathi, Sindhi, Sanskrit, Malayalam, Urdu, Assamese, Telugu, Dogri, Gujarati, Kashmiri, Punjabi, Santali, Maithili, Hindi, Tamil, Kannada)

  • See Upvoters5
  • Downloads291
  • File Size3 MB
  • Views5,590

About Model

IndicLID, is a language identifier for all 22 Indian languages listed in the Indian constitution in both native-script and romanized text. IndicLID is the first LID for romanized text in Indian languages. It is a two stage classifier that is ensemble of a fast linear classifier and a slower classifier finetuned from a pre-trained LM. It can predict 47 classes (24 native-script classes and 21 roman-script classes plus English and Others). IndicLID is evaluated on Bhasha-Abhijnaanam benchmark which is released alnog with this work. For native-script text, IndicLID has better language coverage than existing LIDs and is competitive or better than other LIDs. IndicLID model is 10 times faster and 4 times smaller than the NLLB model also establish a strong baseline results on the roman-script text.

Bhashini-AI4Bharat Textual Language Detection v1.0

Metadata Metadata

MIT

AI4Bharat

OCR (Optical Character Recognition) Model

Other

Open

Sector Agnostic

06/07/26 16:08:10

3 MB

Tags Tags

  • Multilingual
  • AI4Bharat
  • NLP
  • Bhashini
  • Text Processing
  • Deep Learning
  • Transformer
  • Text Language Detection

compile_final_pilot_1.py ( 1.81 KB )


To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.

License Control License Control

MIT

Version Control Version Control

FolderVersion 2(3 MB)
  • admin·1 year(s) ago
    • chevron_rightFolder
      Benchmark
      • undefined
        compile_final_pilot_1.py
      • undefined
        create_benchmark_extra.py
      • undefined
        create_benchmark.py
    • chevron_rightFolder
      deployement
    • chevron_rightFolder
      filter_Dakshina
    • chevron_rightFolder
      final_runs_ACL_inference
    • chevron_rightFolder
      final_runs_train
    • chevron_rightFolder
      Inference
    • chevron_rightFolder
      nueral_net
    • chevron_rightFolder
      preprocess_indiccorp
    • text/markdown
      README.md

More Models from Digital India BHASHINI Division More Models from Digital India BHASHINI Division

SPRING-INX-DATA2VEC-AQC-GUJARATI
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
ssl
Low-resource languages
SSL_finetunning
Data2vec_aqc
spring_lab
IITM
gujarati
  • See Upvoters0
  • Downloads2
  • File Size3.52 GB
  • Views31
Updated 30 day(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-HINDI
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
Low-resource languages
ssl
IITM
spring_lab
Data2vec_aqc
SSL_finetunning
hindi
  • See Upvoters0
  • Downloads0
  • File Size3.53 GB
  • Views30
Updated 30 day(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-MANIPURI
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
Manipuri
Low Resource Languages
SSL_finetunning
Data2vec_aqc
spring_lab
IITM
ssl
  • See Upvoters0
  • Downloads1
  • File Size3.52 GB
  • Views22
Updated 30 day(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-ASSAMESE
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
SSL_finetunning
Assamese
ssl
IITM
spring_lab
Data2vec_aqc
Low-resource languages
  • See Upvoters0
  • Downloads1
  • File Size3.52 GB
  • Views16
Updated 30 day(s) ago

DIGITAL INDIA BHASHINI DIVISION

IndicXlit
A Transformer-based multilingual transliteration model
NLP
Language Modeling
Multilingual Translation
Machine Translation
Regional Languages
Indian Languages
transliteration
  • See Upvoters0
  • Downloads54
  • File Size3.94 MB
  • Views1,368
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

Indic Trans2
AI4Bharat's Indic-Trans-v2 is a multilingual Transformer (~1.1BM) NMT model trained on Samanantar v2 dataset which is the largest publicly available parallel corpora collection for languages of India at the time of writing (23 March 2023). We currently release two models - Indic to English and English to Indic and support all the 22 scheduled languages of India.
Regional Languages
NLP
Indic-TransV2
Indian Languages
Machine Translation
Multilingual Translation
Bilingual Translation
Language Modeling
Computational Linguistics
Machine Translation
  • See Upvoters1
  • Downloads95
  • File Size214.60 KB
  • Views2,802
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

Bhashini - Fastspeech2 Model using (HS)
Text-to-speech models trained using FastPitch and HiFi-GAN vocoder, separately for each language. Supports both 'female' and 'male' voices.
Transformer
Text to Speech
Language Detection
Multilingual
NLP
Text Processing
  • See Upvoters0
  • Downloads112
  • File Size286.72 MB
  • Views2,150
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

Bhashini - IndicNER
IndicNER is a multilingual Named Entity Recognition model fine-tuned on 11 Indian languages to identify named entities in text
NLP
Foreigners
Multilingual
Transformer
Token Classification
Pytorch
Samanantar
Bert
NER
  • See Upvoters2
  • Downloads191
  • File Size591.28 MB
  • Views2,945
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

Bhashini-AI4Bharat Textual Language Detection v1.0
Detect language from provided text, Currently supports 23 languages (English, Bangla, Manipuri, Bodo, Konkani, Oriya, Nepali, Marathi, Sindhi, Sanskrit, Malayalam, Urdu, Assamese, Telugu, Dogri, Gujarati, Kashmiri, Punjabi, Santali, Maithili, Hindi, Tamil, Kannada)
NLP
Text Processing
Deep Learning
Transformer
Text Language Detection
Multilingual
AI4Bharat
Bhashini
  • See Upvoters5
  • Downloads291
  • File Size3 MB
  • Views5,591
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-SANSKRIT
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text. The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
low-resource-language
SSL_finetunning
Data2vec_aqc
spring_lab
IITM
ssl
Sanskrit
  • See Upvoters0
  • Downloads5
  • File Size3.52 GB
  • Views224
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION