Indian Flag
Government Of India
A-
A
A+
ORGANISATION

AI4Bharat Textual Language Detection

Detect language from provided text, Currently supports 22 languages

About Model

IndicLID, is a language identifier for all 22 Indian languages listed in the Indian constitution in both native-script and romanized text. IndicLID is the first LID for romanized text in Indian languages. It is a two stage classifier that is ensemble of a fast linear classifier and a slower classifier finetuned from a pre-trained LM. It can predict 47 classes (24 native-script classes and 21 roman-script classes plus English and Others). IndicLID is evaluated on Bhasha-Abhijnaanam benchmark which is released alnog with this work. For native-script text, IndicLID has better language coverage than existing LIDs and is competitive or better than other LIDs. IndicLID model is 10 times faster and 4 times smaller than the NLLB model also establish a strong baseline results on the roman-script text.

AI4Bharat Textual Language Detection

Metadata Metadata

MIT

AI4Bharat

Text Language Detection

N.A.

Open

AI4Bharat

Sector Agnostic

21/02/25 13:21:38

0

Activity Overview Activity Overview

  • Downloads0
  • Redirect 38
  • File Size 0
  • Views 1,279

Tags Tags

  • Multilingual
  • AI4Bharat
  • NLP
  • Bhashini
  • Text Processing
  • Deep Learning
  • Text Language Detection

License Control License Control

MIT

More Models from AI4Bharat More Models from AI4Bharat

AI4Bharat- 500 M - RomanSetu Multilingual Native-to-Roman Model
RomanSetu is a multilingual continual pretrained transformer model designed for transliteration across six Indic languages
Multilingual
Instruction-Tuning
Llama
LLaMA2
  • See Upvoters1
  • Downloads53
  • File Size0
  • Views1,103
Updated 1 year(s) ago

AI4BHARAT

AI4Bharat- 400 M - RomanSetu Multilingual Native-to-Roman Model
RomanSetu is a multilingual continual pretrained transformer model designed for transliteration across six Indic languages
Multilingual
LLaMA2
Llama
Instruction-Tuning
  • See Upvoters1
  • Downloads94
  • File Size0
  • Views1,218
Updated 1 year(s) ago

AI4BHARAT

AI4Bharat- Maithili - IndicConformer Automatic Speech Recognition (ASR) Model
This model takes in mono-channel audio files at a 16,000 Hz sampling rate (WAV format) and outputs the transcribed text of the speech contained in the audio.
NLP
Automatic Speech Recognition
Speech-to-Text
  • See Upvoters0
  • Downloads29
  • File Size0
  • Views843
Updated 1 year(s) ago

AI4BHARAT

AI4Bharat- Konkani - IndicConformer Automatic Speech Recognition (ASR) Model
Automatic Speech Recognition (ASR) model for Konkani speech recognition, processing 16,000 KHz mono WAV audio and transcribing spoken content into text
Speech-to-Text
Automatic Speech Recognition
NLP
  • See Upvoters0
  • Downloads35
  • File Size0
  • Views850
Updated 1 year(s) ago

AI4BHARAT

AI4Bharat- Kashmiri - IndicConformer Automatic Speech Recognition (ASR) Model
This Automatic Speech Recognition (ASR) model transcribes Kashmiri speech from 16,000 KHz mono WAV audio files into text
NLP
Automatic Speech Recognition
Kashmiri
Speech-to-Text
  • See Upvoters0
  • Downloads28
  • File Size0
  • Views967
Updated 1 year(s) ago

AI4BHARAT

AI4Bharat - Romansetu-200M -Multilingual LLM for Indian langauges using romanization
RomanSetu is Efficiently unlocking multilingual (Indian Languages) capabilities of Large Language Models via Romanization.
Instruction-Tuning
Multilingual
LLaMA2
Llama
  • See Upvoters0
  • Downloads6
  • File Size0
  • Views337
Updated 1 year(s) ago

AI4BHARAT

AI4Bharat - Romansetu-100M - Multilingual LLM for Indian langauges using romanization
RomanSetu is Efficiently unlocking multilingual (Indian Languages) capabilities of Large Language Models via Romanization.
Multilingual
Llama
Instruction-Tuning
LLaMA2
  • See Upvoters0
  • Downloads11
  • File Size0
  • Views641
Updated 1 year(s) ago

AI4BHARAT

AI4Bharat- Kannada - IndicConformer Automatic Speech Recognition (ASR) Model
This Kannada Automatic Speech Recognition (ASR) model transcribes 16kHz mono-channel audio into text. It utilizes a Conformer-Large architecture with 120M parameters and a hybrid CTC-RNNT decoder for high-accuracy speech recognition.
NLP
Automatic Speech Recognition
Audio Processing
  • See Upvoters0
  • Downloads25
  • File Size0
  • Views966
Updated 1 year(s) ago

AI4BHARAT

AI4Bharat – Romanized Path – Base to Supervised Fine-Tuning (SFT)
Romansetu model is built on base pretrained model which is supervised fine tuned on instuction-following tasks using romanized Indian languages.
Llama
Instruction-Tuning
Multilingual
LLaMA2
  • See Upvoters0
  • Downloads3
  • File Size0
  • Views251
Updated 1 year(s) ago

AI4BHARAT

AI4Bharat-IndicTrans2 Large-1B -English-to-Hindi (Devanagari) – : Language Translation Model
A large-scale neural machine translation (NMT) model for translating English to Hindi (Devanagari) language, leveraging 1 billion parameters for high-quality translations.
NLP
Large Model
high-quality-translation
low-resource-NLP
cross-lingual
Multilingual
Machine Translation
Transformer
  • See Upvoters0
  • Downloads35
  • File Size0
  • Views1,124
Updated 1 year(s) ago

AI4BHARAT