Indian Flag
Government Of India
A-
A
A+

Northeast Language Identification

NE-LID is a fast and accurate language identification model for Northeast Indian languages using character level features. It is designed for low resource and script diverse text and achieves high accuracy on short sentences.

About Model

NE LID is a sentence level language identification model developed for low resource languages of Northeast India. The model supports ten languages including Assamese, Bodo, Garo, Hindi, Khasi, Kokborok, Meitei, Mizo, Naga and Nyishi. The model is trained using a fastText supervised classifier with character n gram features which makes it robust to spelling variation short text and multiple writing scripts. A balanced dataset with two thousand sentences per language was used with a stratified train dev test split. Extensive evaluation shows that the model achieves around 99% accuracy and outperforms transformer based language models for this task. Experiments indicate that character level approaches are more effective than subword based transformers for language identification in low resource and script diverse settings. The model is suitable for use in language routing data filtering preprocessing for machine translation speech recognition and other downstream language technology applications in the Northeast India context.

Northeast Language Identification

Metadata Metadata

Attribution 4.0 International (CC BY- 4.0)

MWirelabs

Classification Model

Other

Open

MWire Labs

Social

14/01/26 15:04:29

Badal Nyalang

0

Activity Overview Activity Overview

  • Downloads1
  • Redirect 11
  • File Size 0
  • Views 444

Tags Tags

  • language identification
  • fasttext
  • northeast-india
  • low-resource
  • Multilingual
  • MWire Labs
  • fastText

License Control License Control

Attribution 4.0 International (CC BY- 4.0)

Related Models Related Models

NE-OCR
NE-OCR is a multilingual Optical Character Recognition model developed by MWire Labs to accurately recognize printed text from documents in Northeast Indian languages. The model supports Assamese, Bodo, English, Garo, Hindi, Khasi, Kokborok, Meitei (Bengali script), Meitei (Meitei Mayek script), Mizo, Nagamese, and Nyishi. It is designed to enable reliable digitization of books, newspapers, government records, educational materials, and cultural archives from Northeast India where mainstream OCR
OCR
northeast-india
doctr
vitstr
Mizo
Garo
khasi
Nyishi
Kokborok
Nagamese
BODO
Meitei
Optical Character Recognition
Multilingual OCR
Northeast India OCR
Printed Text Recognition
  • See Upvoters0
  • Downloads5
  • File Size0
  • Views60
Updated 29 day(s) ago

MWIRE LABS

NE-BERT
NE-BERT is Northeast India's first domain-specific multilingual foundation model. Built on the ModernBERT architecture and trained on 8.3 million sentences, it supports 9 regional languages: Assamese, Khasi, Garo, Manipuri (Meitei), Mizo, Nyishi, Nagamese, Kokborok, and Pnar. It achieves State-of-the-Art performance on regional benchmarks and offers 1.6x faster inference, bridging the digital divide for low-resource languages.
modernbert
Masked Language Modeling
northeast-india
low-resource-NLP
northeast bert
mwirelabs
token-efficiency
Assamese
Garo
Nyishi
Meitei
Nagamese
khasi
A'chik
Mizo
kokborok
Pnar
  • See Upvoters0
  • Downloads18
  • File Size0
  • Views471
Updated 4 month(s) ago

MWIRE LABS

More Models from MWire Labs More Models from MWire Labs

Mizo OCR - Text Recognition for Mizo Language
OCR model for the Mizo language achieving 90.68% character accuracy on synthetic and curated printed text
OCR
low-resource
Image-to-Text
trocr
northeast-india
Mizo
  • See Upvoters0
  • Downloads2
  • File Size0
  • Views112
Updated 29 day(s) ago

MWIRE LABS

NE-OCR
NE-OCR is a multilingual Optical Character Recognition model developed by MWire Labs to accurately recognize printed text from documents in Northeast Indian languages. The model supports Assamese, Bodo, English, Garo, Hindi, Khasi, Kokborok, Meitei (Bengali script), Meitei (Meitei Mayek script), Mizo, Nagamese, and Nyishi. It is designed to enable reliable digitization of books, newspapers, government records, educational materials, and cultural archives from Northeast India where mainstream OCR
Mizo
Garo
khasi
Nyishi
Kokborok
Nagamese
Printed Text Recognition
Northeast India OCR
Multilingual OCR
Optical Character Recognition
Meitei
BODO
OCR
northeast-india
doctr
vitstr
  • See Upvoters0
  • Downloads5
  • File Size0
  • Views60
Updated 29 day(s) ago

MWIRE LABS

Nagamese Speech-to-Text
Automatic Speech Recognition (ASR) model for Nagamese speech, designed to transcribe spoken Nagamese into text for real-world usage.
ASR
Speech Recognition
low-resource-language
Nagamese
whisper
Automatic Speech Recognition
  • See Upvoters0
  • Downloads2
  • File Size0
  • Views63
Updated 29 day(s) ago

MWIRE LABS

Garo OCR - Text Recognition for Garo
OCR model for the Garo language achieving 93.13% character accuracy.
florence-2
Garo
northeast-india
Image-to-Text
OCR
  • See Upvoters0
  • Downloads0
  • File Size0
  • Views105
Updated 29 day(s) ago

MWIRE LABS

Northeast Language Identification
NE-LID is a fast and accurate language identification model for Northeast Indian languages using character level features. It is designed for low resource and script diverse text and achieves high accuracy on short sentences.
fasttext
fastText
language identification
MWire Labs
Multilingual
low-resource
northeast-india
  • See Upvoters1
  • Downloads11
  • File Size0
  • Views445
Updated 3 month(s) ago

MWIRE LABS

NortheastNER
NortheastNER is a token classification model built on XLM-RoBERTa and fine-tuned on ~25k sentences from gazetteers, news, and cultural texts across Northeast India. It detects region-specific entities, places, tribes, festivals, tourist sites, flora, fauna, and experimental local names; ideal for low-resource NER, regional search, cultural analytics, and knowledge graph applications.
Northeast India
Token Classification
NER
northeast-india
low-resource
XLM-RoBERTa
Meghalaya
Conservation
  • See Upvoters0
  • Downloads9
  • File Size0
  • Views252
Updated 4 month(s) ago

MWIRE LABS

Kren-M
Northeast India's first AI language model. Kren-M is a 2.6B parameter bilingual model for Khasi-English, built on Gemma-2-2B. Features Kren-NE custom tokenizer covering 7 NE languages (Khasi, Garo, Mizo, Assamese, Manipuri, Nagamese, Nyishi) with 35.7% efficiency gain. Trained on 5.43M Khasi sentences. Capabilities: bidirectional translation, natural conversation, cultural context. Designed for language preservation across Northeast India
bilingual
Instruction-Tuning
continued-pretraining
low-resource
northeast-india
khasi
Tokenizer
Foundational model
Northeast India Languages
Kren-M
Northeast India
Indian Languages
Garo
  • See Upvoters0
  • Downloads24
  • File Size0
  • Views624
Updated 4 month(s) ago

MWIRE LABS

NE-BERT
NE-BERT is Northeast India's first domain-specific multilingual foundation model. Built on the ModernBERT architecture and trained on 8.3 million sentences, it supports 9 regional languages: Assamese, Khasi, Garo, Manipuri (Meitei), Mizo, Nyishi, Nagamese, Kokborok, and Pnar. It achieves State-of-the-Art performance on regional benchmarks and offers 1.6x faster inference, bridging the digital divide for low-resource languages.
Pnar
modernbert
Masked Language Modeling
northeast-india
low-resource-NLP
northeast bert
mwirelabs
token-efficiency
Assamese
Garo
Nyishi
Meitei
Nagamese
khasi
A'chik
Mizo
kokborok
  • See Upvoters0
  • Downloads18
  • File Size0
  • Views471
Updated 4 month(s) ago

MWIRE LABS

KhasiBERT
Khasi language model trained on 3.6M sentences using RoBERTa architecture. 110M parameters. Supports NLP tasks for Khasi text processing.
Meghalaya
roberta
Fill-Mask
khasi
Bert
masked-lm
foundational-model
low-resource
Indian Language
austroasiatic
kha
autotrain_compatible
endpoints_compatible
region:us
digital-india
safetensors
  • See Upvoters1
  • Downloads22
  • File Size0
  • Views732
Updated 7 month(s) ago

MWIRE LABS

Khasi English Semantic Search Model
Khasi-English semantic search model, trained on 66,794 pairs with 0.69-0.74 similarity. ~90MB, supports Meghalaya tourism/culture. By MWirelabs
khasi-culture
text-embeddings-inference
autotrain_compatible
license:cc0-1.0
kha
en
Sentence Similarity
cross-lingual
semantic search
khasi
safetensors
sentence-transformers
Meghalaya
  • See Upvoters0
  • Downloads23
  • File Size0
  • Views651
Updated 7 month(s) ago

MWIRE LABS