Indian Flag
Government Of India
A-
A
A+
ORGANISATION

NE-OCR

NE-OCR is a multilingual Optical Character Recognition model developed by MWire Labs to accurately recognize printed text from documents in Northeast Indian languages. The model supports Assamese, Bodo, English, Garo, Hindi, Khasi, Kokborok, Meitei (Bengali script), Meitei (Meitei Mayek script), Mizo, Nagamese, and Nyishi. It is designed to enable reliable digitization of books, newspapers, government records, educational materials, and cultural archives from Northeast India where mainstream OCR

About Model

NE-OCR is a multilingual OCR system built specifically for the linguistic diversity of Northeast India. Many global OCR engines struggle with regional scripts, mixed-language documents, and low-resource languages commonly used across the region. NE-OCR addresses this gap by providing accurate printed text recognition across multiple languages and scripts used in Northeast India. The model supports the following languages: Assamese Bodo English Garo Hindi Khasi Kokborok Meitei (Bengali script) Meitei (Meitei Mayek script) Mizo Nagamese Nyishi These languages represent several major writing systems used in the region including Latin, Bengali, Devanagari, and Meitei Mayek. Documents across Northeast India frequently contain mixed-language text due to administrative, educational, and cultural practices. NE-OCR is designed to handle such multilingual documents effectively. The model is trained using a diverse dataset of printed text collected from books, newspapers, scanned documents, educational publications, and institutional records from across the region. The system uses modern transformer-based OCR architecture combining visual feature extraction with sequence recognition models to generate accurate text outputs. NE-OCR was developed by MWire Labs as part of a broader effort to build AI infrastructure for indigenous and low-resource languages of Northeast India. The goal is to enable governments, researchers, startups, and institutions to digitize documents, preserve cultural knowledge, and build language technologies on top of reliable OCR systems. Benchmark Performance NE-OCR was evaluated against widely used OCR systems including EasyOCR, Tesseract 5, TrOCR-large, and Chandra across multiple languages from Northeast India. The results show strong performance across different scripts used in the region. Selected benchmark examples: Language: Assamese Script: Bengali NE-OCR: 97.46% EasyOCR: 32.25% Tesseract 5: 8.79% TrOCR-large: 0.80% Chandra: 57.83% Language: Khasi Script: Latin NE-OCR: 98.85% EasyOCR: 77.78% Tesseract 5: 80.72% TrOCR-large: 93.22% Chandra: 94.15% Language: Meitei (Meitei Mayek) Script: Meitei Mayek NE-OCR: 95.56% EasyOCR: 2.50% Tesseract 5: 2.24% TrOCR-large: 2.45% Chandra: 2.57% Language: Mizo Script: Latin NE-OCR: 95.96% EasyOCR: 67.62% Tesseract 5: 68.44% TrOCR-large: 84.58% Chandra: 92.96% Language: Hindi Script: Devanagari NE-OCR: 97.69% EasyOCR: 49.54% Tesseract 5: 41.48% TrOCR-large: 1.27% Chandra: 85.78% Across the full benchmark covering twelve languages, NE-OCR achieved an average recognition accuracy of 94.99 percent, significantly outperforming commonly used OCR systems on regional languages. Typical use cases include digitization of historical archives, processing of government records, OCR for newspapers and educational materials, creation of datasets for language technology research, and document search or knowledge extraction systems. NE-OCR is designed to serve as a foundational component for document AI systems in Northeast India. By enabling reliable OCR across regional languages and scripts, the model helps unlock large volumes of printed knowledge that remain difficult to digitize with existing OCR tools. The project is developed and maintained by MWire Labs as part of its mission to advance artificial intelligence technologies for the languages of Northeast India.

NE-OCR

Metadata Metadata

Attribution 4.0 International (CC BY- 4.0)

MWirelabs

Transformers

PyTorch

Open

MWire Labs

Sector Agnostic

07/03/26 12:14:44

Badal Nyalang

0

Activity Overview Activity Overview

  • Downloads0
  • Redirect 8
  • File Size 0
  • Views 203

Tags Tags

  • OCR
  • northeast-india
  • doctr
  • vitstr
  • Mizo
  • Garo
  • khasi
  • Nyishi
  • Kokborok
  • Nagamese
  • BODO
  • Meitei
  • Optical Character Recognition
  • Multilingual OCR
  • Northeast India OCR
  • Printed Text Recognition

License Control License Control

Attribution 4.0 International (CC BY- 4.0)

Related Models Related Models

Northeast Language Identification
NE-LID is a fast and accurate language identification model for Northeast Indian languages using character level features. It is designed for low resource and script diverse text and achieves high accuracy on short sentences.
language identification
fasttext
northeast-india
low-resource
Multilingual
MWire Labs
fastText
  • See Upvoters1
  • Downloads13
  • File Size0
  • Views867
Updated 6 month(s) ago

MWIRE LABS

Related Articles

Building Sovereign Language AI for Northeast India: The NE-Stack Story
Northeast India is home to over 220 languages spoken by 45 million people, yet remains mostly absent from modern AI systems. MWire Labs, a northeast India AI startup based in Shillong, Meghalaya, is building the NE-Stack, a foundational suite of northeast AI models covering speech recognition, machine translation, OCR, TTS, and vision-language understanding across 11+ indigenous languages. This article documents the technical and strategic approach behind northeast AI development for one of the world's most linguistically diverse and digitally underserved regions.
Building Sovereign Language AI for Northeast India: The NE-Stack Story
Badal NyalangBadal Nyalang
  • See Upvoters1
  • Read Time1 min read

More Models from MWire Labs More Models from MWire Labs

Northeast STT Multilingual Speech to Text Model
A multilingual Speech-to-Text (STT) model for eight Northeast Indian languages, fine-tuned from Whisper Medium using over 150,000 speech-text pairs from public and institutional datasets. The model expands speech recognition support for low-resource indigenous languages, including Khasi, Garo, Mizo, Kokborok, Nagamese, Assamese, Chakma, and Wancho.
low-resource
northeast-india
Northeast India Languages
whisper
Speech processing
Multilingual speech
Speech to Text
Multilingual
Automatic Speech Recognition
  • See Upvoters0
  • Downloads0
  • File Size0
  • Views41
Updated 22 day(s) ago

MWIRE LABS

NE-SpeechEmbed
NE-SpeechEmbed is a multilingual speech-text embedding model by MWire Labs for Northeast Indian languages. The model supports semantic speech search, cross-modal retrieval, and audio-text embeddings across Khasi, Garo, Mizo, Nagamese, Kokborok, Assamese, Wancho, and Chakma.
speech-text-retrieval
speech-embeddings
speech-language-model
mwire-labs
embeddings
whisper
low-resource-languages
northeast-india
retrieval
low-resource
XLM-RoBERTa
speech
multilingual
audio-text-retrieval
  • See Upvoters0
  • Downloads0
  • File Size0
  • Views23
Updated 25 day(s) ago

MWIRE LABS

NE-Embed
NE-Embed is a multilingual text embedding model for Northeast Indian languages, enabling semantic search, retrieval, and RAG across 10 languages including Khasi, Garo, Meitei, Bodo, Mizo, Assamese, Nyishi, Kokborok, Pnar, and Nagamese. Fine-tuned on LaBSE with 201,738 parallel pairs.
Multilingual
embeddings
northeast-india
sentence-transformers
retrieval
low-resource
rag
  • See Upvoters0
  • Downloads0
  • File Size0
  • Views25
Updated 25 day(s) ago

MWIRE LABS

Mizo OCR - Text Recognition for Mizo Language
OCR model for the Mizo language achieving 90.68% character accuracy on synthetic and curated printed text
Image-to-Text
trocr
Mizo
northeast-india
low-resource
OCR
  • See Upvoters0
  • Downloads2
  • File Size0
  • Views303
Updated 4 month(s) ago

MWIRE LABS

NE-OCR
NE-OCR is a multilingual Optical Character Recognition model developed by MWire Labs to accurately recognize printed text from documents in Northeast Indian languages. The model supports Assamese, Bodo, English, Garo, Hindi, Khasi, Kokborok, Meitei (Bengali script), Meitei (Meitei Mayek script), Mizo, Nagamese, and Nyishi. It is designed to enable reliable digitization of books, newspapers, government records, educational materials, and cultural archives from Northeast India where mainstream OCR
northeast-india
khasi
Mizo
Optical Character Recognition
OCR
doctr
vitstr
Multilingual OCR
Northeast India OCR
Printed Text Recognition
BODO
Garo
Kokborok
Nyishi
Meitei
Nagamese
  • See Upvoters0
  • Downloads8
  • File Size0
  • Views204
Updated 4 month(s) ago

MWIRE LABS

Nagamese Speech-to-Text
Automatic Speech Recognition (ASR) model for Nagamese speech, designed to transcribe spoken Nagamese into text for real-world usage.
Nagamese
Automatic Speech Recognition
ASR
Speech Recognition
whisper
low-resource-language
  • See Upvoters0
  • Downloads3
  • File Size0
  • Views209
Updated 4 month(s) ago

MWIRE LABS

Garo OCR - Text Recognition for Garo
OCR model for the Garo language achieving 93.13% character accuracy.
florence-2
northeast-india
Garo
OCR
Image-to-Text
  • See Upvoters0
  • Downloads0
  • File Size0
  • Views269
Updated 4 month(s) ago

MWIRE LABS

Northeast Language Identification
NE-LID is a fast and accurate language identification model for Northeast Indian languages using character level features. It is designed for low resource and script diverse text and achieves high accuracy on short sentences.
language identification
fastText
Multilingual
fasttext
MWire Labs
northeast-india
low-resource
  • See Upvoters1
  • Downloads13
  • File Size0
  • Views867
Updated 6 month(s) ago

MWIRE LABS

NortheastNER
NortheastNER is a token classification model built on XLM-RoBERTa and fine-tuned on ~25k sentences from gazetteers, news, and cultural texts across Northeast India. It detects region-specific entities, places, tribes, festivals, tourist sites, flora, fauna, and experimental local names; ideal for low-resource NER, regional search, cultural analytics, and knowledge graph applications.
Token Classification
Conservation
Northeast India
Meghalaya
northeast-india
low-resource
XLM-RoBERTa
NER
  • See Upvoters0
  • Downloads14
  • File Size0
  • Views366
Updated 8 month(s) ago

MWIRE LABS

Kren-M
Northeast India's first AI language model. Kren-M is a 2.6B parameter bilingual model for Khasi-English, built on Gemma-2-2B. Features Kren-NE custom tokenizer covering 7 NE languages (Khasi, Garo, Mizo, Assamese, Manipuri, Nagamese, Nyishi) with 35.7% efficiency gain. Trained on 5.43M Khasi sentences. Capabilities: bidirectional translation, natural conversation, cultural context. Designed for language preservation across Northeast India
northeast-india
Garo
low-resource
Instruction-Tuning
Indian Languages
Tokenizer
Foundational model
Northeast India Languages
Kren-M
bilingual
continued-pretraining
Northeast India
khasi
  • See Upvoters0
  • Downloads38
  • File Size0
  • Views1,022
Updated 8 month(s) ago

MWIRE LABS