Indian Flag
Government Of India
A-
A
A+
ORGANISATION

KhasiBERT

Khasi language model trained on 3.6M sentences using RoBERTa architecture. 110M parameters. Supports NLP tasks for Khasi text processing.

About Model

KhasiBERT is the first transformer-based language model developed specifically for the Khasi language of Meghalaya. This foundational model addresses the computational linguistics needs for Meghalaya's 1.4+ million Khasi speakers and enables digital language processing infrastructure for the state.

The model implements the RoBERTa architecture with 12 transformer layers, 768 hidden dimensions, and 12 attention heads, totaling 110,652,416 parameters. KhasiBERT was trained using masked language modeling on a corpus of 3,621,116 Khasi sentences collected and preprocessed to represent diverse linguistic patterns of the language.

A custom Byte-Level BPE tokenizer with 32,000 vocabulary tokens was developed specifically for the Khasi language to handle its Austroasiatic linguistic characteristics effectively. Training utilized mixed-precision (FP16) optimization with a batch size of 24, learning rate of 5e-5, and AdamW optimizer over 150,880 training steps.

The model supports standard transformer operations including masked token prediction, contextualized embeddings generation, and can be fine-tuned for downstream tasks such as text classification, sentiment analysis, named entity recognition, and question answering in Khasi. Input sequences are limited to 512 tokens with standard special tokens for beginning-of-sequence, end-of-sequence, padding, unknown tokens, and masking.

Technical specifications include GELU activation functions, 0.1 dropout probability, layer normalization with epsilon 1e-12, and positional embeddings up to 514 positions. The model weights are distributed in Safetensors format and are compatible with the Transformers library ecosystem for integration with existing NLP workflows.

KhasiBERT

Metadata Metadata

Creative Commons Attribution Non Commercial 4.0

MWirelabs

Text Generation

PyTorch

Open

MWire Labs

Education and Skill Development

04/09/25 02:55:59

Badal Nyalang

0

Activity Overview Activity Overview

  • Downloads1
  • Redirect 33
  • File Size 0
  • Views 1,217

Tags Tags

  • Bert
  • safetensors
  • endpoints_compatible
  • autotrain_compatible
  • Fill-Mask
  • Indian Language
  • low-resource
  • region:us
  • kha
  • khasi
  • Meghalaya
  • austroasiatic
  • masked-lm
  • roberta
  • foundational-model
  • digital-india

License Control License Control

Creative Commons Attribution Non Commercial 4.0

Related Models Related Models

Kren-M
Northeast India's first AI language model. Kren-M is a 2.6B parameter bilingual model for Khasi-English, built on Gemma-2-2B. Features Kren-NE custom tokenizer covering 7 NE languages (Khasi, Garo, Mizo, Assamese, Manipuri, Nagamese, Nyishi) with 35.7% efficiency gain. Trained on 5.43M Khasi sentences. Capabilities: bidirectional translation, natural conversation, cultural context. Designed for language preservation across Northeast India
Tokenizer
low-resource
Garo
northeast-india
khasi
Northeast India
continued-pretraining
bilingual
Kren-M
Northeast India Languages
Foundational model
Indian Languages
Instruction-Tuning
  • See Upvoters0
  • Downloads40
  • File Size0
  • Views1,142
Updated 10 month(s) ago

MWIRE LABS

Kren v1: Khasi Generative Language Model
Kren v1 is the first Khasi generative language model, trained on 1M lines, pioneering encoder-to-decoder adaptation for low-resource AI.
Low-Resource NLP
khasi
khasi-culture
Meghalaya
Indigenous Language
Northeast India
Encoder-to-Decoder
AI for Culture
MWire Labs
Natural Language Processing
  • See Upvoters0
  • Downloads0
  • File Size390.67 MB
  • Views8
Updated 1 year(s) ago

MWIRE LABS

More Models from MWire Labs More Models from MWire Labs

Northeast STT Multilingual Speech to Text Model
A multilingual Speech-to-Text (STT) model for eight Northeast Indian languages, fine-tuned from Whisper Medium using over 150,000 speech-text pairs from public and institutional datasets. The model expands speech recognition support for low-resource indigenous languages, including Khasi, Garo, Mizo, Kokborok, Nagamese, Assamese, Chakma, and Wancho.
low-resource
northeast-india
Northeast India Languages
whisper
Speech processing
Multilingual speech
Speech to Text
Multilingual
Automatic Speech Recognition
  • See Upvoters0
  • Downloads2
  • File Size0
  • Views107
Updated 2 month(s) ago

MWIRE LABS

NE-SpeechEmbed
NE-SpeechEmbed is a multilingual speech-text embedding model by MWire Labs for Northeast Indian languages. The model supports semantic speech search, cross-modal retrieval, and audio-text embeddings across Khasi, Garo, Mizo, Nagamese, Kokborok, Assamese, Wancho, and Chakma.
speech-text-retrieval
speech-embeddings
speech-language-model
mwire-labs
embeddings
whisper
low-resource-languages
northeast-india
retrieval
low-resource
XLM-RoBERTa
speech
multilingual
audio-text-retrieval
  • See Upvoters0
  • Downloads0
  • File Size0
  • Views51
Updated 2 month(s) ago

MWIRE LABS

NE-Embed
NE-Embed is a multilingual text embedding model for Northeast Indian languages, enabling semantic search, retrieval, and RAG across 10 languages including Khasi, Garo, Meitei, Bodo, Mizo, Assamese, Nyishi, Kokborok, Pnar, and Nagamese. Fine-tuned on LaBSE with 201,738 parallel pairs.
Multilingual
embeddings
northeast-india
sentence-transformers
retrieval
low-resource
rag
  • See Upvoters0
  • Downloads0
  • File Size0
  • Views59
Updated 2 month(s) ago

MWIRE LABS

Mizo OCR - Text Recognition for Mizo Language
OCR model for the Mizo language achieving 90.68% character accuracy on synthetic and curated printed text
Image-to-Text
trocr
Mizo
northeast-india
low-resource
OCR
  • See Upvoters0
  • Downloads4
  • File Size0
  • Views398
Updated 6 month(s) ago

MWIRE LABS

NE-OCR
NE-OCR is a multilingual Optical Character Recognition model developed by MWire Labs to accurately recognize printed text from documents in Northeast Indian languages. The model supports Assamese, Bodo, English, Garo, Hindi, Khasi, Kokborok, Meitei (Bengali script), Meitei (Meitei Mayek script), Mizo, Nagamese, and Nyishi. It is designed to enable reliable digitization of books, newspapers, government records, educational materials, and cultural archives from Northeast India where mainstream OCR
northeast-india
khasi
Mizo
Optical Character Recognition
OCR
doctr
vitstr
Multilingual OCR
Northeast India OCR
Printed Text Recognition
BODO
Garo
Kokborok
Nyishi
Meitei
Nagamese
  • See Upvoters0
  • Downloads9
  • File Size0
  • Views269
Updated 6 month(s) ago

MWIRE LABS

Nagamese Speech-to-Text
Automatic Speech Recognition (ASR) model for Nagamese speech, designed to transcribe spoken Nagamese into text for real-world usage.
Nagamese
Automatic Speech Recognition
ASR
Speech Recognition
whisper
low-resource-language
  • See Upvoters0
  • Downloads3
  • File Size0
  • Views252
Updated 6 month(s) ago

MWIRE LABS

Garo OCR - Text Recognition for Garo
OCR model for the Garo language achieving 93.13% character accuracy.
florence-2
northeast-india
Garo
OCR
Image-to-Text
  • See Upvoters0
  • Downloads0
  • File Size0
  • Views340
Updated 6 month(s) ago

MWIRE LABS

Northeast Language Identification
NE-LID is a fast and accurate language identification model for Northeast Indian languages using character level features. It is designed for low resource and script diverse text and achieves high accuracy on short sentences.
language identification
fastText
Multilingual
fasttext
MWire Labs
northeast-india
low-resource
  • See Upvoters1
  • Downloads13
  • File Size0
  • Views1,072
Updated 8 month(s) ago

MWIRE LABS

NortheastNER
NortheastNER is a token classification model built on XLM-RoBERTa and fine-tuned on ~25k sentences from gazetteers, news, and cultural texts across Northeast India. It detects region-specific entities, places, tribes, festivals, tourist sites, flora, fauna, and experimental local names; ideal for low-resource NER, regional search, cultural analytics, and knowledge graph applications.
Token Classification
Conservation
Northeast India
Meghalaya
northeast-india
low-resource
XLM-RoBERTa
NER
  • See Upvoters0
  • Downloads15
  • File Size0
  • Views429
Updated 10 month(s) ago

MWIRE LABS

Kren-M
Northeast India's first AI language model. Kren-M is a 2.6B parameter bilingual model for Khasi-English, built on Gemma-2-2B. Features Kren-NE custom tokenizer covering 7 NE languages (Khasi, Garo, Mizo, Assamese, Manipuri, Nagamese, Nyishi) with 35.7% efficiency gain. Trained on 5.43M Khasi sentences. Capabilities: bidirectional translation, natural conversation, cultural context. Designed for language preservation across Northeast India
northeast-india
Garo
low-resource
Instruction-Tuning
Indian Languages
Tokenizer
Foundational model
Northeast India Languages
Kren-M
bilingual
continued-pretraining
Northeast India
khasi
  • See Upvoters0
  • Downloads40
  • File Size0
  • Views1,142
Updated 10 month(s) ago

MWIRE LABS