Indian Flag
Government Of India
A-
A
A+
ORGANISATION

Khasi English Semantic Search Model

Khasi-English semantic search model, trained on 66,794 pairs with 0.69-0.74 similarity. ~90MB, supports Meghalaya tourism/culture. By MWirelabs

About Model

Developed by MWirelabs, this model is the first production-ready semantic search system for Khasi-English language pairs, celebrating Northeast India’s linguistic diversity, with a special focus on Meghalaya. Trained on a curated corpus of 66,794 English-Khasi translation pairs (63,909 Khasi sentences, 65,239 English sentences, 65,241 parallel pairs), it utilizes the lightweight MiniLM-L6-v2 architecture (~22.7M parameters, ~90MB). The model achieves cosine similarity scores of 0.69-0.74, showcasing effective cross-lingual alignment for Khasi, a low-resource Austroasiatic language spoken primarily in Meghalaya.

The dataset, sourced from cleaned Khasi texts, historical documents, bilingual translations, and cultural/administrative materials from Meghalaya, was preprocessed for anonymization.

Key use cases include cross-lingual document similarity, cultural content discovery (e.g., Meghalaya’s Khasi folklore), and educational tools for the region’s tourism and heritage sectors. The lightweight design supports deployment on low-resource devices, enhancing accessibility in Meghalaya. Ethical considerations emphasize respect for Khasi heritage, encouraging collaboration with Meghalaya’s local communities.

This pioneering effort by MWirelabs, released under Creative Commons CC0 1.0, positions the organization as a leader in Meghalaya and Northeast India’s AI innovation, building on the Khasi-English Word Embeddings model.

Citation: @misc{kajingiathuhsearch2025, title={KaJingïathuhSearch2025: Khasi-English Semantic Search Model}, author={MWirelabs}, year={2025}, publisher={Hugging Face}, howpublished={\url{https://huggingface.co/MWirelabs/khasi-english-semantic-search}} }

Khasi English Semantic Search Model

Metadata Metadata

CC0 1.0 Public Domain

MWirelabs

Multilingual Language Model

PyTorch

Open

MWire Labs

Arts, Culture and Tourism

18/09/25 10:17:07

Badal Nyalang

0

Activity Overview Activity Overview

  • Downloads0
  • Redirect 26
  • File Size 0
  • Views 1,020

Tags Tags

  • sentence-transformers
  • safetensors
  • khasi
  • semantic search
  • cross-lingual
  • Sentence Similarity
  • en
  • kha
  • license:cc0-1.0
  • autotrain_compatible
  • text-embeddings-inference
  • khasi-culture
  • Meghalaya

License Control License Control

CC0 1.0 Public Domain

Related Models Related Models

NE-Embed
NE-Embed is a multilingual text embedding model for Northeast Indian languages, enabling semantic search, retrieval, and RAG across 10 languages including Khasi, Garo, Meitei, Bodo, Mizo, Assamese, Nyishi, Kokborok, Pnar, and Nagamese. Fine-tuned on LaBSE with 201,738 parallel pairs.
sentence-transformers
embeddings
retrieval
northeast-india
low-resource
Multilingual
rag
  • See Upvoters0
  • Downloads0
  • File Size0
  • Views37
Updated 1 month(s) ago

MWIRE LABS

More Models from MWire Labs More Models from MWire Labs

Northeast STT Multilingual Speech to Text Model
A multilingual Speech-to-Text (STT) model for eight Northeast Indian languages, fine-tuned from Whisper Medium using over 150,000 speech-text pairs from public and institutional datasets. The model expands speech recognition support for low-resource indigenous languages, including Khasi, Garo, Mizo, Kokborok, Nagamese, Assamese, Chakma, and Wancho.
northeast-india
low-resource
Multilingual
Northeast India Languages
Speech processing
Multilingual speech
Speech to Text
whisper
Automatic Speech Recognition
  • See Upvoters0
  • Downloads2
  • File Size0
  • Views76
Updated 1 month(s) ago

MWIRE LABS

NE-SpeechEmbed
NE-SpeechEmbed is a multilingual speech-text embedding model by MWire Labs for Northeast Indian languages. The model supports semantic speech search, cross-modal retrieval, and audio-text embeddings across Khasi, Garo, Mizo, Nagamese, Kokborok, Assamese, Wancho, and Chakma.
northeast-india
embeddings
whisper
XLM-RoBERTa
speech
mwire-labs
speech-language-model
multilingual
audio-text-retrieval
speech-text-retrieval
speech-embeddings
low-resource-languages
retrieval
low-resource
  • See Upvoters0
  • Downloads0
  • File Size0
  • Views32
Updated 1 month(s) ago

MWIRE LABS

NE-Embed
NE-Embed is a multilingual text embedding model for Northeast Indian languages, enabling semantic search, retrieval, and RAG across 10 languages including Khasi, Garo, Meitei, Bodo, Mizo, Assamese, Nyishi, Kokborok, Pnar, and Nagamese. Fine-tuned on LaBSE with 201,738 parallel pairs.
sentence-transformers
rag
Multilingual
low-resource
northeast-india
retrieval
embeddings
  • See Upvoters0
  • Downloads0
  • File Size0
  • Views37
Updated 1 month(s) ago

MWIRE LABS

Mizo OCR - Text Recognition for Mizo Language
OCR model for the Mizo language achieving 90.68% character accuracy on synthetic and curated printed text
OCR
low-resource
Image-to-Text
trocr
northeast-india
Mizo
  • See Upvoters0
  • Downloads3
  • File Size0
  • Views345
Updated 5 month(s) ago

MWIRE LABS

NE-OCR
NE-OCR is a multilingual Optical Character Recognition model developed by MWire Labs to accurately recognize printed text from documents in Northeast Indian languages. The model supports Assamese, Bodo, English, Garo, Hindi, Khasi, Kokborok, Meitei (Bengali script), Meitei (Meitei Mayek script), Mizo, Nagamese, and Nyishi. It is designed to enable reliable digitization of books, newspapers, government records, educational materials, and cultural archives from Northeast India where mainstream OCR
khasi
Nyishi
Kokborok
Nagamese
OCR
Meitei
Optical Character Recognition
Multilingual OCR
Northeast India OCR
Printed Text Recognition
BODO
northeast-india
doctr
vitstr
Mizo
Garo
  • See Upvoters0
  • Downloads8
  • File Size0
  • Views221
Updated 5 month(s) ago

MWIRE LABS

Nagamese Speech-to-Text
Automatic Speech Recognition (ASR) model for Nagamese speech, designed to transcribe spoken Nagamese into text for real-world usage.
low-resource-language
Automatic Speech Recognition
whisper
Nagamese
Speech Recognition
ASR
  • See Upvoters0
  • Downloads3
  • File Size0
  • Views225
Updated 5 month(s) ago

MWIRE LABS

Garo OCR - Text Recognition for Garo
OCR model for the Garo language achieving 93.13% character accuracy.
Image-to-Text
northeast-india
Garo
florence-2
OCR
  • See Upvoters0
  • Downloads0
  • File Size0
  • Views298
Updated 5 month(s) ago

MWIRE LABS

Northeast Language Identification
NE-LID is a fast and accurate language identification model for Northeast Indian languages using character level features. It is designed for low resource and script diverse text and achieves high accuracy on short sentences.
fasttext
fastText
language identification
MWire Labs
Multilingual
low-resource
northeast-india
  • See Upvoters1
  • Downloads13
  • File Size0
  • Views930
Updated 7 month(s) ago

MWIRE LABS

NortheastNER
NortheastNER is a token classification model built on XLM-RoBERTa and fine-tuned on ~25k sentences from gazetteers, news, and cultural texts across Northeast India. It detects region-specific entities, places, tribes, festivals, tourist sites, flora, fauna, and experimental local names; ideal for low-resource NER, regional search, cultural analytics, and knowledge graph applications.
Token Classification
Northeast India
Conservation
Meghalaya
XLM-RoBERTa
low-resource
northeast-india
NER
  • See Upvoters0
  • Downloads14
  • File Size0
  • Views384
Updated 8 month(s) ago

MWIRE LABS

Kren-M
Northeast India's first AI language model. Kren-M is a 2.6B parameter bilingual model for Khasi-English, built on Gemma-2-2B. Features Kren-NE custom tokenizer covering 7 NE languages (Khasi, Garo, Mizo, Assamese, Manipuri, Nagamese, Nyishi) with 35.7% efficiency gain. Trained on 5.43M Khasi sentences. Capabilities: bidirectional translation, natural conversation, cultural context. Designed for language preservation across Northeast India
Instruction-Tuning
continued-pretraining
low-resource
northeast-india
khasi
Tokenizer
Foundational model
Northeast India Languages
Kren-M
Northeast India
Indian Languages
Garo
bilingual
  • See Upvoters0
  • Downloads40
  • File Size0
  • Views1,074
Updated 8 month(s) ago

MWIRE LABS