Indian Flag
Government Of India
A-
A
A+
ORGANISATION

Mizomade-Mizo-EN-NLLB-S1

Fine-tuned pre trained model for Mizo-English machine translation, trained on Mizo-English parallel datasets to improve translation performance for the Mizo language.

About Model

Mizomade-Mizo-EN-S1 is a fine-tuned NLLB-200 distilled 600M neural machine translation model developed for Mizo-English translation. The model is based on Meta AI's facebook/nllb-200-distilled-600M model and was fine-tuned on Mizo-English parallel translation data to improve translation quality for the Mizo language, a low-resource language. The model supports Mizo-English and English-Mizo machine translation and is intended primarily for research, evaluation, and development of Mizo language technologies. Model Architecture: NLLB-200 / M2M100ForConditionalGeneration Base Model: facebook/nllb-200-distilled-600M Task: Machine Translation Language Pair: Mizo ↔ English Model Size: +600M parameters Fine-tuning: Supervised fine-tuning on Mizo-English parallel data License: CC-BY-NC-4.0 Training hyperparameters The following hyperparameters were used during training: learning_rate: 2e-05 train_batch_size: 4 eval_batch_size: 8 seed: 42 gradient_accumulation_steps: 8 total_train_batch_size: 32 optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments lr_scheduler_type: cosine num_epochs: 3 label_smoothing_factor: 0.1

Mizomade-Mizo-EN-NLLB-S1

Metadata Metadata

Attribution 4.0 International (CC BY- 4.0)

C VANLALDINPUIA, C LALHMANGAIHA, C LALRUATDIKA

Machine Translation Model

Transformers

Restricted

Governance and Administration

07/10/26 05:59:51

1.18 GB

config.json ( 820 Bytes )


To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.

Activity Overview Activity Overview

  • Downloads0
  • Downloads 0
  • File Size 1.18 GB
  • Views 5

Tags Tags

  • Machine Translation
  • Mizoram
  • Mizo

License Control License Control

Attribution 4.0 International (CC BY- 4.0)

Version Control Version Control

FolderVersion 1(1.18 GB)
  • admin·2 day(s) ago
    • application/json
      config.json
    • application/json
      generation_config.json
    • undefined
      model.safetensors
    • text/markdown
      README.md
    • application/json
      tokenizer_config.json
    • application/json
      tokenizer.json
    • undefined
      training_args.bin

Related Datasets Related Datasets

Updated 3 month(s) ago
Tawng Mizo English Translation
Tawng Mizo English Translation
Information
Tawng Mizo–English Translation Dataset is a high-quality parallel corpus designed to support machine translation, multilingual AI, natural language processing (NLP), and language preservation efforts for the Mizo language. The dataset contains carefully curated sentence pairs aligned between Mizo and English, covering a wide range of domains, writing styles, and real-world usage scenarios.
Mizoram
Mizo
Machine translation
Cultural Preservation
  • See Upvoters0
  • Downloads1
  • File Size12.78 MB
  • Views173

DIGITAL INDIA BHASHINI DIVISION

More Models from Digital India BHASHINI Division More Models from Digital India BHASHINI Division

COILD Multilingual Translation
COILD Multilingual Translation Model (COILD-Mul-MT)
Translation
iit-patna
coild
Indic Languages
  • See Upvoters0
  • Downloads0
  • File Size0
  • Views4
Updated 1 day(s) ago

DIGITAL INDIA BHASHINI DIVISION

Mizomade-Mizo-EN-NLLB-S1
Fine-tuned pre trained model for Mizo-English machine translation, trained on Mizo-English parallel datasets to improve translation performance for the Mizo language.
Machine Translation
Mizoram
Mizo
  • See Upvoters0
  • Downloads0
  • File Size1.18 GB
  • Views6
Updated 1 day(s) ago

DIGITAL INDIA BHASHINI DIVISION

BHASHINI IISc Sourashtra VITS TTS Models
VITS text-to-speech models for Sourashtra, a low-resource Indo-Aryan language. There are male and female voices, and each voice accepts text in either Tamil script or Sourashtra script.
low-resource
tts
sourashtra
Text to Speech
VITS
  • See Upvoters0
  • Downloads2
  • File Size0
  • Views54
Updated 8 day(s) ago

DIGITAL INDIA BHASHINI DIVISION

MIZO LANGUAGE OCR
This dataset is designed for training and evaluating Optical Character Recognition (OCR) models capable of extracting text from document images. It contains 14,000 labeled samples consisting of image filenames and their corresponding ground-truth text annotations, along with metadata such as font type, font size, and document category. The dataset includes diverse document styles, including printed books and scanned documents, enabling models to learn robust text recognition under varying visual
Mizo
Mizoram
  • See Upvoters0
  • Downloads1
  • File Size93.73 MB
  • Views27
Updated 23 day(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-GUJARATI
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
IITM
ssl
gujarati
SSL_finetunning
Low-resource languages
Data2vec_aqc
spring_lab
  • See Upvoters0
  • Downloads4
  • File Size3.52 GB
  • Views44
Updated 2 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-HINDI
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
IITM
Low-resource languages
hindi
SSL_finetunning
Data2vec_aqc
spring_lab
ssl
  • See Upvoters0
  • Downloads2
  • File Size3.53 GB
  • Views44
Updated 2 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-MANIPURI
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
Manipuri
Low Resource Languages
SSL_finetunning
Data2vec_aqc
spring_lab
IITM
ssl
  • See Upvoters0
  • Downloads2
  • File Size3.52 GB
  • Views39
Updated 2 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-ASSAMESE
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
Data2vec_aqc
spring_lab
IITM
ssl
Assamese
Low-resource languages
SSL_finetunning
  • See Upvoters0
  • Downloads3
  • File Size3.52 GB
  • Views27
Updated 2 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

IndicXlit
A Transformer-based multilingual transliteration model
transliteration
Language Modeling
Multilingual Translation
Machine Translation
Regional Languages
Indian Languages
NLP
  • See Upvoters0
  • Downloads69
  • File Size3.94 MB
  • Views1,594
Updated 3 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

Indic Trans2
AI4Bharat's Indic-Trans-v2 is a multilingual Transformer (~1.1BM) NMT model trained on Samanantar v2 dataset which is the largest publicly available parallel corpora collection for languages of India at the time of writing (23 March 2023). We currently release two models - Indic to English and English to Indic and support all the 22 scheduled languages of India.
NLP
Machine Translation
Language Modeling
Bilingual Translation
Multilingual Translation
Machine Translation
Regional Languages
Indian Languages
Indic-TransV2
Computational Linguistics
  • See Upvoters1
  • Downloads99
  • File Size214.60 KB
  • Views3,265
Updated 3 month(s) ago

DIGITAL INDIA BHASHINI DIVISION