Indian Flag
Government Of India
A-
A
A+
ORGANISATION

AI4Bharat - MultiIndic WikiBio Structured Summarization Model

MultiIndicWikiBioSS is a multilingual sequence-to-sequence model fine-tuned on the IndicWikiBio dataset for nine Indian languages

About Model

MultiIndicWikiBioSS is a multilingual, sequence-to-sequence model fine-tuned from an IndicBARTSS checkpoint on the IndicWikiBio dataset, supporting biography generation across nine Indian languages: Hindi, Bengali, Punjabi, Tamil, and Telugu. Unlike models such as mBART50 and mT5, MultiIndicWikiBioSS retains native scripts, removing the need for script mapping to or from Devanagari. It is optimized for efficiency, being significantly smaller and computationally less expensive for fine-tuning and decoding than mBART and mT5-base. Trained on a dataset of 34,653 examples, the model provides strong multilingual capabilities tailored to Indic languages, making it a valuable tool for biography generation tasks.

AI4Bharat - MultiIndic WikiBio Structured Summarization Model

Metadata Metadata

MIT

Aman Kumar and Himani Shrotriya and Prachi Sahu and Raj Dabre and Ratish Puduppully and Anoop Kunchukuttan and Amogh Mishra and Mitesh M. Khapra and Pratyush Kumar

Summarization Model

N.A.

Open

AI4Bharat

Sector Agnostic

21/02/25 13:21:56

0

Activity Overview Activity Overview

  • Downloads0
  • Redirect 5
  • File Size 0
  • Views 139

Tags Tags

  • Transformers
  • Multilingual
  • NLP
  • Text2Text Generation

License Control License Control

MIT

More Models from AI4Bharat More Models from AI4Bharat

AI4Bharat- 500 M - RomanSetu Multilingual Native-to-Roman Model
RomanSetu is a multilingual continual pretrained transformer model designed for transliteration across six Indic languages
Multilingual
Instruction-Tuning
Llama
LLaMA2
  • See Upvoters1
  • Downloads53
  • File Size0
  • Views1,103
Updated 1 year(s) ago

AI4BHARAT

AI4Bharat- 400 M - RomanSetu Multilingual Native-to-Roman Model
RomanSetu is a multilingual continual pretrained transformer model designed for transliteration across six Indic languages
Multilingual
LLaMA2
Llama
Instruction-Tuning
  • See Upvoters1
  • Downloads94
  • File Size0
  • Views1,218
Updated 1 year(s) ago

AI4BHARAT

AI4Bharat- Maithili - IndicConformer Automatic Speech Recognition (ASR) Model
This model takes in mono-channel audio files at a 16,000 Hz sampling rate (WAV format) and outputs the transcribed text of the speech contained in the audio.
NLP
Automatic Speech Recognition
Speech-to-Text
  • See Upvoters0
  • Downloads29
  • File Size0
  • Views843
Updated 1 year(s) ago

AI4BHARAT

AI4Bharat- Konkani - IndicConformer Automatic Speech Recognition (ASR) Model
Automatic Speech Recognition (ASR) model for Konkani speech recognition, processing 16,000 KHz mono WAV audio and transcribing spoken content into text
Speech-to-Text
Automatic Speech Recognition
NLP
  • See Upvoters0
  • Downloads35
  • File Size0
  • Views850
Updated 1 year(s) ago

AI4BHARAT

AI4Bharat- Kashmiri - IndicConformer Automatic Speech Recognition (ASR) Model
This Automatic Speech Recognition (ASR) model transcribes Kashmiri speech from 16,000 KHz mono WAV audio files into text
NLP
Automatic Speech Recognition
Kashmiri
Speech-to-Text
  • See Upvoters0
  • Downloads28
  • File Size0
  • Views967
Updated 1 year(s) ago

AI4BHARAT

AI4Bharat - Romansetu-200M -Multilingual LLM for Indian langauges using romanization
RomanSetu is Efficiently unlocking multilingual (Indian Languages) capabilities of Large Language Models via Romanization.
Instruction-Tuning
Multilingual
LLaMA2
Llama
  • See Upvoters0
  • Downloads6
  • File Size0
  • Views337
Updated 1 year(s) ago

AI4BHARAT

AI4Bharat - Romansetu-100M - Multilingual LLM for Indian langauges using romanization
RomanSetu is Efficiently unlocking multilingual (Indian Languages) capabilities of Large Language Models via Romanization.
Multilingual
Llama
Instruction-Tuning
LLaMA2
  • See Upvoters0
  • Downloads11
  • File Size0
  • Views641
Updated 1 year(s) ago

AI4BHARAT

AI4Bharat- Kannada - IndicConformer Automatic Speech Recognition (ASR) Model
This Kannada Automatic Speech Recognition (ASR) model transcribes 16kHz mono-channel audio into text. It utilizes a Conformer-Large architecture with 120M parameters and a hybrid CTC-RNNT decoder for high-accuracy speech recognition.
NLP
Automatic Speech Recognition
Audio Processing
  • See Upvoters0
  • Downloads25
  • File Size0
  • Views966
Updated 1 year(s) ago

AI4BHARAT

AI4Bharat – Romanized Path – Base to Supervised Fine-Tuning (SFT)
Romansetu model is built on base pretrained model which is supervised fine tuned on instuction-following tasks using romanized Indian languages.
Llama
Instruction-Tuning
Multilingual
LLaMA2
  • See Upvoters0
  • Downloads3
  • File Size0
  • Views251
Updated 1 year(s) ago

AI4BHARAT

AI4Bharat-IndicTrans2 Large-1B -English-to-Hindi (Devanagari) – : Language Translation Model
A large-scale neural machine translation (NMT) model for translating English to Hindi (Devanagari) language, leveraging 1 billion parameters for high-quality translations.
NLP
Large Model
high-quality-translation
low-resource-NLP
cross-lingual
Multilingual
Machine Translation
Transformer
  • See Upvoters0
  • Downloads35
  • File Size0
  • Views1,124
Updated 1 year(s) ago

AI4BHARAT