Indian Flag
Government Of India
A-
A
A+
ORGANISATION

Bhashini - Fastspeech2 Model using (HS)

Text-to-speech models trained using FastPitch and HiFi-GAN vocoder, separately for each language. Supports both 'female' and 'male' voices.

  • See Upvoters0
  • Downloads113
  • File Size286.72 MB
  • Views2,205

About Model

This repository contains a Fastspeech2 Model for 16 Indian languages (male and female both) implemented using the Hybrid Segmentation (HS) for speech synthesis. The model is capable of generating mel-spectrograms from text inputs and can be used to synthesize speech. 
Fs2 is composed of 6 feed-forward Transformer blocks with multi-head self-attention and 1D convolution on both phoneme encoder and mel-spectrogram decoder. In each feed-forward Transformer, the hidden size of multi-head attention is set to 256 and the number of head is set to 2. The kernel size of 1D convolution in the two-layer convolution network is set to 9 and 1, and the input/output size of the number of channels in the first and the second layer is 256/1024 and 1024/256. The duration predictor and variance adaptor, which are composed of stacks of several convolution networks and the final linear projection layer. The convolution layers of the duration predictor and variance adaptor are set to 2 and 5, the kernel size is set to 3, the input/output size of all layers is 256/256, and the dropout rate is set to 0.5.

Bhashini - Fastspeech2 Model using (HS)

Metadata Metadata

MIT

SMT Lab IIT Madras

Speech Synthesis (TTS) Model

Other

Open

Sector Agnostic

06/07/26 16:10:49

286.72 MB

Tags Tags

  • Multilingual
  • NLP
  • Text Processing
  • Transformer
  • Text to Speech
  • Language Detection

assamese ( 2 directories )


Directory
female

1 directories

Directory
male

1 directories

License Control License Control

MIT

Version Control Version Control

FolderVersion 2(286.72 MB)
  • admin·1 year(s) ago
    • chevron_rightFolder
      assamese
      • chevron_rightFolder
        female
      • chevron_rightFolder
        male
    • chevron_rightFolder
      bengali
    • chevron_rightFolder
      bodo
    • chevron_rightFolder
      charmap
    • chevron_rightFolder
      english
    • undefined
      .gitattributes
    • undefined
      api.py
    • undefined
      app.py
    • undefined
      environment.yml
    • undefined
      get_phone_mapped_python.py
    • more_horiz 25 more

More Models from Digital India BHASHINI Division More Models from Digital India BHASHINI Division

SPRING-INX-DATA2VEC-AQC-GUJARATI
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
ssl
Low-resource languages
SSL_finetunning
Data2vec_aqc
spring_lab
IITM
gujarati
  • See Upvoters0
  • Downloads2
  • File Size3.52 GB
  • Views32
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-HINDI
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
Low-resource languages
ssl
IITM
spring_lab
Data2vec_aqc
SSL_finetunning
hindi
  • See Upvoters0
  • Downloads1
  • File Size3.53 GB
  • Views31
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-MANIPURI
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
Manipuri
Low Resource Languages
SSL_finetunning
Data2vec_aqc
spring_lab
IITM
ssl
  • See Upvoters0
  • Downloads1
  • File Size3.52 GB
  • Views26
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-ASSAMESE
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text.The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
SSL_finetunning
Assamese
ssl
IITM
spring_lab
Data2vec_aqc
Low-resource languages
  • See Upvoters0
  • Downloads1
  • File Size3.52 GB
  • Views18
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

IndicXlit
A Transformer-based multilingual transliteration model
NLP
Language Modeling
Multilingual Translation
Machine Translation
Regional Languages
Indian Languages
transliteration
  • See Upvoters0
  • Downloads64
  • File Size3.94 MB
  • Views1,430
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

Indic Trans2
AI4Bharat's Indic-Trans-v2 is a multilingual Transformer (~1.1BM) NMT model trained on Samanantar v2 dataset which is the largest publicly available parallel corpora collection for languages of India at the time of writing (23 March 2023). We currently release two models - Indic to English and English to Indic and support all the 22 scheduled languages of India.
Regional Languages
NLP
Indic-TransV2
Indian Languages
Machine Translation
Multilingual Translation
Bilingual Translation
Language Modeling
Computational Linguistics
Machine Translation
  • See Upvoters1
  • Downloads95
  • File Size214.60 KB
  • Views2,880
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

Bhashini - Fastspeech2 Model using (HS)
Text-to-speech models trained using FastPitch and HiFi-GAN vocoder, separately for each language. Supports both 'female' and 'male' voices.
Transformer
Text to Speech
Language Detection
Multilingual
NLP
Text Processing
  • See Upvoters0
  • Downloads113
  • File Size286.72 MB
  • Views2,206
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

Bhashini - IndicNER
IndicNER is a multilingual Named Entity Recognition model fine-tuned on 11 Indian languages to identify named entities in text
NLP
Foreigners
Multilingual
Transformer
Token Classification
Pytorch
Samanantar
Bert
NER
  • See Upvoters2
  • Downloads196
  • File Size591.28 MB
  • Views2,991
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

Bhashini-AI4Bharat Textual Language Detection v1.0
Detect language from provided text, Currently supports 23 languages (English, Bangla, Manipuri, Bodo, Konkani, Oriya, Nepali, Marathi, Sindhi, Sanskrit, Malayalam, Urdu, Assamese, Telugu, Dogri, Gujarati, Kashmiri, Punjabi, Santali, Maithili, Hindi, Tamil, Kannada)
NLP
Text Processing
Deep Learning
Transformer
Text Language Detection
Multilingual
AI4Bharat
Bhashini
  • See Upvoters5
  • Downloads291
  • File Size3 MB
  • Views5,683
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION

SPRING-INX-DATA2VEC-AQC-SANSKRIT
Automatic Speech Recognition (ASR) model for speech recognition, processing audio and transcribing spoken content into text. The inference code, installation requirements, and usage instructions are available in the SPRING Lab, IIT Madras GitHub repository: https://github.com/Speech-Lab-IITM/Fairseq-Inference
low-resource-language
SSL_finetunning
Data2vec_aqc
spring_lab
IITM
ssl
Sanskrit
  • See Upvoters0
  • Downloads5
  • File Size3.52 GB
  • Views227
Updated 1 month(s) ago

DIGITAL INDIA BHASHINI DIVISION