Indic-mobile is a 0.5B parameter language model built completely from scratch — no fine-tuning, no adapter on top of an existing checkpoint. Every weight was pretrained from zero, purpose-built for all 22 officially recognized Indian languages and designed for efficient deployment on mobile and edge devices.
Indic-mobile is a 0.5B parameter language model built completely from scratch — no fine-tuning, no adapter on top of an existing checkpoint. Every weight was pretrained from zero, purpose-built for all 22 officially recognized Indian languages and designed for efficient deployment on mobile and edge devices. Supported Languages Indic-mobile covers all 22 languages recognized under the 8th Schedule of the Indian Constitution: Assamese, Bengali, Bodo, Dogri, Gujarati, Hindi, Kannada, Kashmiri, Konkani, Maithili, Malayalam, Manipuri, Marathi, Nepali, Odia, Punjabi, Sanskrit, Santali, Sindhi, Tamil, Telugu, Urdu Why Indic-mobile? India has 1.4 billion people and 22 officially recognized languages — yet most language models were never built with this diversity in mind. Indic-mobile is designed to change that: Built from scratch — not a fine-tune or adapter on an existing English-centric model Truly multilingual — trained across all 22 Indian languages from the ground up Mobile-first — 0.5B parameters means it runs efficiently on edge devices and smartphones Open source — weights, architecture, and everything else, freely available Model Architecture Architecture: Custom (trained from scratch) Parameters: 0.5B Precision: BF16 Training: Pretrained from scratch (no base model used) Objective: Causal language modeling across 22 Indic languages Intended Uses Direct Use Text generation in any of the 22 official Indian languages Multilingual Indic chatbots and assistants On-device / mobile NLP applications Low-resource language research and experimentation Downstream Use Fine-tuning for specific Indic language tasks (classification, summarization, translation, QA) Integration into larger Indic NLP pipelines RAG (Retrieval-Augmented Generation) systems for Indian language content Out-of-Scope Use High-stakes decision making without human oversight Generation of harmful, misleading, or abusive content in any language Tasks requiring deep factual accuracy without verification Bias, Risks, and Limitations As a small 0.5B model, it may struggle with complex reasoning or long-form generation compared to larger models Training data distribution across all 22 languages may not be perfectly balanced; lower-resource languages may underperform Like all language models, it may reflect biases present in the training data Not intended for use in safety-critical or high-stakes applications without further evaluation and fine-tuning Recommendations Users should evaluate the model on their specific use case and language before deployment, particularly for lower-resource Indic languages.
Apache 2.0
Purushottam Kumar
Multimodal Language Model
Transformers
Open
Social
26/07/26 12:40:25
953.22 MB
To preview this file, you need to be a registered user. Please complete the registration process to gain access and continue viewing the content.
Apache 2.0
© 2026 - Copyright AIKosh. All rights reserved.