Video developed by or in collaboration with ICAR institutes to all stakeholders
This dataset contains all the videos developed by or in collaboration with ICAR institutes to all stakeholders. Meta data 4200+ videos is available at present. It also contains the URL of the video through which video can be downloaded/accessed. The videoa are available in different languages such as Assamese, Bengali, English, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, Telugu. The dataset is intended to support AI use cases including summarisation, metadata generation, transcription to different languages, enabling researchers, developers, and agricultural extension workers to build multilingual recommendation engines, searchable video libraries, and NLP-based farmer chatbots. Each record includes title, organization, source URL, keywords, language, and subject domain. The dataset is released under the ICAR Data Use License.
This Dataset Is Designed To Enable The Development Of Multilingual Agricultural Ai Systems For India's Diverse Linguistic Landscape. It Supports Training And Fine-tuning Of Nlp, Small Language Models (Slms), And Large Language Models (Llms) For Tasks Such As Hindi-english-indic Language Agricultural Query Understanding, Semantic Video Retrieval, Metadata Enrichment, Topic Classification, Automated Tagging, Summarization, And Recommendation Engines. Researchers Can Build Agricultural Knowledge Graphs, Smart E-learning Platforms, And Ai-powered Agritech Solutions Focused On Soil Health, Sustainable Farming, Irrigation Management, And Farmer Education. The Dataset Is Also Suitable For Cross-lingual Information Retrieval, Agricultural Ontology Creation, And Domain-specific Model Evaluation. By Providing Metadata In 12+ Languages, It Directly Addresses The Language Barrier In Indian Digital Agriculture, Making Extension Content Accessible To Farmers In Their Native Tongues.
ICAR Data Use License
© 2026 - Copyright AIKosh. All rights reserved.