
Sanskrit Tibetan Parallel Dataset
SansTib is a large-scale sentence-aligned parallel corpus of Sanskrit and Classical Tibetan Buddhist texts containing approximately 317,289 sentence pairs. The dataset also includes manually aligned evaluation datasets and a bilingual sentence embedding model for cross-lingual language understanding.
The Purpose Of This Dataset Is To Provide A Large Sentence-aligned Sanskrit–classical Tibetan Parallel Corpus For Cross-lingual Language Understanding And Translation Research. It Supports The Development And Evaluation Of Machine Translation Systems, Bilingual Sentence Embedding Models, Multilingual Information Retrieval, Sentence Alignment Algorithms, Semantic Similarity Models, Cross-lingual Representation Learning, Digital Philology, And Computational Research On Buddhist Literature. The Dataset Can Further Be Used To Develop Multilingual Dictionaries, Retrieval Systems, Language Resources For Low-resource Classical Languages, And Other Downstream Nlp Applications
Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
© 2026 - Copyright AIKosh. All rights reserved.