
Sanskrit Tibetan Parallel Dataset
SansTib is a large-scale sentence-aligned parallel corpus of Sanskrit and Classical Tibetan Buddhist texts containing approximately 317,289 sentence pairs. The dataset also includes manually aligned evaluation datasets and a bilingual sentence embedding model for cross-lingual language understanding.
The purpose of this dataset is to provide a large sentence-aligned Sanskrit–Classical Tibetan parallel corpus for cross-lingual language understanding and translation research. It supports the development and evaluation of machine translation systems, bilingual sentence embedding models, multilingual information retrieval, sentence alignment algorithms, semantic similarity models, cross-lingual representation learning, digital philology, and computational research on Buddhist literature. The dataset can further be used to develop multilingual dictionaries, retrieval systems, language resources for low-resource classical languages, and other downstream NLP applications
Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
© 2026 - Copyright AIKosh. All rights reserved.