**Santham** is a high-quality, curated parallel corpus for Sanskrit-Tamil machine translation. Segmentation of potery data obtained from Sanskrit Heritage segmneter aligned with human translation.
Segmentation of potery data obtained from Sanskrit Heritage segmneter aligned with human translation. To cite this work, please use the following BibTeX entry: @inproceedings{s-etal-2026-santham, title = "Santham: A Curated {S}anskrit{--}{T}amil Dataset with Anvaya and Segmentation for Building and Evaluating Machine Translation", author = "S, Prasanna Venkatesh T and Shetye, Ketaki Mangesh and Arjunasamy, Vishnuraj and Sahu, Ayush Kumar and Krishnan, Sriram and Krishnamurthy, Parameswari", editor = "Satuluri, Pavankumar and Goyal, Pawan", booktitle = "Proceedings of the 8th International {S}anskrit Computational Linguistics Symposium", month = mar, year = "2026", address = "IIT Roorkee, Roorkee, India", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2026.iscls-1.5/", pages = "65--80" }
For Training Models Specifically With Segmented Version Of Poetry Text.
Attribution 4.0 International (CC BY- 4.0)
1 files
© 2026 - Copyright AIKosh. All rights reserved.