Indian Flag
Government Of India
A-
A
A+
ORGANISATION
Santham-Anvaya-Parallel

Santham-Anvaya-Parallel

**Santham** is a high-quality, curated parallel corpus for Sanskrit-Tamil machine translation. This dataset also contains 1000 benchmark sentences.

About Dataset

Sanskrit poetry frequently relies on complex metrical order that hinder direct translation. This repository contains *anvaya* (prose-order reordered) data mapped to poetry to serve as an intermediate translation bridge. | *Anvaya* | 10,146 | Poetry data mapped to *anvaya* (reordered) data. | To cite this work, please use the following BibTeX entry: @inproceedings{s-etal-2026-santham, title = "Santham: A Curated {S}anskrit{--}{T}amil Dataset with Anvaya and Segmentation for Building and Evaluating Machine Translation", author = "S, Prasanna Venkatesh T and Shetye, Ketaki Mangesh and Arjunasamy, Vishnuraj and Sahu, Ayush Kumar and Krishnan, Sriram and Krishnamurthy, Parameswari", editor = "Satuluri, Pavankumar and Goyal, Pawan", booktitle = "Proceedings of the 8th International {S}anskrit Computational Linguistics Symposium", month = mar, year = "2026", address = "IIT Roorkee, Roorkee, India", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2026.iscls-1.5/", pages = "65--80" }

Purpose of Dataset

Translation Of Sanskrit Tamil Dataset Using Anvaya As Source.

Activity Overview Activity Overview

  • Downloads0
  • Downloads 15
  • File Size 1.32 MB
  • Views 199

Tags Tags

  • Tamil
  • Parallel Corpus
  • Sanskrit
  • parallel sentences
  • language:tam
  • language:san
  • anvaya
  • Sanskrit-Tamil

License Control License Control

Attribution 4.0 International (CC BY- 4.0)

santham-anvaya ( 1 directories )


Directory
santham-anvaya

2 files

Data Quality Score BetaData Quality Score Beta

Version Control Version Control

FolderVersion 1(1.32 MB)
  • Nagaraju V·4 month(s) ago
    • chevron_rightFolder
      santham-anvaya
      • chevron_rightFolder
        santham-anvaya