Skip to main navigation Skip to search Skip to main content

IL-ILGOV-2024: a translation benchmark for Hindi-to-12 languages in the governance domain

  • Vandan Mujadia*
  • , Rao B. Ashwath
  • , Dipti Misra Sharma
  • *Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

Abstract

Benchmark corpora for translation are crucial for assessing machine translation (MT) systems. These corpora allow researchers and developers to evaluate the performance, accuracy, and efficiency of their translation models. In this context, our work introduces a new translation benchmark for translations between Indian languages, with Hindi as the source language, spanning 12 languages (Assamese, Bangla, English, Gujarati, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, Telugu, and Urdu) in the governance domain. Translations are sourced online, automatically aligned, and then undergo human validation and correction for alignment and translation. Additional human validation, further refines these translations for benchmark use. The corpus comprises 1304 n-way parallel sentences and three bilingual sets (dev, devtest, and test), each containing 1000 to 3000 sentences, with Hindi serving as the source language. We offer these corpus to the research community for testing and evaluating MT systems in the governance domain and between Indian languages.

Original languageEnglish
Pages (from-to)3851-3872
Number of pages22
JournalLanguage Resources and Evaluation
Volume59
Issue number4
DOIs
Publication statusPublished - 12-2025

All Science Journal Classification (ASJC) codes

  • Language and Linguistics
  • Education
  • Linguistics and Language
  • Library and Information Sciences

Fingerprint

Dive into the research topics of 'IL-ILGOV-2024: a translation benchmark for Hindi-to-12 languages in the governance domain'. Together they form a unique fingerprint.

Cite this