DigestAI news desk

Cut through the AI noise.

Research

Researchers release COILD corpus for Indian language machine translation

A team of researchers has introduced COILD, a new parallel corpus and benchmark for machine translation between Indian languages. The dataset includes over 1.16 million human-translated and verified sentence pairs across 20 language pairs, covering families like Indo-Aryan, Dravidian, Tibeto-Burman, and Austro-Asiatic. The data comes from licensed repositories and spans eight real-world domains,…

1 source primary source

Key points

  • COILD corpus contains 1.16 million human-translated sentence pairs across 20 Indian language pairs
  • Data sourced from licensed repositories covering eight real-world domains like legal and medical texts
  • Fine-tuning IndicTrans2-Distilled and NLLB-200 models improved translation accuracy per researchers

The researchers also created a domain-centric benchmark of 2,000 expert-verified sentences for consistent evaluation. Testing fine-tuned models—IndicTrans2-Distilled and NLLB-200—showed improved translation quality across metrics and human reviews. The work aims to support better multilingual AI for Indian languages, where existing datasets often lack depth or cultural relevance.

The story so far

2 episodes →
  1. Researchers release COILD corpus for Indian language machine translationthis story
Read the original at arXiv cs.CL · by Kshetrimayum Boynao Singh, Nitin Kumar Mishra, Palash Pratim Dutta, Atai Waris Khan, Aparna Kaushik, Avinash Kumar, Deeksha, Deepak Kumar, Saroj Kumar Jha, Saloka Sengupta, Anansa Roy, Umalatha Kannoth, Saifulla Samar, Meena Sharma, Manpreet Kaur, Jyoti Sharma, Ashwini Vaidya, Muralikrishna SN, Md primary sourceOpen source ↗
Topics · follow one to build your own front page
IndicTrans2-DistilledNLLB-200

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories