{"version":1,"type":"story","url":"https://digestai.news/story/researchers-release-coild-corpus-for-indian-language-machine-translati","json":"https://digestai.news/story/researchers-release-coild-corpus-for-indian-language-machine-translati.json","markdown":"https://digestai.news/story/researchers-release-coild-corpus-for-indian-language-machine-translati.md","slug":"researchers-release-coild-corpus-for-indian-language-machine-translati","headline":"Researchers release COILD corpus for Indian language machine translation","summary":"A team of researchers has introduced **COILD**, a new parallel corpus and benchmark for machine translation between Indian languages. The dataset includes over **1.16 million human-translated and verified sentence pairs** across **20 language pairs**, covering families like Indo-Aryan, Dravidian, Tibeto-Burman, and Austro-Asiatic. The data comes from **licensed repositories** and spans **eight real-world domains**, such as legal, medical, and literary texts, to reflect cultural and linguistic diversity beyond English-centric models.\n\nThe researchers also created a **domain-centric benchmark** of **2,000 expert-verified sentences** for consistent evaluation. Testing fine-tuned models—**IndicTrans2-Distilled** and **NLLB-200**—showed improved translation quality across metrics and human reviews. The work aims to support better multilingual AI for Indian languages, where existing datasets often lack depth or cultural relevance.","keyPoints":["COILD corpus contains 1.16 million human-translated sentence pairs across 20 Indian language pairs","Data sourced from licensed repositories covering eight real-world domains like legal and medical texts","Fine-tuning IndicTrans2-Distilled and NLLB-200 models improved translation accuracy per researchers"],"whyItMatters":"COILD addresses a critical gap in AI training data for Indian languages, enabling more accurate and culturally relevant machine translation beyond English-centric models.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":["IndicTrans2-Distilled","NLLB-200"],"people":[]},"firstPublishedAt":"2026-09-25T04:00:00Z","updatedAt":"2026-09-25T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.CL","title":"COILD: An Indic-Centric Parallel Corpus and Benchmark for Machine Translation Across Indian Languages","url":"https://arxiv.org/abs/2609.28826","publishedAt":"2026-09-25T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":{"title":"Global Efforts Expand Machine Translation Corpora","url":"https://digestai.news/thread/new-wolofarabic-corpus-boosts-machine-translation-accuracy","storyCount":2},"cite":{"text":"Digest AI, \"Researchers release COILD corpus for Indian language machine translation\", 25 September 2026, https://digestai.news/story/researchers-release-coild-corpus-for-indian-language-machine-translati","publisher":"Digest AI","title":"Researchers release COILD corpus for Indian language machine translation","datePublished":"2026-09-25T04:00:00Z","url":"https://digestai.news/story/researchers-release-coild-corpus-for-indian-language-machine-translati"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}