Global Efforts Expand Machine Translation Corpora
Researchers are developing specialized datasets to improve machine translation accuracy for underrepresented languages. The latest release of the COILD corpus for Indian languages follows a similar initiative for Wolof-Arabic translation.
-
Researchers release COILD corpus for Indian language machine translation
A team of researchers has introduced COILD, a new parallel corpus and benchmark for machine translation between Indian languages. The dataset includes over 1.16 million human-translated and verified…
1 source primary source -
New Wolof‑Arabic Corpus Boosts Machine Translation Accuracy
Researchers have released MudawanSn, a curated collection of 1,271 sentence pairs that map Wolof text to Modern Standard Arabic. The data were extracted from the MasakhaNER news corpus and cover…
1 source primary source