DigestAI news desk
Research updated

New Wolof‑Arabic Corpus Boosts Machine Translation Accuracy

Researchers have released MudawanSn, a curated collection of 1,271 sentence pairs that map Wolof text to Modern Standard Arabic. The data were extracted from the MasakhaNER news corpus and cover topics such as politics, society, religion, and sports in Senegal. Unlike broader multilingual datasets like FLORES‑200 or NTREX, MudawanSn is the first publicly available parallel corpus specifically…

1 source primary source

Key points

  • MudawanSn: 1,271 manually aligned Wolof–Arabic sentence pairs from MasakhaNER, first public parallel corpus for this low‑resource pair.
  • Fine‑tuning AfriNLLB‑12 on MudawanSn yields 7.76 BLEU (Wolof→Arabic) and 8.75 BLEU (Arabic→Wolof), outperforming baselines.
  • Dataset released under CC BY‑NC, available on Hugging Face and GitHub, enabling research and benchmarking for Wolof‑Arabic MT.

The authors evaluated four machine‑translation systems—NLLB‑200 (600 M parameters), mT5‑base, and two AfriNLLB variants—on both translation directions. Fine‑tuning with MudawanSn produced marked gains, with the AfriNLLB‑12 model achieving 7.76 BLEU and 30.72 chrF++ for Wolof‑to‑Arabic and 8.75 BLEU and 33.08 chrF++ for Arabic‑to‑Wolof. These scores surpass the baselines by several points, demonstrating the corpus’s value for improving performance on this low‑resource pair.

MudawanSn is released under a CC BY‑NC license and can be accessed on Hugging Face and GitHub. The dataset’s high‑quality alignment and manual translation make it a valuable benchmark for future research on Wolof‑Arabic machine translation and low‑resource language technology.

Read the original at arXiv cs.CL · by Mouhamed Mbaye, Thierno Diop primary source Open source ↗
Topics · follow one to build your own front page
NLLB-200mT5-baseAfriNLLB-12

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

More in Research

All →

Related stories