New Wolof‑Arabic Corpus Boosts Machine Translation Accuracy
Researchers have released MudawanSn, a curated collection of 1,271 sentence pairs that map Wolof text to Modern Standard Arabic. The data were extracted from the MasakhaNER news corpus and cover topics such as politics, society, religion, and sports in Senegal. Unlike broader multilingual datasets like FLORES‑200 or NTREX, MudawanSn is the first publicly available parallel corpus specifically…
Key points
- MudawanSn: 1,271 manually aligned Wolof–Arabic sentence pairs from MasakhaNER, first public parallel corpus for this low‑resource pair.
- Fine‑tuning AfriNLLB‑12 on MudawanSn yields 7.76 BLEU (Wolof→Arabic) and 8.75 BLEU (Arabic→Wolof), outperforming baselines.
- Dataset released under CC BY‑NC, available on Hugging Face and GitHub, enabling research and benchmarking for Wolof‑Arabic MT.
The authors evaluated four machine‑translation systems—NLLB‑200 (600 M parameters), mT5‑base, and two AfriNLLB variants—on both translation directions. Fine‑tuning with MudawanSn produced marked gains, with the AfriNLLB‑12 model achieving 7.76 BLEU and 30.72 chrF++ for Wolof‑to‑Arabic and 8.75 BLEU and 33.08 chrF++ for Arabic‑to‑Wolof. These scores surpass the baselines by several points, demonstrating the corpus’s value for improving performance on this low‑resource pair.
MudawanSn is released under a CC BY‑NC license and can be accessed on Hugging Face and GitHub. The dataset’s high‑quality alignment and manual translation make it a valuable benchmark for future research on Wolof‑Arabic machine translation and low‑resource language technology.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Research
All →- Legal LLMs' Hallucinations Should Be Evaluated as Warrant Failures · 1 src
- Benchmarking LLMs for Key-Value Extraction in Noisy OCR Documents · 1 src
- Study Shows Relation Facts Trigger Earlier Than Entity Facts in Language Models · 1 src
- LLM-Enhanced Model Improves Extubation Failure Prediction Using Therapy Notes · 1 src
- Study maps four-stage pipeline for LLM math word problems, isolates failure point · 1 src
Comments
via GitHub Discussions