Qwen2.5-Omni-3B adapters boost entity recall in accented conversational ASR
A new three‑stage pipeline targets accented English speech from India, Indonesia and Latin America, where standard ASR systems optimized for word error rate often miss named entities and filler words. First, heuristic SQL filters curate training data that is 2.8 × richer in entities than random samples. Second, regional LoRA adapters fine‑tuned on the Qwen2.5‑Omni‑3B model generate both verbatim…
Key points
- Heuristic SQL filters produce training data with 2.8× higher entity density than random sampling.
- LoRA adapters fine‑tuned on Qwen2.5‑Omni‑3B raise entity recall to 80‑85% and filler recall to 76‑86% on accented speech.
- Pipeline matches a zero‑shot 30B model’s performance with only 3B parameters and 10× fewer resources.
On a 6 k‑utterance test set the pipeline reaches 80‑85 % entity recall (up from 53‑55 %) and 76‑86 % filler recall (up from under 5 %), while keeping word error rate between 6 % and 10 %. It outperforms Whisper and a commercial ASR on entity recall and matches a zero‑shot 30 B‑parameter model with ten‑fold fewer parameters. Paired bootstrap tests attribute 2.8‑4.2 percentage‑point gains in entity recall to data curation alone (p < 0.0001).
The story so far
2 episodes →- Qwen2.5-Omni-3B adapters boost entity recall in accented conversational ASRthis story
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Research
All →- TatBLiMP benchmarks Tatar linguistic minimal pairs · 1 src
- Recursive language models generalize out of domain, study shows · 1 src
- Researchers propose AI-GRACE framework for operationalizing agentic AI use cases · 1 src
- Embed‑TTT improves rule induction in ARC‑like tasks, authors say · 1 src
- Researchers introduce SpecOpt for agentic molecule specificity optimization · 1 src
Comments
via GitHub Discussions