DigestAI news desk

AI news, digested. Every story with its sources, every hour.

Research

Qwen2.5-Omni-3B adapters boost entity recall in accented conversational ASR

A new three‑stage pipeline targets accented English speech from India, Indonesia and Latin America, where standard ASR systems optimized for word error rate often miss named entities and filler words. First, heuristic SQL filters curate training data that is 2.8 × richer in entities than random samples. Second, regional LoRA adapters fine‑tuned on the Qwen2.5‑Omni‑3B model generate both verbatim…

1 source primary source

Key points

  • Heuristic SQL filters produce training data with 2.8× higher entity density than random sampling.
  • LoRA adapters fine‑tuned on Qwen2.5‑Omni‑3B raise entity recall to 80‑85% and filler recall to 76‑86% on accented speech.
  • Pipeline matches a zero‑shot 30B model’s performance with only 3B parameters and 10× fewer resources.

On a 6 k‑utterance test set the pipeline reaches 80‑85 % entity recall (up from 53‑55 %) and 76‑86 % filler recall (up from under 5 %), while keeping word error rate between 6 % and 10 %. It outperforms Whisper and a commercial ASR on entity recall and matches a zero‑shot 30 B‑parameter model with ten‑fold fewer parameters. Paired bootstrap tests attribute 2.8‑4.2 percentage‑point gains in entity recall to data curation alone (p < 0.0001).

The story so far

2 episodes →
  1. Qwen2.5-Omni-3B adapters boost entity recall in accented conversational ASRthis story
Read the original atarXiv cs.CL · by Fiza Husain, Ankit Pandey, Yash Singh primary sourceOpen source ↗
Topics · follow one to build your own front page
Qwen2.5-Omni-3BWhisper

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

More in Research

All →

Related stories