Researchers test pseudo‑labeling to improve ASR on noisy police audio
Researchers evaluated pseudo‑labeling to adapt large ASR models to noisy police communication data from Baltimore and Chicago. They found that standard confidence metrics such as log‑probabilities and STAR scores cannot reliably separate high‑from low‑quality pseudo‑labels.
Key points
- Pseudo‑labeling improves ASR on noisy police audio but internal confidence metrics fail to distinguish quality
- External LLM‑as‑judge filtering reduces WER more than internal metrics
- Cross‑model pseudo‑labeling shows promise for future ASR adaptation
To improve filtering, the team introduced an external LLM‑as‑judge approach that discards contextually implausible transcripts. This method reduced word‑error‑rate (WER) more than internal metrics, though a gap remains compared to an oracle filter. They also explored a cross‑model pseudo‑labeling strategy, where one model is fine‑tuned with pseudo‑labels generated by another, and identified it as a promising direction for future work.
The story so far
3 episodes →- Researchers test pseudo‑labeling to improve ASR on noisy police audiothis story
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Generative AI & Models
All →- typeSafe AI launches Jev, a judgment‑only model, priced at $0.042 per million tokens · 2 src
- OpenAI and Anthropic cut API prices for GPT-6 Sol and Claude Opus 5.5 on September 22 · 15 src
- OpenAI reportedly preparing GPT-6 Cyber model for select customers · 8 src
- Google DeepMind says Gemini 4 is nearing launch · 1 src
- NaiveAI releases Naive-N0.5-Flash 309B MoE model with 1M context under MIT license · 1 src
Comments
via GitHub Discussions