DigestAI news desk

Cut through the AI noise.

Research

PTC-Bias improves speech LLM accuracy with phoneme-level bias correction

Researchers introduced PTC-Bias, a two-stage method for reducing errors in speech large language models (SpeechLLMs). The framework uses phoneme-level temporal competition to refine rare-word recognition. In the first stage, PTC Retrieval decodes phonemes frame-by-frame and narrows down candidate pronunciations into a shortlist. The second stage, PTC Correction, compares retrieved candidates…

1 source primary source

Key points

  • PTC-Bias uses phoneme-level temporal competition to filter and correct rare-word errors in SpeechLLMs
  • Tests on LibriSpeech show up to 23.9% relative reduction in bias-word error rates for 2000-word lists
  • Method requires no additional SpeechLLM forward passes, relying only on phoneme posteriors

Tests on the LibriSpeech dataset show PTC-Bias cuts bias-word error rates by up to 23.4% (test-clean) and 23.9% (test-other) compared to CTC-Filter, while keeping unweighted error rates stable. The method works with bias lists of up to 2000 words and two SpeechLLMs: Prompt-SLAM-ASR-7B and an unspecified second model. No extra SpeechLLM forward passes are needed, as both stages reuse phoneme posteriors.

Read the original at arXiv cs.CL · by Zhiqi Ai, Han Cheng, Shiyi Mu, Yongjin Zhou, Shugong Xu primary sourceOpen source ↗
Topics · follow one to build your own front page
Prompt-SLAM-ASR-7B

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories