Researchers introduce FD-VAD for streaming full-duplex speech endpoint detection
A new arXiv paper presents FD-VAD, an ASR-free system for semantic endpoint detection in streaming full-duplex voice interaction. The model determines whether a pause signals hesitation or a completed turn by mapping causal audio windows directly to Continue/Stop decisions, without relying on automatic speech recognition or dialogue state tracking.
Key points
- FD-VAD performs semantic endpoint detection directly from streaming audio without ASR
- Achieves 0.853 EOT recall at FP≤0.10 on TurnBench dev set in zero-shot setting
- Uses frozen speech encoder, modality adapter, and adapted language model with last-chunk training
FD-VAD combines a frozen speech encoder, a lightweight modality adapter, and a parameter-efficiently adapted language model trained with a last-chunk objective for streaming inference. It adds confidence-gated endpoint commitment to balance interruption risk against delay, and boundary-focused hard-negative sampling to improve decisions at ambiguous turn boundaries. On the TurnBench dev set, FD-VAD achieves an EOT recall of 0.853 at a false positive rate ≤0.10 in a zero-shot setting, outperforming strong streaming and non-streaming semantic turn classifiers across in-domain and conversational evaluations.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Researchers release ArgGYM benchmark for testing defeasible reasoning in AI models · 1 src
- Researchers release SimTrace for generating synthetic user behavior data · 1 src
- MetaPersona framework uses 11,000+ studies to build synthetic populations for AI tasks · 1 src
- Researchers propose DLFP controller to cut AI inference latency by up to 30% · 1 src
- Researchers introduce GoldiMask to improve diffusion language model fine-tuning · 1 src
Comments
via GitHub Discussions