{"version":1,"type":"story","url":"https://digestai.news/story/researchers-introduce-fd-vad-for-streaming-full-duplex-speech-endpoint","json":"https://digestai.news/story/researchers-introduce-fd-vad-for-streaming-full-duplex-speech-endpoint.json","markdown":"https://digestai.news/story/researchers-introduce-fd-vad-for-streaming-full-duplex-speech-endpoint.md","slug":"researchers-introduce-fd-vad-for-streaming-full-duplex-speech-endpoint","headline":"Researchers introduce FD-VAD for streaming full-duplex speech endpoint detection","summary":"A new arXiv paper presents FD-VAD, an ASR-free system for semantic endpoint detection in streaming full-duplex voice interaction. The model determines whether a pause signals hesitation or a completed turn by mapping causal audio windows directly to Continue/Stop decisions, without relying on automatic speech recognition or dialogue state tracking.\n\nFD-VAD combines a frozen speech encoder, a lightweight modality adapter, and a parameter-efficiently adapted language model trained with a last-chunk objective for streaming inference. It adds confidence-gated endpoint commitment to balance interruption risk against delay, and boundary-focused hard-negative sampling to improve decisions at ambiguous turn boundaries. On the TurnBench dev set, FD-VAD achieves an EOT recall of 0.853 at a false positive rate ≤0.10 in a zero-shot setting, outperforming strong streaming and non-streaming semantic turn classifiers across in-domain and conversational evaluations.","keyPoints":["FD-VAD performs semantic endpoint detection directly from streaming audio without ASR","Achieves 0.853 EOT recall at FP≤0.10 on TurnBench dev set in zero-shot setting","Uses frozen speech encoder, modality adapter, and adapted language model with last-chunk training"],"whyItMatters":"Enables natural turn-taking in full-duplex voice AI by detecting conversational intent from audio alone, removing ASR dependency and reducing latency.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":["FD-VAD"],"people":[]},"firstPublishedAt":"2026-09-30T04:00:00Z","updatedAt":"2026-09-30T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.CL","title":"FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech","url":"https://arxiv.org/abs/2609.35791","publishedAt":"2026-09-30T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Researchers introduce FD-VAD for streaming full-duplex speech endpoint detection\", 30 September 2026, https://digestai.news/story/researchers-introduce-fd-vad-for-streaming-full-duplex-speech-endpoint","publisher":"Digest AI","title":"Researchers introduce FD-VAD for streaming full-duplex speech endpoint detection","datePublished":"2026-09-30T04:00:00Z","url":"https://digestai.news/story/researchers-introduce-fd-vad-for-streaming-full-duplex-speech-endpoint"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}