OpenAI launches GPT‑Live‑1 API enabling simultaneous speech and listening
OpenAI has opened its new GPT‑Live‑1 speech model to developers via an API that can both listen and speak at the same time, a capability called full‑duplex. The model, already integrated into ChatGPT, lets developers pair it with different backend engines to balance reasoning depth, speed, and cost for various applications. Priced at $0.05 per minute, the service targets high‑value use cases…
Key points
- GPT‑Live‑1 offers full‑duplex speech at $0.05 per minute, enabling simultaneous listening and speaking
- Benchmark scores: 80.1% interactivity, 0.8 s latency, 87% tool‑calling accuracy, 32% banking pass rate
- Yelp uses the model for phone reservations, noting improved call handling
Early adopters such as Yelp are using GPT‑Live‑1 for phone‑based reservation calls, reporting smoother handling and higher satisfaction. In OpenAI’s internal benchmarks, the model outperforms its predecessor GPT‑Realtime‑2.1, achieving an 80.1% full‑duplex interactivity score versus 45.4%, cutting turn‑taking latency to 0.8 seconds, and raising tool‑calling accuracy to 87%. A banking voice‑support test shows a 32% pass rate, up from 12.4% previously. The release also adds twelve new voices covering diverse accents, dialects, and languages, and provides automatic speech‑to‑text transcripts alongside response text.
OpenAI's GPT-Live-1 API lets developers build apps that talk and listen at the same time
The Decoder · 10 September 2026
OpenAI's GPT-Live-1 API lets developers build apps that talk and listen at the same time
OpenAI is making GPT-Live-1 available to developers as an API. The speech model can listen and talk at the same time, a feature known as "full-duplex," and is already running inside ChatGPT. Developers can pair it with different backend models depending on the task, matching reasoning depth, speed, and cost to each use case. At $0.05 per minute, it's not cheap. Yelp is using the model for phone-based reservations and reports better call handling, according to CTO Alex Levy.
On OpenAI's benchmarks, GPT-Live-1 pulls well ahead of its predecessors. In full-duplex interactivity tests, it scores 80.1 percent compared to 45.4 percent for GPT-Realtime-2.1. Turn-taking latency drops to 0.8 seconds from 1.4 seconds. Tool-calling accuracy jumps to 87 percent from 60 percent. In a banking voice support benchmark, GPT-Live-1 hits a 32 percent pass rate, up from 12.4 percent for the previous model.
GPT-Live-1 also ships with twelve new voices spanning different accents, dialects, and languages. It provides ASR transcripts and response text out of the box. Full details will be available in the API documentation.
This text was published by The Decoder and written by Matthias Bastian. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
2 sources- Hacker News discussion · 46 points news.ycombinator.com
- Build more natural voice experiences with GPT‐Live‐1 in the API Primary source · OpenAI ·
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
Comments
via GitHub Discussions