GPT‑Live‑1 brings ChatGPT’s natural, full-duplex conversations to the API, with more control over how voice agents speak and act.
We’re launching GPT‑Live‑1 in the API, giving developers a powerful, natural voice model for building voice-enabled apps and business workflows. First introduced in ChatGPT, GPT‑Live‑1 is capable of listening and speaking at the same time, and, as seen with Codex and ChatGPT Work(opens in a new window), can delegate deeper reasoning and actions to the models and tools it is paired with.
For the API release of GPT‑Live‑1, we’ve focused on new capabilities that let developers steer and customize voice experiences around their users, workflows, and goals. A core GPT‑Live‑1 strength, smooth interruption handling, is already delivering business impact: in early evaluations, Speak found that GPT‑Live‑1 gave learners more time to think before the language tutor responded, cutting interruptions by almost 80% versus previous turn-based systems.
Interruption handling: Improves interruption handling via a single model that reasons over incoming and outgoing audio together, avoiding the latency and brittle handoffs of chained STT–LLM–TTS architectures.
Reasoning & tool calling delegation: GPT‑Live‑1 can delegate reasoning and tool calls to a backend text model like GPT‑6 Astra or a third-party model.
Tone, pace, and style: Lets developers shape an agent’s tone, pace, and conversational style through the system prompt.
Silent context management & background noise: Better handles background noise and silence without interrupting the conversation or narrating every step out loud.
Telephony support: Enables deployment of full-duplex voice agents for phone calls, from restaurant reservations to customer support.
Source link







