On September 15, 2026, Google introduced two live voice models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, which move the voice assistant from a demo capability to a production-ready product with published API pricing and external speech quality scores.

image

What happened

The announcement was published on September 15, 2026, and includes two model variants with the identifiers gemini-3.8-live and gemini-3.8-live-extended-thinking. The models work with 97 languages, automatically switching languages mid-conversation, process video and camera images in near real time, and execute tool and API calls in the background without interrupting the user's speech. The Gemini 3.8 Live Extended Thinking version reasons and speaks simultaneously, narrating intermediate steps of multi-step tasks in the form of a live progress narrative. According to Google, the extended version took 1st place overall in the Artificial Analysis Speech to Speech Quality Index with a score of 82.6, scored 68.6% on τ-Voice, 35.1% on Sierra τ-Voice-banking, and 97.7% on Big Bench Audio, while Gemini 3.8 Live became 2nd in the Speech Agent Arena. The API cost per 1 million tokens is $0.75 for input and $4.50 for output for text, and $3/$12 for audio.

Context

The release belongs to the class of native speech-to-speech models, where a single model directly perceives and produces speech, replacing the traditional chain of a speech recognizer, text model, and synthesizer, in which each dialogue pass goes through several stages. The quality scores cited in the announcement rely on external platforms, the Artificial Analysis and Speech Agent Arena indices, as well as the τ-Voice, Sierra τ-Voice-banking, and Big Bench Audio benchmarks, so the claimed level can be re-verified by independent measurements.

Why this matters for the industry

For the industry, the release sets a benchmark in both key categories of the real-time voice agent market: published API rates become a price benchmark, and the models' benchmark results become a quality benchmark against which competitors will now be measured. Production-ready voice models with background tool execution and spoken reasoning give companies a ready-made infrastructure layer for multi-step voice processes, such as ordering, booking, and support, without building their own voice stack from cascading components. At the same time, the gap between the model's result on the general-system τ-Voice benchmark and the domain-specific Sierra τ-Voice-banking shows a noticeable drop in quality in applied scenarios like banking, which limits the application of such agents in critical business processes until this gap is closed.

Why this matters for users

Developers can try gemini-3.8-live and gemini-3.8-live-extended-thinking for free right now in Google AI Studio, connect them to their own voice agents via the Gemini API, run an A/B comparison with their current STT/LLM/TTS cascade on latency, intelligibility, and cost, and calculate the cost of their own voice scenarios based on the published token prices. Regular users will start encountering these models as they are rolled out into Google products: Search Live, Gemini Live, and Workspace.

What is still unknown / limitations

All benchmark numbers in the announcement are cited with the reference "according to Google," the measurement methodology has not been published, and there is no independent replication of the results. There is no public technical report in the sources describing the architecture and training method of the Extended Thinking mode, so the novelty of the mode is currently confirmed only by the announcement itself. The 97.7% result on Big Bench Audio likely reflects saturation of this benchmark and is weakly informative about the model's actual capabilities. Actual latency, reliability, and voice quality in real-world scenarios have not yet been published and must be verified by independent testing.

Sources

Author

Look at AI, editorial team