On September 15, 2026, Google released the live voice dialogue models Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, available in 97 languages. The main novelty is the Extended Thinking mode, in which the model reasons and simultaneously speaks with the user.


What happened
The released models understand audio, images, and camera video in near real-time and can switch between 97 supported languages during a conversation. They call tools and APIs in the background without interrupting the dialogue. According to Google, the base Gemini 3.8 Live took second place in the Speech Agent Arena, while Extended Thinking took first place in the Speech to Speech Quality Index (Artificial Analysis) with a score of 82.6; it also recorded 68.6% on τ-Voice, 35.1% on τ-Voice-banking (Sierra), and 97.7% on Big Bench Audio. During reasoning, the Extended Thinking version gives verbal cues like “Let me check that…” and voices the progress of background tasks. The API pricing on the paid tier per 1 million tokens is: $0.75 for text input and $4.50 for text output, $3.00 for audio input, and $12.00 for audio output, with audio additionally billed at $0.005/minute on input and $0.018/minute on output. Both models are already available in the Gemini API and Google AI Studio, with a free tier.
Context
The main user inconvenience with voice assistants has long been the pause between a question and an answer — the moment when the model “thought and went silent.” The Extended Thinking mode is aimed precisely at this: instead of silence, the model speaks while thinking, voicing its actions. The release also marked a shift in the culture of evaluating voice AI: previously, vendors selectively published quality metrics or hid them entirely in press releases, but now public speech-to-speech benchmarks were released alongside the announcement. The infrastructure for integration is ready: the LiveKit, Vercel, LangChain, and Agora ecosystems are supported, and a private preview of Gemini Enterprise is available for enterprises.
Why this matters for the industry
For developers and companies, the release moves voice agents from the category of experimental demos to a deployable product: a public API, free tier, and published prices allow immediately calculating the unit economics of a voice agent and comparing it with text APIs and existing voice solutions. Support, sales, and banking scenarios, which were previously expensive, become economically comparable to text dialogues, while deployment in Search Live, Gemini Live, and Workspace (Docs Live, Gmail Live, Keep Live) expands the industry's reference base. The publication of benchmarks set a quality benchmark: the next generations of voice models will be evaluated by the same speech-to-speech and agentic metrics, not by marketing demos, while the 35.1% result on τ-Voice-banking indicates the upper limit of current reliability in complex scenarios.
Why this matters for users
For readers, the models are immediately available: the free tier in the Gemini API and Google AI Studio allows independently testing the “reasons while speaking” pattern and background API calls, while Google AI Pro/Ultra subscribers get the functionality in Gemini Live, Gmail, Docs, and Keep. In practice, this is a voice dialogue without pauses for thinking: the assistant can answer questions while performing a task and voices its progress.
What is still unknown / limitations
All benchmark figures cited are based on Google's data; independent verification is not yet available. Data on latency, rate limits, and SLA have not been published, and there is no paper with ablations, model size, or error analysis. The 35.1% result on τ-Voice-banking indicates the upper limit of reliability in complex agentic scenarios, and leaderboard positions are given without absolute values for other participants, so the scale of the lead cannot be assessed. The choice of figures highlighted is part of the narrative, and there is a risk of cherry-picking. The claim that parallel reasoning removes the main barrier for voice assistants remains a marketing interpretation: the quality of responses in “thinking and speaking” mode has not been independently measured.
Sources
- Gemini 3.8 Live & Gemini 3.8 Live Extended Thinking — The Keyword (Google's official blog)
- Gemini 3.8 Live Extended Thinking powers Gemini Live, Gmail — 9to5Google
Author
Look at AI, editorial team
