On September 15, 2026, Google released two live models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, which process continuous streams of audio, video, and text, respond in real time by voice, and support 97 languages with automatic switching directly in the conversation.


What happened
Both versions see an image from the camera in near real time and launch tools and API calls in the background without interrupting speech. Gemini 3.8 Live Extended Thinking reasons and speaks simultaneously, accompanying background multi-step tasks with live commentary. According to Google, Extended Thinking took 1st place in the Speech to Speech Quality Index from Artificial Analysis with a score of 82.6, showed 68.6% on tau-Voice, 35.1% on tau-Voice-banking from Sierra, and 97.7% on Big Bench Audio; Gemini 3.8 Live took 2nd place in the Speech Agent Arena. The prices of the paid Gemini API tier per 1 million tokens are the same for both models: input — text $0.75, audio $3; output — text $4.50, audio $12. Both versions are available in the Gemini API and Google AI Studio, including free access, and a private preview of Gemini Enterprise is open to enterprise customers. Google announced integrations via the Live API with Agora, LiveKit, Pipecat, LangChain, and Vercel, as well as partnerships with Salesforce and Genspark.
Context
To understand the release, it is important to know what familiar problem it solves: traditional voice assistants performed a task and remained silent while the user waited for the result. Gemini 3.8 Live introduces the opposite scenario, in which the assistant continues to speak while tools and API calls are executed in the background, and in the Extended Thinking version, reasoning proceeds in parallel with speech. At the same time, the release is primarily of an engineering-product nature: the scientific novelty of the models is not confirmed by a technical report and ablations, and at the time of the announcement, the main value of the release lies in the capabilities, the ready-made API, and the price, not in research results.
Why this matters for the industry
Developers of voice agents have received a production tool that removes the main limitation of the category: the dialogue stops falling silent during task execution. The published unified pricing for audio and text tokens sets a benchmark for the economics of voice agents, around which competitors will have to build their own rates. Ready-made integrations and partnerships accelerate the appearance of voice interfaces in products and compress the path from idea to a working demo to a few days, as a proprietary speech recognition and synthesis stack for a pilot is no longer needed.
Why this matters for users
The models can be tried for free in Google AI Studio and via the Gemini API, and you can build your own voice agent without your own ASR and TTS stack. Regular users will get a new voice mode in Gemini Live, Search Live, and in Workspace (Docs Live, Gmail Live, Keep Live) — with an understanding of what is happening on the screen and in the frame.
What is still unknown / limitations
All the metrics provided are figures claimed by Google: there is no independent replication, no description of the evaluation protocol, and no failure analysis. There is no technical report, ablations, or description of the architecture, so claims about production-readiness and a platform shift remain marketing formulations, not verified results. There is no per-language breakdown of quality, latency measurements, or disclosed end-to-end delay. The result of 97.7% on Big Bench Audio does not automatically transfer to reliability in complex production scenarios, and the gap between 68.6% on tau-Voice and 35.1% on tau-Voice-banking shows that complex multi-step tool tasks in the voice channel are still far from saturated. Corporate access is currently limited to the private preview of Gemini Enterprise without public SLAs.
Sources
Author
Look at AI, editorial team
