Gemini 3.8 & 3.8 Live Extended Thinking
Gemini 3.8 Live and 3.8 Live Extended Thinking are real-time voice and multimodal dialogue models engineered for concurrent reasoning, background tool calling, and live interactive voice workflows.
Gemini 3.8 Live and 3.8 Live Extended Thinking are real-time voice and multimodal dialogue models engineered for concurrent reasoning, background tool calling, and live interactive voice workflows.
What the product does and how it is positioned
Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are dialogue models designed for near real-time voice interactions, visual grounding, and multi-step reasoning.
Gemini 3.8 Live focuses on scalable conversational intelligence with automatic language switching across 97 languages, while 3.8 Live Extended Thinking delivers simultaneous parallel reasoning, early conversational cues, and continuous background task narration.
Source-supported ways to use the product
Guiding employee onboarding sessions dynamically by answering questions using real-time visual context.
Coordinating multi-step bookings and running background function calls without interrupting live spoken conversation.
Transforming hand-drawn visual sketches and spoken instructions into functional React components.
Gemini 3.8 Live Extended Thinking addresses interruptions in voice workflows by reasoning and speaking at the same time. Rather than pausing dialogue while processing complex requests, the model introduces conversational cues like verbal acknowledgments and progress narration.
Background execution capabilities allow both models to make API calls, run external tools, and handle multi-step actions asynchronously. Spoken interaction continues without pauses while underlying tasks finalize.
Checks to run with your own material and workflow
What was checked and when
Answers based on the source-checked product record
It reasons and speaks simultaneously, using natural verbal acknowledgments and live progress narration to keep conversation flowing while multi-step tasks run.
The models automatically detect and transition between 97 supported languages mid-conversation without manual setting adjustments.
All audio output generated by the models is embedded with an imperceptible SynthID watermark to ensure the media remains detectable as AI-generated content.
Developers can access both models through the Gemini API and Google AI Studio, as well as through supported partner streaming platforms.
Yes, Gemini 3.8 Live processes visual inputs in near real-time, allowing the model to ground conversations and answer questions based on visual context.