Google Unveils Gemini 3.8 Live Models With Real-Time Reasoning for Voice Agents

Reviewed byNidhi Govil

6 Sources

Share

Google introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking as its most advanced voice processing models. The Extended Thinking variant scored 82.6 on the Speech to Speech Quality Index, claiming the top spot. Both models support 97 languages, process visual inputs in real-time, and execute API calls in the background while maintaining natural conversations.

Google Expands Gemini 3.8 Family With Advanced Voice Models

Google has introduced two new AI speech models—Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking—designed to transform how voice agents handle complex tasks

1

2

. These production-ready voice assistants arrive just weeks after the Gemini 3.7 Flash launch, marking Google's rapid advancement in conversational AI technology. Both models feature near real-time reasoning capabilities that enable more fluid, intelligent interactions between users and AI systems

3

.

Source: 9to5Google

Source: 9to5Google

Gemini 3.8 Live is built for scale and cost efficiency, priced at $0.005 per minute for audio inputs and $0.018 per minute for outputs

4

. Meanwhile, Gemini 3.8 Live Extended Thinking targets high-complexity tasks with increased intelligence and multi-step reasoning, though it includes additional charges for reasoning tokens and supplementary inputs like video and documents.

Record-Breaking Performance on Industry Benchmarks

Gemini 3.8 Live Extended Thinking has captured the number one spot on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6, surpassing competitors like GPT-Live-1-Astra and Grok Voice Think Fast 2.0

3

4

. The model demonstrates exceptional agentic task completion capabilities with a 68.6% score on τ-Voice and 35.1% on the τ-Voice-banking benchmark, which evaluates real-time, audio-native conversational AI agents on complex banking and fintech customer support scenarios

1

.

The Extended Thinking variant also achieved 97.7% on Big Bench Audio, showcasing strong reasoning capabilities while maintaining competitive pricing compared to other frontier models

2

. Gemini 3.8 Live secured second place in the Speech Agent Arena and first place on ServiceNow's EVA-Bench, demonstrating high user preference and enterprise-grade performance

4

.

Simultaneous Reasoning and Speaking Transforms User Experience

A defining feature of these live dialogue models is their ability to reason and speak simultaneously, addressing latency issues that have plagued voice-based AI systems

4

. Gemini 3.8 Live Extended Thinking maintains uninterrupted conversational flow by using early verbal cues like "Let me check that..." to acknowledge prompts naturally while processing requests in the background

2

.

Both models can execute tools and API calls behind the scenes, allowing enterprise voice agents to continue conversations while simultaneously completing assigned tasks

1

4

. This creates a more natural, human-like experience without constant interruptions as the AI searches for information or performs actions.

Real-Time Visual Input Processing and Language Detection

Gemini 3.8 Live processes visual inputs in near real-time, enabling visual grounding that enhances contextual understanding during conversations

1

2

. The model can detect and transition between 97 supported languages mid-conversation, making it accessible to global users and suitable for multilingual enterprise environments

1

4

.

Source: PYMNTS

Source: PYMNTS

This automatic language detection capability allows users to switch languages naturally without manual configuration, while the real-time visual input processing enables the AI to understand and respond to what users are showing through their cameras or screens.

Availability Across Google Workspace and Developer Platforms

Gemini 3.8 Live powers Search Live in AI Mode, while Gemini 3.8 Live Extended Thinking is available in Gemini Live for all users

1

2

. Subscribers with AI Pro and Ultra plans can access the Extended Thinking model in Google Workspace applications including Gmail Live for conversational search, Docs Live for draft generation and editing, and Keep Live for note creation

1

2

.

Developers can integrate both models through the Gemini API and Google AI Studio, while enterprise users have access via private preview in Gemini Enterprise

1

4

. The models will eventually roll out to Gemini Enterprise for Customer Experience and Google Workspace business customers. Google partners including Vercel, Agora, LiveKit, Pipecat, Fishjam, and Vision Agents enable developers to build applications using these AI speech models

4

.

SynthID Watermarking Addresses AI-Generated Content Concerns

All audio generated by Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking carries invisible SynthID watermarking

1

4

. Google's digital watermarking and detection tool embeds signals into AI-generated content that can be used to identify synthetic audio and combat misinformation. This transparency measure becomes increasingly important as large language models and conversational AI systems generate more realistic human-like speech.

Voice Technology Becomes Infrastructure for Agentic Commerce

The launch reflects broader industry momentum toward voice-native AI infrastructure

5

. Voice technology is emerging as foundational middleware between large language models and end-user transactions, reducing friction at critical moments when consumers are ready to make purchases. The past year has seen accelerated investment in voice-native AI infrastructure, signaling investor confidence in high-fidelity conversational AI as essential rather than supplementary technology

5

.

Source: SiliconANGLE

Source: SiliconANGLE

Voice AI is shifting from answering questions to taking action, with competitors like Anthropic's Claude voice mode and OpenAI's Presence enabling task completion through conversation

5

. Google's new models position the company to compete in this evolving landscape where enterprise voice agents handle customer support, sales, and complex workflows through natural dialogue.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved