Google Launches Gemini 3.5 Transcribe with 70% Faster AI Transcription

Reviewed byNidhi Govil

5 Sources

Share

Google has launched Gemini 3.5 Transcribe, an AI-powered speech-to-text model that delivers 70% faster transcription with a 5.5% error rate. The model automatically removes filler words, supports 85 languages, and handles custom vocabulary while converting raw audio into polished, formatted text.

News article

Google Releases Gemini 3.5 Transcribe for Enhanced Speech-to-Text

Google has announced Gemini 3.5 Transcribe, an AI-powered speech-to-text model designed to transform voice input into polished, formatted text

1

2

. The new model already powers the Rambler feature on Pixel 11 devices and the Gemini app on macOS, with broader rollout planned across the Google ecosystem

3

. This release comes as the industry awaits the promised Gemini 3.5 Pro model originally slated for June

2

. The AI transcription model represents a significant upgrade over Google's previous voice-to-text engine, Chirp 3, particularly in multilingual performance and accuracy

2

5

.

Performance Improvements Drive Faster Voice-to-Text Input

Gemini 3.5 Transcribe delivers approximately 70% faster processing from voice to final transcribed text compared to its predecessor

1

. The live speech recognition error rate has dropped to 5.5%, down from Chirp 3's 7.32%

1

. While this improvement may seem incremental, error reduction in speech-to-text systems significantly impacts user experience, as manually correcting transcription mistakes disrupts workflow. The model supports more than 85 languages and can detect them automatically, making it suitable for multilingual environments

2

3

. Unlike conventional speech recognition models that struggle with background noise and complex jargon, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished text

5

.

AI Model Removes Filler Words and Handles Custom Vocabulary

The new model goes beyond simple transcription by intelligently processing speech patterns. Gemini 3.5 Transcribe automatically removes filler words like "um" and "uh" from transcriptions, turning ramblings into structured text

1

2

. Users can provide custom vocabulary to the model, allowing it to adapt to unique spelling requirements and specialized jargon without manual editing

2

. The system can also edit text on the fly when users correct themselves mid-speech

1

. For pre-recorded audio, the model offers speaker attribution for up to three speakers along with word-level timestamps

2

3

. Google claims the model excels at capturing alphanumeric strings such as order numbers and postal codes

3

.

Gemini Live Updates Improve Interruption Handling

Alongside Gemini 3.5 Transcribe, Google introduced Gemini 3.5 Live and Gemini 3.5 Live Experimental models

2

4

. Gemini 3.5 Live addresses a persistent frustration with voice AI by better handling mid-sentence interruptions, processing live visuals, and blending multiple languages on the fly

4

. The experimental version narrates its progress step by step in real time while tackling complex reasoning tasks

2

. These updates aim to deliver more natural and responsive conversations, with the ability to access background tools—a capability currently missing from Gemini Live

4

.

Developer Access Through Gemini API and Integration Plans

Developers can now access Gemini 3.5 Transcribe through the Gemini API in Google AI Studio and Antigravity, Google's agentic development platform

1

3

. The model is available through two separate APIs: real-time streaming for interactive voice apps with sub-second latency, and pre-recorded audio processing for transcribing meetings and call logs

5

. AI Studio's build model now includes Gemini 3.5 Transcribe, enabling developers to create applications using AI-optimized voice-to-text

1

. In the Gemini app for macOS, the model combines natural voice input with screen awareness, allowing it to summarize local files, repurpose text between apps, or generate images using spoken instructions

4

.

Chrome Integration Expands Voice Input Capabilities

Google plans to bring Gemini 3.5 Transcribe to the Chrome browser soon, enabling users to input text into any web field using voice

1

3

. This integration will allow voice-based replies to emails, social media posts, and prompts for Gemini itself

3

4

. The Rambler feature on Android, currently available on Pixel 11 phones, will expand to more Gemini Intelligence devices later this year

1

. Gboard's Rambler uses Gemini 3.5 Transcribe to convert rambling speech into well-formatted text, with voice-based editing capabilities

4

. The rollout has begun for macOS Gemini app users and Android users in select countries and languages, with broader availability expected in the coming days

2

4

. Future integration is planned for Search Live, Gemini Live, Docs, Keep, and Gmail

3

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved