8 Sources
[1]
Google announces Gemini 3.5 Transcribe for AI-powered speech-to-text
While we wait (possibly in vain) for Gemini 3.5 Pro to launch, Google is releasing a different model in the 3.5 branch. The company has announced Gemini 3.5 Transcribe, an AI model designed to streamline voice input by editing out "ums" and corrections, outputting polished AI text. This model
[2]
Google's new AI transcription edits out your 'ums' and 'ahs'
Google has updated Gemini Audio with some new Gemini 3.5 models, introducing new transcription capabilities that automatically detect specialized jargon and more than 85 languages. Gemini 3.5 Live, 3.5 Live Experimental, and 3.5 Transcribe are designed to provide better precision for Google's
[3]
Google says its latest Gemini transcription model can turn your ramblings into structured text - Engadget
The company says Gemini 3.5 Transcribe will soon let you use speech-to-text in any web field in Chrome. Google has revealed a new AI audio model that it says offers improved speech recognition and transcription. Gemini 3.5 Transcribe is joining Gemini 3.5 Live and Gemini 3.5 Live Experimental in
[4]
Google just fixed one of the most annoying things about talking to AI -- Gemini Live can finally handle interruptions
* Gemini Live is being upgraded to Gemini Live 3.5, which handles interruptions better * The new model can also process live visuals and blend multiple languages on the fly * It also promises more natural and responsive conversations and the ability to access background tools Google has upgraded
[5]
Intelligent transcription with Gemini 3.5 Transcribe
Luke Leonhard Chief of Staff, Gemini Audio, on behalf of Gemini Audio Team Today, we're introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions. Unlike conventional speech recognition models that struggle with background noise,
[6]
Google launches Gemini 3.5 Transcribe for audio transcription
Google has introduced Gemini 3.5 Transcribe, an AI audio model that the company says improves speech recognition, transcription accuracy and multilingual speech-to-text. Google said the model can automatically detect more than 85 languages and convert unstructured speech into formatted text. The
[7]
Gemini 3.5 Transcribe: How AI Turns Long Conversations into Searchable Text
Word-level timestamps and speaker identification turn plain audio into organized records. Users can jump straight to a keyword, moment, or speaker. There is no need to replay an entire recording. Google measured the model's performance using artificial analysis benchmarks and the FLEURS dataset,
[8]
Google unveils Gemini 3.5 Transcribe speech-to-text model By Investing.com
Investing.com - Google (NASDAQ:GOOGL) launched Gemini 3.5 Transcribe on Wednesday, a speech-to-text model designed for real-time transcription with improved accuracy and reduced latency compared to its predecessor. The model converts audio into formatted text and handles background noise,
Share
Copy Link
Google has launched Gemini 3.5 Transcribe, an AI-powered speech-to-text model that delivers 70% faster transcription with a 5.5% error rate. The model automatically removes filler words, supports 85 languages, and handles custom vocabulary while converting raw audio into polished, formatted text.

Google has announced Gemini 3.5 Transcribe, an AI-powered speech-to-text model designed to transform voice input into polished, formatted text
1
2
. The new model already powers the Rambler feature on Pixel 11 devices and the Gemini app on macOS, with broader rollout planned across the Google ecosystem3
. This release comes as the industry awaits the promised Gemini 3.5 Pro model originally slated for June2
. The AI transcription model represents a significant upgrade over Google's previous voice-to-text engine, Chirp 3, particularly in multilingual performance and accuracy2
5
.Gemini 3.5 Transcribe delivers approximately 70% faster processing from voice to final transcribed text compared to its predecessor
1
. The live speech recognition error rate has dropped to 5.5%, down from Chirp 3's 7.32%1
. While this improvement may seem incremental, error reduction in speech-to-text systems significantly impacts user experience, as manually correcting transcription mistakes disrupts workflow. The model supports more than 85 languages and can detect them automatically, making it suitable for multilingual environments2
3
. Unlike conventional speech recognition models that struggle with background noise and complex jargon, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished text5
.The new model goes beyond simple transcription by intelligently processing speech patterns. Gemini 3.5 Transcribe automatically removes filler words like "um" and "uh" from transcriptions, turning ramblings into structured text
1
2
. Users can provide custom vocabulary to the model, allowing it to adapt to unique spelling requirements and specialized jargon without manual editing2
. The system can also edit text on the fly when users correct themselves mid-speech1
. For pre-recorded audio, the model offers speaker attribution for up to three speakers along with word-level timestamps2
3
. Google claims the model excels at capturing alphanumeric strings such as order numbers and postal codes3
.Alongside Gemini 3.5 Transcribe, Google introduced Gemini 3.5 Live and Gemini 3.5 Live Experimental models
2
4
. Gemini 3.5 Live addresses a persistent frustration with voice AI by better handling mid-sentence interruptions, processing live visuals, and blending multiple languages on the fly4
. The experimental version narrates its progress step by step in real time while tackling complex reasoning tasks2
. These updates aim to deliver more natural and responsive conversations, with the ability to access background tools—a capability currently missing from Gemini Live4
.Related Stories
Developers can now access Gemini 3.5 Transcribe through the Gemini API in Google AI Studio and Antigravity, Google's agentic development platform
1
3
. The model is available through two separate APIs: real-time streaming for interactive voice apps with sub-second latency, and pre-recorded audio processing for transcribing meetings and call logs5
. AI Studio's build model now includes Gemini 3.5 Transcribe, enabling developers to create applications using AI-optimized voice-to-text1
. In the Gemini app for macOS, the model combines natural voice input with screen awareness, allowing it to summarize local files, repurpose text between apps, or generate images using spoken instructions4
.Google plans to bring Gemini 3.5 Transcribe to the Chrome browser soon, enabling users to input text into any web field using voice
1
3
. This integration will allow voice-based replies to emails, social media posts, and prompts for Gemini itself3
4
. The Rambler feature on Android, currently available on Pixel 11 phones, will expand to more Gemini Intelligence devices later this year1
. Gboard's Rambler uses Gemini 3.5 Transcribe to convert rambling speech into well-formatted text, with voice-based editing capabilities4
. The rollout has begun for macOS Gemini app users and Android users in select countries and languages, with broader availability expected in the coming days2
4
. Future integration is planned for Search Live, Gemini Live, Docs, Keep, and Gmail3
.Summarized by
Navi
[3]
[4]
09 Sept 2025•Technology

16 Sept 2026•Technology

10 Jun 2026•Technology

1
Science and Research

2
Technology

3
Policy and Regulation
