5 Sources
[1]
Google announces Gemini 3.5 Transcribe for AI-powered speech-to-text
While we wait (possibly in vain) for Gemini 3.5 Pro to launch, Google is releasing a different model in the 3.5 branch. The company has announced Gemini 3.5 Transcribe, an AI model designed to streamline voice input by editing out "ums" and corrections, outputting polished AI text. This model already powers the Gboard "Rambler" feature on the Pixel 11, but it's about to appear throughout the Google ecosystem. According to Google, Gemini 3.5 Transcribe is much faster and more accurate than its previous voice-to-text engine, known as Chirp 3. The new AI model should be about 70 percent faster from voice to final transcribed text, and the live-speech error rate has dropped to 5.5 percent. That's only a little better than Chirp 3, which Google measures at 7.32 percent. Still, it's a pain to fix typos when you're using voice input, so any improvement here is beneficial. The new model isn't just better at hearing the words you say; it's also allegedly better able to get at the heart of what you meant to say. As you speak, Gemini 3.5 Transcribe can remove the awkward "ums" and "uhs" that clutter your stream of consciousness. It can also edit text on the fly if you need to correct yourself and refer to your provided custom vocabulary to handle "specialized jargon." All this works in 85 languages and with up to three speakers in pre-recorded audio. The drawback, of course, is that you're relying on the AI to accurately get the gist of your speech. For short blocks of text, the model seems good at cleaning up inconsistencies and verbal stumbles (based on my testing with Rambler), but the AI does technically change the wording of what you said, and that may not be appropriate for all situations. Talk to the robot Gemini 3.5 Transcribe is already live in a few places, including Rambler in Gboard. However, that's still limited to Pixel 11 phones. Google says it will expand Rambler to "more Gemini Intelligence devices" later this year. If you're using the Gemini app on macOS, your voice input will also benefit from Gemini 3.5 Transcribe starting today. Developers will be seeing plenty of Gemini 3.5 Transcribe as well. Starting today, Antigravity will have the new model, with full access to screen context and chat history (with your permission). AI Studio's build model also has Gemini 3.5 Transcribe now, allowing you to vibe code apps with AI-optimized voice-to-text. Developers can also access the model from the Gemini API. If you're not using those specific apps or devices, you won't see Gemini 3.5 Transcribe just yet, but it's coming. Google says it plans to bring the AI transcription feature to the Chrome browser "soon." When it's available, you'll be able to input text into any web field. Use it to write emails, leave angry comments, or even prompt AI chatbots by voice.
[2]
Google's new AI transcription edits out your 'ums' and 'ahs'
Google has updated Gemini Audio with some new Gemini 3.5 models, introducing new transcription capabilities that automatically detect specialized jargon and more than 85 languages. Gemini 3.5 Live, 3.5 Live Experimental, and 3.5 Transcribe are designed to provide better precision for Google's voice-controlled AI features, without struggling with background noise or when your speech is interrupted. Gemini 3.5 Transcribe is a completely new addition to the Gemini family, and its introduction comes as we're still waiting for Google to release the Gemini 3.5 Pro model that it promised to roll out in June. Google says that 3.5 Transcribe "represents a major advancement from our previous transcription model, Chirp 3," especially regarding multilingual performance and wording error rates. The transcription model allows users to "edit naturally with just your voice," according to Google, and can automatically format text and remove filler words like "um" and "uh." Users can provide a customized vocabulary to the model, allowing 3.5 Transcribe to automatically adapt transcription to unique spelling requirements and specialized jargon to prevent those words from being edited manually. It can also attribute speech for up to three speakers in pre-recorded audio, alongside providing word-level timestamps. Transcribe is launching alongside 3.5 Live and 3.5 Live Experimental, which build on the existing speech recognition tech that powers Gemini's voice chat mode. Gemini 3.5 Live is better at handling mid-sentence interruptions, language recognition, and live visual processing, while Gemini 3.5 Live Experimental goes further by narrating its progress step by step in real time while it tackles reasoning on more complex tasks. These Gemini Audio updates are rolling out starting today in English for all macOS Gemini app users, and the Rambler dictation feature on Android in select countries and languages. It's also available for developers in public preview in the Gemini API via AI Studio and Antigravity. Google says that Chrome support is coming soon.
[3]
Google says its latest Gemini transcription model can turn your ramblings into structured text - Engadget
The company says Gemini 3.5 Transcribe will soon let you use speech-to-text in any web field in Chrome. Google has revealed a new AI audio model that it says offers improved speech recognition and transcription. Gemini 3.5 Transcribe is joining Gemini 3.5 Live and Gemini 3.5 Live Experimental in the Gemini Audio family. The company says Gemini 3.5 Transcribe is more adept at speech-to-text than earlier models, with greater precision and the ability to automatically detect more than 85 languages. It says the model can adapt a user's unstructured speech into formatted text. You can use it to make edits with voice commands, and it removes filler words from transcribed speech too. Google says the model can learn custom vocabulary and unique spellings, and is adept at capturing alphanumeric strings such as order numbers and postal codes. Moreover, the company claims Gemini 3.5 Transcribe can attribute speech to up to three speakers with word-level timestamps based on pre-recorded audio. That could be handy for, say, transcribing a podcast. It may not come as a surprise that Gemini 3.5 Transcribe powers the Rambler feature on Android devices like the Pixel 11-series phones as well as the Gemini app on macOS (where it can work with other Gemini models to carry out agentic tasks). Intriguingly, Google says you'll soon be able to use text-to-speech in any field on a web page in Chrome. For instance, you'll be able to dictate replies and posts, and use your voice to prompt Gemini. Gemini 3.5 Transcribe is also available in Google Antigravity, the company's agentic development platform, and it's coming to Search Live, Gemini Live, Docs, Keep and Gmail. Developers will be able to tap into the model via APIs.
[4]
Google just fixed one of the most annoying things about talking to AI -- Gemini Live can finally handle interruptions
* Gemini Live is being upgraded to Gemini Live 3.5, which handles interruptions better * The new model can also process live visuals and blend multiple languages on the fly * It also promises more natural and responsive conversations and the ability to access background tools Google has upgraded its Gemini Live voice mode with a new model, Gemini 3.5 Live, which it says delivers a major improvement over the existing Gemini 3.1 Live model. The new 3.5 model will be better at handling mid-sentence interruptions, processing live visuals, blending multiple languages on the fly, and triggering background tools. Gemini Live currently lacks the ability to access and use other tools, so this looks like a big step forward in its functionality. Gemini Live is already one of the best things about using Gemini on your phone, letting you have natural, human-like conversations with the Google AI, but it's not perfect. In fact, Gemini Live is pretty prone to errors. Its flaws are probably different depending on how your voice sounds, but for me, I find it occasionally thinks I'm talking Italian and switches languages, and sometimes if I interrupt it, then it gets confused and will reset itself back to its default voice, not the one I've chosen (the Ursa voice is my favorite). Hopefully the new Gemini 3.5 Live will see an end to these problems, and I'll bring you a hands-on test as soon as I have access to the new model. Smarter translation To coincide with the release of Gemini 3.5 Live, Google is also releasing two developer models -- Gemini 3.5 Live Experimental and Gemini 3.5 Transcribe, which developers can access in the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform. Some of these powerful smart transcription features are already making their way into consumer products. On Android, Gboard's Rambler feature uses Gemini 3.5 Transcribe to turn rambling speech into well-formatted text, automatically removing filler words such as "um" and "ah." You can then use your voice to correct misspellings, edit the text, or change its writing style. In the Gemini app for macOS, the model combines natural voice input with an awareness of what's on your screen. It can call on other Gemini models in the background to summarize local files, repurpose text between apps, or generate images at your cursor, all using spoken instructions. Perhaps most interestingly, the technology is coming soon to Chrome. You'll be able to talk to type in any text field on the web, whether you're replying to an email, drafting a social post, or writing a prompt for Gemini. Gemini 3.5 Live, Gemini 3.5 Live Experimental, and Gemini 3.5 Transcribe are all being announced today. It's not clear when they will roll out to everyone, but we'd expect it to be over the next few days. Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.
[5]
Intelligent transcription with Gemini 3.5 Transcribe
Luke Leonhard Chief of Staff, Gemini Audio, on behalf of Gemini Audio Team Today, we're introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions. Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text. Across our products like the Gemini app and on Android, we've seen consumers already benefiting from this transcription model with new voice capabilities like Rambler on Android and in the Gemini app on macOS. Now, developers can build similar capabilities with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform. We've built 3.5 Transcribe to plug seamlessly into your developer workflows, whether you're building voice agents, real-time captioning tools, or post-call analytics pipelines. The model is available across two separate APIs: * Real-time streaming: Delivers continuous, bidirectional streaming with sub-second latency for interactive voice apps via the Live API using . * Pre-recorded audio processing: Transcribes recorded audio, meetings, call logs, and more with speaker attribution and word-level timestamps via the Interactions API using . Get more precise and intelligent transcription Gemini 3.5 Transcribe is designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary, so you can execute tasks with your voice.
Share
Copy Link
Google has launched Gemini 3.5 Transcribe, an AI-powered speech-to-text model that delivers 70% faster transcription with a 5.5% error rate. The model automatically removes filler words, supports 85 languages, and handles custom vocabulary while converting raw audio into polished, formatted text.

Google has announced Gemini 3.5 Transcribe, an AI-powered speech-to-text model designed to transform voice input into polished, formatted text
1
2
. The new model already powers the Rambler feature on Pixel 11 devices and the Gemini app on macOS, with broader rollout planned across the Google ecosystem3
. This release comes as the industry awaits the promised Gemini 3.5 Pro model originally slated for June2
. The AI transcription model represents a significant upgrade over Google's previous voice-to-text engine, Chirp 3, particularly in multilingual performance and accuracy2
5
.Gemini 3.5 Transcribe delivers approximately 70% faster processing from voice to final transcribed text compared to its predecessor
1
. The live speech recognition error rate has dropped to 5.5%, down from Chirp 3's 7.32%1
. While this improvement may seem incremental, error reduction in speech-to-text systems significantly impacts user experience, as manually correcting transcription mistakes disrupts workflow. The model supports more than 85 languages and can detect them automatically, making it suitable for multilingual environments2
3
. Unlike conventional speech recognition models that struggle with background noise and complex jargon, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished text5
.The new model goes beyond simple transcription by intelligently processing speech patterns. Gemini 3.5 Transcribe automatically removes filler words like "um" and "uh" from transcriptions, turning ramblings into structured text
1
2
. Users can provide custom vocabulary to the model, allowing it to adapt to unique spelling requirements and specialized jargon without manual editing2
. The system can also edit text on the fly when users correct themselves mid-speech1
. For pre-recorded audio, the model offers speaker attribution for up to three speakers along with word-level timestamps2
3
. Google claims the model excels at capturing alphanumeric strings such as order numbers and postal codes3
.Alongside Gemini 3.5 Transcribe, Google introduced Gemini 3.5 Live and Gemini 3.5 Live Experimental models
2
4
. Gemini 3.5 Live addresses a persistent frustration with voice AI by better handling mid-sentence interruptions, processing live visuals, and blending multiple languages on the fly4
. The experimental version narrates its progress step by step in real time while tackling complex reasoning tasks2
. These updates aim to deliver more natural and responsive conversations, with the ability to access background tools—a capability currently missing from Gemini Live4
.Related Stories
Developers can now access Gemini 3.5 Transcribe through the Gemini API in Google AI Studio and Antigravity, Google's agentic development platform
1
3
. The model is available through two separate APIs: real-time streaming for interactive voice apps with sub-second latency, and pre-recorded audio processing for transcribing meetings and call logs5
. AI Studio's build model now includes Gemini 3.5 Transcribe, enabling developers to create applications using AI-optimized voice-to-text1
. In the Gemini app for macOS, the model combines natural voice input with screen awareness, allowing it to summarize local files, repurpose text between apps, or generate images using spoken instructions4
.Google plans to bring Gemini 3.5 Transcribe to the Chrome browser soon, enabling users to input text into any web field using voice
1
3
. This integration will allow voice-based replies to emails, social media posts, and prompts for Gemini itself3
4
. The Rambler feature on Android, currently available on Pixel 11 phones, will expand to more Gemini Intelligence devices later this year1
. Gboard's Rambler uses Gemini 3.5 Transcribe to convert rambling speech into well-formatted text, with voice-based editing capabilities4
. The rollout has begun for macOS Gemini app users and Android users in select countries and languages, with broader availability expected in the coming days2
4
. Future integration is planned for Search Live, Gemini Live, Docs, Keep, and Gmail3
.Summarized by
Navi
[3]
[4]
29 Jul 2026•Technology

10 Jun 2026•Technology

12 Dec 2025•Technology

1
Technology

2
Policy and Regulation

3
Technology
