2 Sources
[1]
Gemini 3.8 TTS can design voices from prompts and clone them
"Today, we're introducing two new text-to-speech models to the Gemini family, transforming voice generation from static presets into a dynamic creative studio," Google wrote in its announcement. Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on Wednesday. Both are rolling out
[2]
Google launches two benchmark-topping speech generation models
Google launches two benchmark-topping speech generation models Google LLC today made two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, available through its cloud platform. The algorithms have highly similar application programming interfaces, which makes using
Share
Copy Link
Google launched Gemini 3.8 Flash TTS and Flash-Lite TTS, two text-to-speech models that transform voice generation from static presets into dynamic creation. Flash TTS designs custom voices from plain language prompts across 100 languages and dialects, while Flash-Lite TTS offers cost-efficient, high-volume speech generation for dubbing and voice agents.

Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on Wednesday, introducing two text-to-speech models that shift voice generation from static presets into what the company calls a dynamic creative studio
1
. Both speech generation models are now rolling out through the Gemini API and Google AI Studio, joining the Gemini 3.8 family that Google launched earlier this month1
. Flash TTS targets creative applications such as games, audiobooks and podcasts, while Flash-Lite TTS serves as the cost-efficient option optimized for dubbing and voice agents1
2
. The algorithms have highly similar application programming interfaces, making it relatively simple for developers to use them side-by-side2
.With Flash TTS, users can design a new voice from scratch by describing its role, accent and characteristics in plain language. The model supports voice generation across more than 100 languages and dialects, with Flash TTS specifically handling 130 languages and Flash-Lite TTS supporting 101
1
2
. Google's examples of custom voice creation include a high-energy DJ from Melbourne and a Japanese dragon1
. Both models offer access to a library of more than 2,000 prepackaged voices, including regional varieties such as Mexican Spanish, Quebec French and Scots English1
2
. Custom voices can be saved for reuse across projects, and a voice remixing feature that adjusts timbre, pitch, pace and accent is coming soon1
.Flash TTS can recreate a voice from a 30-second audio sample, but Google requires users to supply a verbal consent recording from the voice owner
1
2
. The system checks that the consent recording matches the reference speaker before creating the voice, adding a layer of protection against unauthorized voice cloning1
. Developers can customize parameters such as an AI speaker's vocal timbre, accent and pacing when working with these AI-driven text-to-speech models2
. Google plans to add a third customization option down the line that will make it possible to create a new voice by modifying one of the prepackaged options2
.Both models let users direct delivery line by line with stage directions, enabling developers to add oratory cues to every line of the script that the models read out loud
1
2
. These cues generate audio elements such as non-lexical vocalizations and pacing shifts2
. The models support multi-speaker scenes from a single script and long-form audio lasting hours1
. Scripts can include cues such as laughs, sighs and short interjections like "mhm," expanding the range of dynamic voice generation capabilities1
.Every clip generated by Google's Gemini Audio models carries a SynthID watermark, according to the company
1
. The audio watermarking is inaudible to humans but can be picked up by AI detection tools2
. Replicated voices also carry C2PA content credentials, which specify when a file was generated, whether it has been modified since and related details1
2
. These measures add transparency and traceability to AI-generated audio content.Related Stories
Google says Flash TTS took first place on Hume AI's Voice Design Benchmark with a score of 71.4
1
. Flash TTS and Flash-Lite TTS ranked first and second on Hume AI's Overall Quality Index, demonstrating strong performance across multiple evaluation criteria1
2
. In blind preference tests on Voice Arena, a benchmark that measures text-to-speech algorithms' output quality based on human feedback, the models secured top positions in languages including Japanese, Hindi and Mexican Spanish1
2
. The market already has specialists such as ElevenLabs and India's Murf AI, and Adobe added speech generation to Firefly in August, making Google's entry a significant competitive move1
.Google named Figma, HeyGen, Wondercraft, Linguana, 99.co and Ollang as companies integrating the new models
1
. Flash TTS is also coming to Gemini Notebook, and Flash-Lite TTS to Google Vids1
. Access for enterprises through Gemini Enterprise is coming soon1
. According to Google staffers Leland Rechis and Alan Cowen, "These models enable creators, developers, and enterprises to create richer, more expressive audio experiences, while enabling improved user experiences in products like Gemini Notebook and Google Vids"2
. Flash TTS and Flash-Lite TTS are part of a broader lineup of audio processing models, as Google previously released algorithms optimized for voice agents, transcription and translation2
.Summarized by
Navi
[1]
[2]
16 Apr 2026•Technology

04 Jun 2025•Technology

31 Jan 2025•Technology

1
Technology

2
Policy and Regulation

3
Technology
