Google launched Gemini 3.8 Flash TTS and Flash-Lite TTS, two text-to-speech models that transform voice generation from static presets into dynamic creation. Flash TTS designs custom voices from plain language prompts across 100 languages and dialects, while Flash-Lite TTS offers cost-efficient, high-volume speech generation for dubbing and voice agents.

News article

Google Transforms Voice Generation with Gemini 3.8 Flash TTS Models

Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on Wednesday, introducing two text-to-speech models that shift voice generation from static presets into what the company calls a dynamic creative studio

1

. Both speech generation models are now rolling out through the Gemini API and Google AI Studio, joining the Gemini 3.8 family that Google launched earlier this month

1

. Flash TTS targets creative applications such as games, audiobooks and podcasts, while Flash-Lite TTS serves as the cost-efficient option optimized for dubbing and voice agents

1

2

. The algorithms have highly similar application programming interfaces, making it relatively simple for developers to use them side-by-side

2

.

Custom Voice Creation Across 100 Languages and Dialects

With Flash TTS, users can design a new voice from scratch by describing its role, accent and characteristics in plain language. The model supports voice generation across more than 100 languages and dialects, with Flash TTS specifically handling 130 languages and Flash-Lite TTS supporting 101

1

2

. Google's examples of custom voice creation include a high-energy DJ from Melbourne and a Japanese dragon

1

. Both models offer access to a library of more than 2,000 prepackaged voices, including regional varieties such as Mexican Spanish, Quebec French and Scots English

1

2

. Custom voices can be saved for reuse across projects, and a voice remixing feature that adjusts timbre, pitch, pace and accent is coming soon

1

.

Clone Voices from 30-Second Audio Samples with Consent Protocols

Flash TTS can recreate a voice from a 30-second audio sample, but Google requires users to supply a verbal consent recording from the voice owner

1

2

. The system checks that the consent recording matches the reference speaker before creating the voice, adding a layer of protection against unauthorized voice cloning

1

. Developers can customize parameters such as an AI speaker's vocal timbre, accent and pacing when working with these AI-driven text-to-speech models

2

. Google plans to add a third customization option down the line that will make it possible to create a new voice by modifying one of the prepackaged options

2

.

Dynamic Voice Generation Features for Multi-Speaker Scenes

Both models let users direct delivery line by line with stage directions, enabling developers to add oratory cues to every line of the script that the models read out loud

1

2

. These cues generate audio elements such as non-lexical vocalizations and pacing shifts

2

. The models support multi-speaker scenes from a single script and long-form audio lasting hours

1

. Scripts can include cues such as laughs, sighs and short interjections like "mhm," expanding the range of dynamic voice generation capabilities

1

.

SynthID Audio Watermarking and C2PA Content Credentials

Every clip generated by Google's Gemini Audio models carries a SynthID watermark, according to the company

1

. The audio watermarking is inaudible to humans but can be picked up by AI detection tools

2

. Replicated voices also carry C2PA content credentials, which specify when a file was generated, whether it has been modified since and related details

1

2

. These measures add transparency and traceability to AI-generated audio content.

Benchmark Performance on Hume AI and Voice Arena Tests

Google says Flash TTS took first place on Hume AI's Voice Design Benchmark with a score of 71.4

1

. Flash TTS and Flash-Lite TTS ranked first and second on Hume AI's Overall Quality Index, demonstrating strong performance across multiple evaluation criteria

1

2

. In blind preference tests on Voice Arena, a benchmark that measures text-to-speech algorithms' output quality based on human feedback, the models secured top positions in languages including Japanese, Hindi and Mexican Spanish

1

2

. The market already has specialists such as ElevenLabs and India's Murf AI, and Adobe added speech generation to Firefly in August, making Google's entry a significant competitive move

1

.

Integration Partners and Enterprise Rollout Plans

Google named Figma, HeyGen, Wondercraft, Linguana, 99.co and Ollang as companies integrating the new models

1

. Flash TTS is also coming to Gemini Notebook, and Flash-Lite TTS to Google Vids

1

. Access for enterprises through Gemini Enterprise is coming soon

1

. According to Google staffers Leland Rechis and Alan Cowen, "These models enable creators, developers, and enterprises to create richer, more expressive audio experiences, while enabling improved user experiences in products like Gemini Notebook and Google Vids"

2

. Flash TTS and Flash-Lite TTS are part of a broader lineup of audio processing models, as Google previously released algorithms optimized for voice agents, transcription and translation

2

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved