Gemini 3.8 Flash TTS
Give your script a real performance: direct emotion line by line, stage two speakers, and render speech in 130 languages with Gemini 3.8 Flash TTS.
Support
Pro AI Tools
Explore elite tools including the Suno AI Music Generator.
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1
FLUX 3 Video Generator

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

A Voice Engine Built for Directed Performance
Launched on September 23, 2026, the Gemini TTS family pairs a creative flagship with a lean, high-throughput tier built for bulk audio.
- A Flagship and a Workhorse, Released TogetherGemini 3.8 Flash TTS carries the expressive, long-form workload, while Flash-Lite TTS keeps spend low for high-volume speech jobs.
- Direction Instead of Preset PickingPer-turn style notes, structured speech metadata and inline vocal events shape tone, pacing, emotion and accent as you write.
- Designed Voices and Consented ReplicationDescribe a voice in everyday language, or recreate a real speaker from a reference clip paired with a matching consent recording.
How to Prompt Gemini 3.8 Flash TTS for Clean Audio
Four habits that keep your transcript clean and let the performance metadata handle the acting.
What Gemini 3.8 Flash TTS Can Do
Performance direction, language coverage and built-in safeguards — a closer look at the flagship tier of the Gemini TTS line.
Turn-by-Turn Performance Control
Per-turn styling plus inline laughs, sighs, coughs, breaths and pauses — closer to coaching a voice actor than clicking a preset.
Voices Described in Plain English
Prompts set age range, personality, accent, vocal texture and role, supported by 2,000+ production voices through the Voices endpoint.
Replication Guarded by Consent
A clean reference recording plus a matching consent recording from the same adult speaker, with SynthID watermarking and C2PA credentials.
Two-Speaker Scenes Without Stitching
A script carries the whole conversation for podcasts, lessons, product demos and game scenes, so no manual line-by-line assembly is needed.
Stable Voice Across Long Recordings
Google documents steady voice identity, timbre, volume and room tone across multi-minute narration and extended dialogue.
130 Languages With Regional Accents
Flash TTS spans 130 languages against Flash-Lite's 101, with regional accents, minority dialects and IPA pronunciation overrides.
Gemini 3.8 Flash TTS: Frequently Asked Questions
Quick answers on pricing, tier selection, benchmark results and the safety rules behind Gemini 3.8 Flash TTS.
What does Gemini 3.8 Flash TTS charge per minute of audio?
Roughly 1.35 cents per audio minute, based on launch rates of $0.50 per million input tokens and $9 per million output tokens.
When should I pick Flash-Lite TTS over Flash TTS?
Choose Flash TTS when acting nuance and long-form audio matter; reach for Flash-Lite TTS when you need bulk output and low latency.
How does it score against rival voice models?
Google reports 71.4 on Hume's Voice Design Benchmark, while Voice Arena places it second with 1,260 Elo.
What changed from the Gemini 3.1 Flash TTS Preview?
Flash-Lite TTS succeeds the 3.1 preview and brings audio output pricing down from $20 to $6 per million tokens.
Can I clone a voice, and which rules apply?
Replication requires a reference clip plus a matching consent recording from the same adult speaker.
Why is the model reading my stage directions aloud?
Your input is treated as a verbatim transcript, so move sustained directions into speech metadata instead.
Run Gemini 3.8 Flash TTS on Your Own Scripts
Try both tiers inside the Gemini API or Google AI Studio — moving from Flash TTS to Flash-Lite TTS takes nothing more than a new model identifier. Weigh batch against priority inference before locking in a production budget.
