The cheapest MiniMax H3 is here,$0.013/s

Gemini 3.8 Flash TTS

Give your script a real performance: direct emotion line by line, stage two speakers, and render speech in 130 languages with Gemini 3.8 Flash TTS.

Gemini 3.8 Flash TTS
Craft expressive speech with the flagship tier, or cut costs on volume work with Flash-Lite TTS
AI Video Prompt Generator

Support

Pro AI Tools

Explore elite tools including the Suno AI Music Generator.

placeholder hero

A Voice Engine Built for Directed Performance

Launched on September 23, 2026, the Gemini TTS family pairs a creative flagship with a lean, high-throughput tier built for bulk audio.

  • A Flagship and a Workhorse, Released Together
    Gemini 3.8 Flash TTS carries the expressive, long-form workload, while Flash-Lite TTS keeps spend low for high-volume speech jobs.
  • Direction Instead of Preset Picking
    Per-turn style notes, structured speech metadata and inline vocal events shape tone, pacing, emotion and accent as you write.
  • Designed Voices and Consented Replication
    Describe a voice in everyday language, or recreate a real speaker from a reference clip paired with a matching consent recording.

How to Prompt Gemini 3.8 Flash TTS for Clean Audio

Four habits that keep your transcript clean and let the performance metadata handle the acting.

What Gemini 3.8 Flash TTS Can Do

Performance direction, language coverage and built-in safeguards — a closer look at the flagship tier of the Gemini TTS line.

Turn-by-Turn Performance Control

Per-turn styling plus inline laughs, sighs, coughs, breaths and pauses — closer to coaching a voice actor than clicking a preset.

Voices Described in Plain English

Prompts set age range, personality, accent, vocal texture and role, supported by 2,000+ production voices through the Voices endpoint.

Replication Guarded by Consent

A clean reference recording plus a matching consent recording from the same adult speaker, with SynthID watermarking and C2PA credentials.

Two-Speaker Scenes Without Stitching

A script carries the whole conversation for podcasts, lessons, product demos and game scenes, so no manual line-by-line assembly is needed.

Stable Voice Across Long Recordings

Google documents steady voice identity, timbre, volume and room tone across multi-minute narration and extended dialogue.

130 Languages With Regional Accents

Flash TTS spans 130 languages against Flash-Lite's 101, with regional accents, minority dialects and IPA pronunciation overrides.

FAQ

Gemini 3.8 Flash TTS: Frequently Asked Questions

Quick answers on pricing, tier selection, benchmark results and the safety rules behind Gemini 3.8 Flash TTS.

1

What does Gemini 3.8 Flash TTS charge per minute of audio?

Roughly 1.35 cents per audio minute, based on launch rates of $0.50 per million input tokens and $9 per million output tokens.

2

When should I pick Flash-Lite TTS over Flash TTS?

Choose Flash TTS when acting nuance and long-form audio matter; reach for Flash-Lite TTS when you need bulk output and low latency.

3

How does it score against rival voice models?

Google reports 71.4 on Hume's Voice Design Benchmark, while Voice Arena places it second with 1,260 Elo.

4

What changed from the Gemini 3.1 Flash TTS Preview?

Flash-Lite TTS succeeds the 3.1 preview and brings audio output pricing down from $20 to $6 per million tokens.

5

Can I clone a voice, and which rules apply?

Replication requires a reference clip plus a matching consent recording from the same adult speaker.

6

Why is the model reading my stage directions aloud?

Your input is treated as a verbatim transcript, so move sustained directions into speech metadata instead.

Run Gemini 3.8 Flash TTS on Your Own Scripts

Try both tiers inside the Gemini API or Google AI Studio — moving from Flash TTS to Flash-Lite TTS takes nothing more than a new model identifier. Weigh batch against priority inference before locking in a production budget.