Gemini 3.8 Flash TTS

Hand Gemini 3.8 Flash TTS a written script and it returns a fully acted voice track, with mood, pacing and 130 languages under your control.

Gemini 3.8 Flash TTS
Shape emotion and pacing take by take, or drop to Flash-Lite TTS for high-volume, low-cost runs
AI Video Prompt Generator

Support

Pro AI Tools

Explore elite tools including the Suno AI Music Generator.

placeholder hero

Gemini 3.8 Flash TTS: A Voice Model Built for Direction, Not Presets

The rollout introduced two tiers under one family: an expressive flagship for crafted audio and a lean engine for bulk narration.

  • A Shared Launch, Two Distinct Workloads
    Gemini 3.8 Flash TTS handles expressive, directed reads, while Flash-Lite TTS keeps per-minute costs down when you generate at scale.
  • You Direct the Read Instead of Choosing a Preset
    Style prompts per turn, structured speech metadata and inline vocal events together control tone, tempo, feeling and accent.
  • Design a Voice From Words, or Replicate With Consent
    Sketch a voice using plain description, or mirror a real speaker by supplying a reference clip alongside a matching consent file.

How to Prompt Gemini 3.8 Flash TTS for Clean Takes

Four habits that keep your transcript readable and let the performance metadata do the acting.

What Gemini 3.8 Flash TTS Can Actually Do

Performance control, voice design, dialogue staging and 130-language reach — the full toolkit inside the flagship Gemini TTS model.

Line-by-Line Performance Control

Per-turn styling plus inline laughs, sighs, coughs, breaths and pauses feels like coaching a performer rather than picking a stock voice.

Voice Design From Plain Description

Prompts can specify age, personality, accent, vocal texture and role, with 2,000+ ready-made voices available through the Voices endpoint.

Replication Guarded by Consent

You need a clean reference clip and a matching consent clip from the same adult speaker, plus SynthID watermarking and C2PA credentials.

Dialogue Built for Two Speakers

Write the exchange once for podcasts, lessons, product walkthroughs or game scenes — no hand-stitching separate lines together.

Steady Voice Across Long Narrations

Google reports consistent identity, timbre, loudness and room tone across multi-minute narration and drawn-out dialogue.

130 Languages With Regional Color

The flagship supports 130 languages versus Flash-Lite's 101, including regional accents, minority dialects and IPA overrides.

FAQ

Gemini 3.8 Flash TTS: Questions People Ask Most

Quick answers about Gemini 3.8 Flash TTS costs, picking between tiers, benchmark scores and safety rules.

1

What does Gemini 3.8 Flash TTS charge per audio minute?

Roughly 1.35 cents per audio minute, based on launch pricing of $0.50 per million input tokens and $9 per million output tokens.

2

When should I pick Flash-Lite TTS instead?

Choose the flagship for acting nuance and long-form pieces; choose Flash-Lite TTS for high volume and low latency.

3

How does it score against rival voice models?

Google cites 71.4 on Hume's Voice Design Benchmark, and Voice Arena ranks it second with 1,260 Elo.

4

What changed from the Gemini 3.1 Flash TTS Preview?

Flash-Lite TTS takes over from the 3.1 preview and lowers audio output pricing from $20 to $6 per million tokens.

5

Can I clone a voice, and what rules apply?

Replication requires a reference clip and a matching consent recording captured from the same adult speaker.

6

Why does the model read my stage directions aloud?

Your input is treated as a verbatim transcript — shift any lasting directions into the speech metadata fields.

Point Gemini 3.8 Flash TTS at Your Next Script

Try both tiers in the Gemini API or Google AI Studio — moving between Gemini 3.8 Flash TTS and Flash-Lite TTS takes nothing more than swapping a model ID. Weigh batch against priority inference before committing a budget.