Gemini 3.8 Flash TTS

Turn a script into a performed voice track: Gemini 3.8 Flash TTS shapes emotion, two-voice exchanges and 130 languages from one prompt.

Gemini 3.8 Flash TTS
Steer delivery in real time with the expressive tier, or keep costs low on bulk audio with Flash-Lite TTS
AI Video Prompt Generator

Support

Pro AI Tools

Explore elite tools including the Suno AI Music Generator.

placeholder hero

A New Standard for AI Voice: What Gemini 3.8 Flash TTS Brings

Launched on September 23, 2026 by Google, the family pairs a detail-rich creative model with a lean engine built for bulk narration.

  • Two Models, Two Very Different Jobs
    The flagship tier chases nuance and long recordings, while Flash-Lite keeps per-minute spending down for large speech batches.
  • Directing the Read, Not Choosing a Preset
    Style notes attached to each turn, structured speech metadata and inline vocal tags govern tone, tempo, feeling and accent.
  • Voice Design Plus Consented Replication
    Sketch a voice with a plain-English prompt, or mirror a real speaker using a reference sample together with their consent recording.

Four Rules for Prompting Gemini 3.8 Flash TTS

Follow these steps to keep spoken text verbatim and let the performance metadata carry the acting.

What Gemini 3.8 Flash TTS Can Actually Do

From turn-by-turn acting control to 130-language coverage, these are the strengths that set the flagship Gemini TTS model apart.

Turn-by-Turn Performance Control

Attach a style to each line and drop in laughs, sighs, coughs, breaths or pauses — it feels like coaching a performer rather than picking a stock voice.

Voices Built from a Written Description

Describe age, character, accent, timbre and role in plain words, with more than 2,000 ready-made voices reachable through the Voices endpoint.

Cloning Behind a Consent Check

You need a clear reference sample plus a consent clip from that same adult, with SynthID watermarking and C2PA credentials on the output.

Conversations for Two Speakers

Write both sides of a podcast, lesson, product walkthrough or game scene and the model handles the back-and-forth without stitching clips by hand.

Consistent Tone Across Long Reads

Google reports that identity, timbre, loudness and room tone hold steady across multi-minute narration and extended exchanges.

Global Coverage and Local Accents

The flagship handles 130 languages versus 101 on Flash-Lite, including regional accents, minority dialects and IPA overrides.

FAQ

Gemini 3.8 Flash TTS: Frequently Asked Questions

Quick answers about cost per minute, picking a tier, benchmark scores and the safety rules around voice replication.

1

How is Gemini 3.8 Flash TTS priced?

Roughly 1.35 cents for every minute of audio, based on launch rates of $0.50 per million input tokens and $9 per million output tokens.

2

Which tier should I pick — Flash or Flash-Lite?

Choose the flagship for acting detail and long recordings; choose Flash-Lite when you need volume and fast turnaround.

3

How does it score against rival voice models?

Google cites 71.4 on Hume's Voice Design Benchmark, and Voice Arena ranks it second with 1,260 Elo.

4

What changed from the Gemini 3.1 Flash TTS Preview?

Flash-Lite takes over from the 3.1 preview and drops audio output pricing from $20 to $6 per million tokens.

5

Can I clone a voice, and what safeguards exist?

Yes, but only with a reference sample and a consent recording supplied by the same adult speaker.

6

Why is the model reading my stage directions aloud?

Because the script is treated as verbatim text — shift any lasting direction into the speech metadata field instead.

Hear Gemini 3.8 Flash TTS Read Your Own Script

Try both tiers in the Gemini API or Google AI Studio — moving between Gemini 3.8 Flash TTS and Flash-Lite TTS takes nothing more than a one-line identifier change. Weigh batch against priority inference before you lock in a production budget.