Gemini 3.8 Flash TTS
Turn a script into a performed voice track: Gemini 3.8 Flash TTS shapes emotion, two-voice exchanges and 130 languages from one prompt.
Support
Pro AI Tools
Explore elite tools including the Suno AI Music Generator.
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1
FLUX 3 Video Generator

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

A New Standard for AI Voice: What Gemini 3.8 Flash TTS Brings
Launched on September 23, 2026 by Google, the family pairs a detail-rich creative model with a lean engine built for bulk narration.
- Two Models, Two Very Different JobsThe flagship tier chases nuance and long recordings, while Flash-Lite keeps per-minute spending down for large speech batches.
- Directing the Read, Not Choosing a PresetStyle notes attached to each turn, structured speech metadata and inline vocal tags govern tone, tempo, feeling and accent.
- Voice Design Plus Consented ReplicationSketch a voice with a plain-English prompt, or mirror a real speaker using a reference sample together with their consent recording.
Four Rules for Prompting Gemini 3.8 Flash TTS
Follow these steps to keep spoken text verbatim and let the performance metadata carry the acting.
What Gemini 3.8 Flash TTS Can Actually Do
From turn-by-turn acting control to 130-language coverage, these are the strengths that set the flagship Gemini TTS model apart.
Turn-by-Turn Performance Control
Attach a style to each line and drop in laughs, sighs, coughs, breaths or pauses — it feels like coaching a performer rather than picking a stock voice.
Voices Built from a Written Description
Describe age, character, accent, timbre and role in plain words, with more than 2,000 ready-made voices reachable through the Voices endpoint.
Cloning Behind a Consent Check
You need a clear reference sample plus a consent clip from that same adult, with SynthID watermarking and C2PA credentials on the output.
Conversations for Two Speakers
Write both sides of a podcast, lesson, product walkthrough or game scene and the model handles the back-and-forth without stitching clips by hand.
Consistent Tone Across Long Reads
Google reports that identity, timbre, loudness and room tone hold steady across multi-minute narration and extended exchanges.
Global Coverage and Local Accents
The flagship handles 130 languages versus 101 on Flash-Lite, including regional accents, minority dialects and IPA overrides.
Gemini 3.8 Flash TTS: Frequently Asked Questions
Quick answers about cost per minute, picking a tier, benchmark scores and the safety rules around voice replication.
How is Gemini 3.8 Flash TTS priced?
Roughly 1.35 cents for every minute of audio, based on launch rates of $0.50 per million input tokens and $9 per million output tokens.
Which tier should I pick — Flash or Flash-Lite?
Choose the flagship for acting detail and long recordings; choose Flash-Lite when you need volume and fast turnaround.
How does it score against rival voice models?
Google cites 71.4 on Hume's Voice Design Benchmark, and Voice Arena ranks it second with 1,260 Elo.
What changed from the Gemini 3.1 Flash TTS Preview?
Flash-Lite takes over from the 3.1 preview and drops audio output pricing from $20 to $6 per million tokens.
Can I clone a voice, and what safeguards exist?
Yes, but only with a reference sample and a consent recording supplied by the same adult speaker.
Why is the model reading my stage directions aloud?
Because the script is treated as verbatim text — shift any lasting direction into the speech metadata field instead.
Hear Gemini 3.8 Flash TTS Read Your Own Script
Try both tiers in the Gemini API or Google AI Studio — moving between Gemini 3.8 Flash TTS and Flash-Lite TTS takes nothing more than a one-line identifier change. Weigh batch against priority inference before you lock in a production budget.
