
What Gemini 3.8 Flash TTS actually is
Gemini 3.8 Flash TTS (gemini-3.8-flash-tts) is Google's current text-to-speech model, and unlike a lot of the audio models floating around, it is stable rather than preview. The id carries no -preview suffix, and Google lists it as the recommended replacement for the older gemini-3.1-flash-tts-preview. It takes text in and produces audio out, nothing else, which is exactly what you want from a dedicated speech model.
The spec sheet is generous. You get 30 prebuilt studio voices (Zephyr, Puck, Kore, Fenrir and friends), each with a descriptor like Bright or Firm, plus an Extended Voice Library of hundreds more you can pull with client.voices.list(). It speaks 130 languages with auto-detected input, and you can steer style, accent, pace and tone with natural-language prompts, or drop inline tags like <laugh> and <sigh> for point-in-time vocal events. It also ships voice design (build a persona from a text description) and voice cloning from a reference clip.
The one spec that bites in production: multi-speaker is capped at two speakers per request. A three-person podcast or a group scene means synthesizing each turn separately and stitching the audio yourself. Output defaults to 24 kHz mono WAV, with mulaw and alaw available for telephony. This sits alongside the text-only Gemini 3.8 Flash model, and the two are billed completely differently, which is where most of the confusion starts.
Gemini 3.8 Flash TTS pricing: the full table
Here is every tier, in USD per 1M tokens, at the 2026 promotional rates. The 2027 rate is in parentheses because, as we will get to, everything doubles.
| Tier | Text input / 1M | Audio output / 1M | Notes |
|---|---|---|---|
| Standard | $0.50 ($1.00) | $9.00 ($18.00) | The default rate |
| Batch | $0.25 ($0.50) | $4.50 ($9.00) | ~50% off, async jobs |
| Flex | $0.25 ($0.50) | $4.50 ($9.00) | Same rate, cheaper caching |
| Priority | $0.90 ($1.80) | $16.20 ($32.40) | ~1.8x, latency-sensitive |
| Free | Free | Free | Standard/priority via AI Studio |
All of these come straight from Google's Gemini API pricing page. Context caching is supported and runs from $0.125 per 1M input-caching tokens on standard, plus a $0.50 per 1M per hour storage fee. The audio output number is the one that matters, because that is what scales with how much speech you generate. Text input is almost a rounding error by comparison.
What is quietly interesting is where $9.00 sits in Google's own lineup. It is cheaper than the Gemini 2.5 Flash TTS it succeeds ($10 audio out), and less than half the price of the 2.5 Pro and 3.1 Flash TTS preview models ($20 audio out). There is also a cheaper sibling, Flash-Lite TTS, at $6 audio out if you can live with 101 languages instead of 130.

What that actually costs you
Per-token pricing is hard to feel, so let's turn it into real money. Audio bills at 25 tokens per second, which means 1,500 tokens a minute and 90,000 tokens an hour. At the 2026 standard rate of $9.00 per 1M, an hour of generated speech is about $0.81. Add the text you fed in (a full hour of narration is only around 12,000 input tokens) and it is well under a cent more.

A few worked examples at the standard rate:
- A 30-minute podcast episode: about $0.41 in audio.
- An 8-hour audiobook: roughly $6.48.
- 10,000 support IVR messages at 15 seconds each: 3.75M audio tokens, so about $33.75.
Run any of those as an overnight batch job and you halve it, which makes the audiobook closer to $3.24. For high-volume, non-interactive generation, batch is the obvious play. The one place the sticker doesn't tell the whole story is anything a human waits on in real time, where you are paying the priority rate and eating latency you cannot fully control.
The catch: prices double on 1 January 2027
This is the part to write on a sticky note. The $0.50 and $9.00 figures are introductory promotional rates. On 1 January 2027, standard text input goes to $1.00 per 1M and audio output to $18.00 per 1M. Every other tier doubles in lockstep: batch and flex audio output to $9.00, priority audio output to $32.40.

So that $0.81 hour of audio becomes about $1.62 in 2027, and the 8-hour audiobook goes from $6.48 to $12.96. It is the same move Google pulled on the text models, and it is a fair one as long as you see it coming. If you are budgeting anything past this year, double the numbers in this post and plan from there. It also nudges you toward locking in whatever caching and batch savings you can now.
How it compares to ElevenLabs, OpenAI, and the rest
Comparing voice models is genuinely tricky, because the vendors don't bill the same way. Amazon, Azure and Deepgram charge per character of input text. OpenAI and Google charge per audio output token, which scales with the length of the generated speech, not the script. ElevenLabs sells subscription credits. To put them on one axis, I converted everything to an estimated cost per hour of speech, using Amazon's own anchor that 1M characters is roughly 23 hours of audio. Treat the per-hour figures as estimates, not contract terms, because they move with speaking pace.
| Vendor | Flagship model | Native price | ~ $/hour of audio | Free tier |
|---|---|---|---|---|
| Gemini 3.8 Flash TTS | gemini-3.8-flash-tts | $9 / 1M audio tokens | ~$0.81 | Free in AI Studio |
| Amazon Polly (Neural) | Neural | $16 / 1M chars | ~$0.70 | 1M chars/mo (12 mo) |
| Azure AI Speech | Neural / HD Flash | $15 / 1M chars | ~$0.65 | 0.5M chars/mo |
| OpenAI | gpt-4o-mini-tts | $12 / 1M audio tokens | ~$0.90 | None |
| Amazon Polly (Generative) | Generative | $30 / 1M chars | ~$1.30 | 100k chars/mo |
| Deepgram | Aura-2 | $30 / 1M chars | ~$1.30 | $200 credit |
| ElevenLabs | Eleven v3 / Flash | ~5c/min (Business) | ~$3.00 | 10k credits/mo |
The takeaway that surprised me: Gemini 3.8 Flash TTS is priced like a commodity neural voice while being one of the expressive generative models. Amazon and Azure's cheaper rows are their older neural voices; their generative-quality tier (Polly Generative) is $1.30 an hour, and Deepgram's flagship lands in the same place. ElevenLabs, still the reference point for voice quality, is roughly four times the price per hour. If cost per hour is your deciding factor and you want expressive output, Gemini is hard to beat right now, at least until the 2027 reset pulls it up to about $1.62.
For the OpenAI side of this, my write-ups on the OpenAI realtime API and gpt-realtime-mini pricing go deeper on where token-billed voice makes sense.
Is the voice any good?
Price only matters if the output is usable, and this is where early reactions get more mixed. The rollout drew some grumbling on quality. One developer on Hacker News put it bluntly:
"I don't know, I guess their roll out is a bit sloppy. It's a bit of a shame, though, since the voices which are available all sound like generic Gemini voices to me. Nothing stands out."
That is one person's ear, and voice preference is subjective, but it lines up with a pattern: Google's TTS is cheap and broad, while ElevenLabs still wins the "does this specific voice sound remarkable" test for a lot of people. Some users are sticking with what they know for now:
"For on the go, I've been using ElevenReader. Technically not free but their free tier has been plenty for me."
So the honest read is this. If you need cheap, multilingual, good-enough narration at scale, Gemini 3.8 Flash TTS is excellent value. If a single, distinctive, brand-defining voice is the product, audition it carefully against ElevenLabs before you commit, and don't assume the price gap means a quality gap in your favor.
Where a cheap voice model fits, and where it doesn't
Here is the reframe I'd want a buyer to leave with. A TTS model is infrastructure. It is a fantastic, cheap way to give an app a voice: an audiobook pipeline, an in-product narrator, an accessibility layer, IVR prompts. For those, $0.81 an hour is a genuinely great deal.
But the moment the use case is customer support, the model is the easy 10% of the problem. The hard part is everything around it: pulling the right answer from your knowledge base, calling the tools that actually resolve a ticket, and getting tested against your real history before it ever speaks to a customer. That is the difference between infrastructure and an employee. Gemini TTS is the former. It will not know your refund policy or triage a ticket, and stitching all of that together yourself is a project, not an API call.
This is the frame we work in every day. We have spent years putting AI on live support queues, and the thing that separates a demo from a rollout is never the model, it is the plumbing and the testing. That is why an AI helpdesk agent is a very different purchase from a voice endpoint, even when both say "AI" on the tin. If you are weighing the human-versus-automation math generally, our AI agent vs human agent cost breakdown is a better starting point than a per-token rate.
Try eesel
eesel is an AI teammate platform, and the current roster includes a ready-to-work AI helpdesk agent that joins your existing support queue. Where Gemini TTS is a model you have to build around, eesel arrives already knowing how to connect to your helpdesk, learn from your past tickets, and get simulated against real historical conversations so you know its resolution rate before it goes live.

If you like working in a terminal, the eesel CLI drives the same teammate and workspace programmatically. You can operate it by hand, wire it into scripts, or let a coding agent like Claude Code or Cursor run it headlessly, so the AI you buy is available both in the dashboard and through the API, the same way you would reach for a model like Gemini TTS. You can try eesel for free and simulate it against your own tickets first.
Frequently Asked Questions
How much does Gemini 3.8 Flash TTS cost?
Is there a free tier for Gemini 3.8 Flash TTS?
How does Gemini 3.8 Flash TTS pricing compare to ElevenLabs?
Why does the Gemini 3.8 Flash TTS price double in January 2027?
Is Gemini 3.8 Flash TTS good enough for customer support voice?

Article by
Rama Adi Nugraha
Rama is a software engineer at eesel AI with two years of experience writing about B2B SaaS, AI tools, and customer support technology. Based in Bali, Indonesia, he brings a developer's perspective to product comparisons — cutting through marketing copy to what the integrations and APIs actually do.








