Gemini 3.8 Flash TTS pricing: what an hour of AI voice actually costs

Rama Adi Nugraha
Written by

Rama Adi Nugraha

Katelin Teen
Reviewed by

Katelin Teen

Last edited September 23, 2026

Expert Verified
Illustration of sound waves and price tags representing Gemini 3.8 Flash TTS pricing

What Gemini 3.8 Flash TTS actually is

Gemini 3.8 Flash TTS (gemini-3.8-flash-tts) is Google's current text-to-speech model, and unlike a lot of the audio models floating around, it is stable rather than preview. The id carries no -preview suffix, and Google lists it as the recommended replacement for the older gemini-3.1-flash-tts-preview. It takes text in and produces audio out, nothing else, which is exactly what you want from a dedicated speech model.

The spec sheet is generous. You get 30 prebuilt studio voices (Zephyr, Puck, Kore, Fenrir and friends), each with a descriptor like Bright or Firm, plus an Extended Voice Library of hundreds more you can pull with client.voices.list(). It speaks 130 languages with auto-detected input, and you can steer style, accent, pace and tone with natural-language prompts, or drop inline tags like <laugh> and <sigh> for point-in-time vocal events. It also ships voice design (build a persona from a text description) and voice cloning from a reference clip.

The one spec that bites in production: multi-speaker is capped at two speakers per request. A three-person podcast or a group scene means synthesizing each turn separately and stitching the audio yourself. Output defaults to 24 kHz mono WAV, with mulaw and alaw available for telephony. This sits alongside the text-only Gemini 3.8 Flash model, and the two are billed completely differently, which is where most of the confusion starts.

Gemini 3.8 Flash TTS pricing: the full table

Here is every tier, in USD per 1M tokens, at the 2026 promotional rates. The 2027 rate is in parentheses because, as we will get to, everything doubles.

TierText input / 1MAudio output / 1MNotes
Standard$0.50 ($1.00)$9.00 ($18.00)The default rate
Batch$0.25 ($0.50)$4.50 ($9.00)~50% off, async jobs
Flex$0.25 ($0.50)$4.50 ($9.00)Same rate, cheaper caching
Priority$0.90 ($1.80)$16.20 ($32.40)~1.8x, latency-sensitive
FreeFreeFreeStandard/priority via AI Studio

All of these come straight from Google's Gemini API pricing page. Context caching is supported and runs from $0.125 per 1M input-caching tokens on standard, plus a $0.50 per 1M per hour storage fee. The audio output number is the one that matters, because that is what scales with how much speech you generate. Text input is almost a rounding error by comparison.

What is quietly interesting is where $9.00 sits in Google's own lineup. It is cheaper than the Gemini 2.5 Flash TTS it succeeds ($10 audio out), and less than half the price of the 2.5 Pro and 3.1 Flash TTS preview models ($20 audio out). There is also a cheaper sibling, Flash-Lite TTS, at $6 audio out if you can live with 101 languages instead of 130.

Bar chart comparing audio output cost per 1M tokens across Gemini text-to-speech models, with 3.8 Flash TTS at $9 highlighted
Bar chart comparing audio output cost per 1M tokens across Gemini text-to-speech models, with 3.8 Flash TTS at $9 highlighted

What that actually costs you

Per-token pricing is hard to feel, so let's turn it into real money. Audio bills at 25 tokens per second, which means 1,500 tokens a minute and 90,000 tokens an hour. At the 2026 standard rate of $9.00 per 1M, an hour of generated speech is about $0.81. Add the text you fed in (a full hour of narration is only around 12,000 input tokens) and it is well under a cent more.

Pipeline showing 25 audio tokens per second becoming 1,500 per minute, $0.0135 per minute, and about $0.81 per hour at the 2026 standard rate
Pipeline showing 25 audio tokens per second becoming 1,500 per minute, $0.0135 per minute, and about $0.81 per hour at the 2026 standard rate

A few worked examples at the standard rate:

  • A 30-minute podcast episode: about $0.41 in audio.
  • An 8-hour audiobook: roughly $6.48.
  • 10,000 support IVR messages at 15 seconds each: 3.75M audio tokens, so about $33.75.

Run any of those as an overnight batch job and you halve it, which makes the audiobook closer to $3.24. For high-volume, non-interactive generation, batch is the obvious play. The one place the sticker doesn't tell the whole story is anything a human waits on in real time, where you are paying the priority rate and eating latency you cannot fully control.

The catch: prices double on 1 January 2027

This is the part to write on a sticky note. The $0.50 and $9.00 figures are introductory promotional rates. On 1 January 2027, standard text input goes to $1.00 per 1M and audio output to $18.00 per 1M. Every other tier doubles in lockstep: batch and flex audio output to $9.00, priority audio output to $32.40.

Before and after comparison showing Gemini 3.8 Flash TTS standard rates doubling from $9 to $18 audio output and $0.50 to $1.00 text input on 1 January 2027
Before and after comparison showing Gemini 3.8 Flash TTS standard rates doubling from $9 to $18 audio output and $0.50 to $1.00 text input on 1 January 2027

So that $0.81 hour of audio becomes about $1.62 in 2027, and the 8-hour audiobook goes from $6.48 to $12.96. It is the same move Google pulled on the text models, and it is a fair one as long as you see it coming. If you are budgeting anything past this year, double the numbers in this post and plan from there. It also nudges you toward locking in whatever caching and batch savings you can now.

How it compares to ElevenLabs, OpenAI, and the rest

Comparing voice models is genuinely tricky, because the vendors don't bill the same way. Amazon, Azure and Deepgram charge per character of input text. OpenAI and Google charge per audio output token, which scales with the length of the generated speech, not the script. ElevenLabs sells subscription credits. To put them on one axis, I converted everything to an estimated cost per hour of speech, using Amazon's own anchor that 1M characters is roughly 23 hours of audio. Treat the per-hour figures as estimates, not contract terms, because they move with speaking pace.

VendorFlagship modelNative price~ $/hour of audioFree tier
Gemini 3.8 Flash TTSgemini-3.8-flash-tts$9 / 1M audio tokens~$0.81Free in AI Studio
Amazon Polly (Neural)Neural$16 / 1M chars~$0.701M chars/mo (12 mo)
Azure AI SpeechNeural / HD Flash$15 / 1M chars~$0.650.5M chars/mo
OpenAIgpt-4o-mini-tts$12 / 1M audio tokens~$0.90None
Amazon Polly (Generative)Generative$30 / 1M chars~$1.30100k chars/mo
DeepgramAura-2$30 / 1M chars~$1.30$200 credit
ElevenLabsEleven v3 / Flash~5c/min (Business)~$3.0010k credits/mo

The takeaway that surprised me: Gemini 3.8 Flash TTS is priced like a commodity neural voice while being one of the expressive generative models. Amazon and Azure's cheaper rows are their older neural voices; their generative-quality tier (Polly Generative) is $1.30 an hour, and Deepgram's flagship lands in the same place. ElevenLabs, still the reference point for voice quality, is roughly four times the price per hour. If cost per hour is your deciding factor and you want expressive output, Gemini is hard to beat right now, at least until the 2027 reset pulls it up to about $1.62.

For the OpenAI side of this, my write-ups on the OpenAI realtime API and gpt-realtime-mini pricing go deeper on where token-billed voice makes sense.

Is the voice any good?

Price only matters if the output is usable, and this is where early reactions get more mixed. The rollout drew some grumbling on quality. One developer on Hacker News put it bluntly:

Hacker News

"I don't know, I guess their roll out is a bit sloppy. It's a bit of a shame, though, since the voices which are available all sound like generic Gemini voices to me. Nothing stands out."

That is one person's ear, and voice preference is subjective, but it lines up with a pattern: Google's TTS is cheap and broad, while ElevenLabs still wins the "does this specific voice sound remarkable" test for a lot of people. Some users are sticking with what they know for now:

Hacker News

"For on the go, I've been using ElevenReader. Technically not free but their free tier has been plenty for me."

So the honest read is this. If you need cheap, multilingual, good-enough narration at scale, Gemini 3.8 Flash TTS is excellent value. If a single, distinctive, brand-defining voice is the product, audition it carefully against ElevenLabs before you commit, and don't assume the price gap means a quality gap in your favor.

Where a cheap voice model fits, and where it doesn't

Here is the reframe I'd want a buyer to leave with. A TTS model is infrastructure. It is a fantastic, cheap way to give an app a voice: an audiobook pipeline, an in-product narrator, an accessibility layer, IVR prompts. For those, $0.81 an hour is a genuinely great deal.

But the moment the use case is customer support, the model is the easy 10% of the problem. The hard part is everything around it: pulling the right answer from your knowledge base, calling the tools that actually resolve a ticket, and getting tested against your real history before it ever speaks to a customer. That is the difference between infrastructure and an employee. Gemini TTS is the former. It will not know your refund policy or triage a ticket, and stitching all of that together yourself is a project, not an API call.

This is the frame we work in every day. We have spent years putting AI on live support queues, and the thing that separates a demo from a rollout is never the model, it is the plumbing and the testing. That is why an AI helpdesk agent is a very different purchase from a voice endpoint, even when both say "AI" on the tin. If you are weighing the human-versus-automation math generally, our AI agent vs human agent cost breakdown is a better starting point than a per-token rate.

Try eesel

eesel is an AI teammate platform, and the current roster includes a ready-to-work AI helpdesk agent that joins your existing support queue. Where Gemini TTS is a model you have to build around, eesel arrives already knowing how to connect to your helpdesk, learn from your past tickets, and get simulated against real historical conversations so you know its resolution rate before it goes live.

eesel AI onboarding screen showing an AI teammate connected to a helpdesk and ready to draft replies
eesel AI onboarding screen showing an AI teammate connected to a helpdesk and ready to draft replies

If you like working in a terminal, the eesel CLI drives the same teammate and workspace programmatically. You can operate it by hand, wire it into scripts, or let a coding agent like Claude Code or Cursor run it headlessly, so the AI you buy is available both in the dashboard and through the API, the same way you would reach for a model like Gemini TTS. You can try eesel for free and simulate it against your own tickets first.

Frequently Asked Questions

How much does Gemini 3.8 Flash TTS cost?
Gemini 3.8 Flash TTS pricing is $0.50 per 1M text input tokens and $9.00 per 1M audio output tokens on the standard paid tier, through 31 December 2026. Because audio bills at 25 tokens per second, that works out to about $0.81 for a full hour of generated speech. From 1 January 2027 both rates double. For the wider range, see my Google Gemini 3 pricing guide.
Is there a free tier for Gemini 3.8 Flash TTS?
Yes. Input and output are free of charge on the standard and priority tiers through Google AI Studio, though the batch and flex tiers are paid-only. The free tier is fine for testing voices, but Google uses free-tier content to improve its products, so read the terms before putting anything sensitive through it. Consumer voice features in the Gemini app need a Google AI subscription.
How does Gemini 3.8 Flash TTS pricing compare to ElevenLabs?
Normalized to audio duration, Gemini 3.8 Flash TTS is around $0.81 per hour of speech, while ElevenLabs advertises low-latency TTS from about 5 cents a minute (roughly $3 an hour) on its Business plan. Gemini is a lot cheaper per hour, but ElevenLabs still leads on voice range and cloning maturity. The catch with any raw model is that price per hour is not price per resolved customer service conversation.
Why does the Gemini 3.8 Flash TTS price double in January 2027?
The $0.50 and $9.00 rates are promotional, the same pattern Google used on earlier Gemini models. On 1 January 2027, standard input goes to $1.00 per 1M and audio output to $18.00 per 1M, and the batch, flex and priority multipliers double alongside them. If you are sizing a 2027 budget, double today's numbers. For the text model's version of this same cliff, see my Gemini 3.8 Flash pricing breakdown.
Is Gemini 3.8 Flash TTS good enough for customer support voice?
It is cheap and expressive, but a voice model only handles the speaking. Support voice also needs retrieval, tools, and testing before it touches a live caller, which is where most projects actually break. That is the job of an AI helpdesk agent rather than a TTS endpoint. If you want the voice layer specifically, pair it with something like Zendesk voice AI or a purpose-built stack.

Share this article

Rama Adi Nugraha

Article by

Rama Adi Nugraha

Rama is a software engineer at eesel AI with two years of experience writing about B2B SaaS, AI tools, and customer support technology. Based in Bali, Indonesia, he brings a developer's perspective to product comparisons — cutting through marketing copy to what the integrations and APIs actually do.

Related Posts

All posts →
Illustration of sound waves and price tags representing Gemini 3.8 Flash TTS pricing
Trending

Gemini 3.8 Flash TTS pricing: what an hour of AI voice actually costs

Gemini 3.8 Flash TTS is $9 per 1M audio tokens, roughly $0.81 an hour of speech. Here is every tier, the 2027 catch, and how it stacks up against ElevenLabs and OpenAI.

Rama Adi NugrahaRama Adi NugrahaSep 24, 2026
Illustration of sound waves, a voice-design control panel and a play button representing Gemini 3.8 Flash TTS
Trending

Gemini 3.8 Flash TTS review: is Google's new voice model worth it?

A hands-on Gemini 3.8 Flash TTS review: how it actually sounds, what you can control, where it beats ElevenLabs on price, and the one benchmark Google did not lead.

Alicia Kirana UtomoAlicia Kirana UtomoSep 24, 2026
Illustration of a team reviewing Gemini 3.8 Flash, with a speed gauge, a rocket, and a verdict checkmark
Trending

Gemini 3.8 Flash review: fast, verbose, and not the upgrade the number implies

A hands-on Gemini 3.8 Flash review: what it's good at, where it falls down, the 13-second catch nobody quoted, and whether to switch from 3.7 Flash.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieSep 8, 2026
Illustration of two people reviewing charts and speed dials around a Gemini spark, representing Gemini 3.8 Flash pricing
Trending

Gemini 3.8 Flash pricing: every rate, the hidden cost, and the catch

Gemini 3.8 Flash costs $0.75/$3.75 per 1M tokens, exactly what 3.7 Flash costs. But the sticker price hides a verbosity tax, and both numbers double on 1 January 2027.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieSep 8, 2026
Illustration of a fast-moving robot coding on a laptop while a person watches, representing Gemini 3.8 Flash
Trending

Gemini 3.8 Flash: what it is, honest benchmarks, and my review

Google shipped Gemini 3.8 Flash on September 2, 2026, three weeks after 3.7. Same price, better scores, and one line of fine print that changes the answer.

Alicia Kirana UtomoAlicia Kirana UtomoSep 3, 2026
Illustration of a developer and a colleague working with a fast AI coding agent
Trending

Gemini 3.7 Flash review: a great model that stopped being cheap

I put Google's Gemini 3.7 Flash against its own benchmarks and its own price list. It is fast and sharp, but it is no longer the cheap high-volume workhorse.

Rama Adi NugrahaRama Adi NugrahaAug 14, 2026
Qwen 3.8 Flash Next alternatives roundup banner
Trending

The 8 best Qwen 3.8 Flash Next alternatives in 2026

The best Qwen 3.8 Flash Next alternatives in 2026, from GLM 5.3 Flash to DeepSeek V4 Flash and Gemini 3.7 Flash, with real pricing and who each one is for.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieAug 30, 2026
Qwen 3.8 Flash Next review banner
Trending

Qwen 3.8 Flash Next review: fast, cheap, and half-baked on purpose

A hands-on review of Qwen 3.8 Flash Next: what it is actually good at, where the under-trained preview shows, and whether it belongs in your stack.

Rama Adi NugrahaRama Adi NugrahaAug 30, 2026
Qwen 3.8 Flash Next launch banner
Trending

Qwen 3.8 Flash Next: Alibaba's open-weight Qwen4 preview, explained

Qwen 3.8 Flash Next is Alibaba's open-weight preview of the Qwen4 architecture. Here is what it is, what it costs, and whether it belongs in your stack.

Alicia Kirana UtomoAlicia Kirana UtomoAug 30, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free