Gemini 3.8 Flash pricing: every rate, the hidden cost, and the catch

Kurnia Kharisma Agung Samiadjie
Written by

Kurnia Kharisma Agung Samiadjie

Katelin Teen
Reviewed by

Katelin Teen

Last edited September 8, 2026

Expert Verified
Illustration of two people reviewing charts and speed dials around a Gemini spark, representing Gemini 3.8 Flash pricing

Gemini 3.8 Flash pricing at a glance

Here is every published rate for gemini-3.8-flash on the Gemini API, standard paid tier, per 1M tokens. I have put the 2027 column right next to today's so the promo is impossible to miss.

MeterThrough 31 Dec 2026From 1 Jan 2027
Input$0.75$1.50
Output (includes thinking)$3.75$7.50
Context cache, read$0.075$0.15
Context cache, storage (per 1M/hr)$0.50$1.00
Batch and Flex, input$0.375$0.75
Batch and Flex, output$1.875$3.75
Priority, input$1.35$2.70
Priority, output$6.75$13.50

A few things that are easy to miss reading Google's page top to bottom:

  • Output pricing includes thinking tokens. You pay the $3.75 rate for the model's reasoning, even though the API only returns a summary of it. On a model tuned to reason more, that is the line that moves your bill.
  • Batch and Flex are exactly half price. If your work is not interactive, an overnight batch job halves every token you spend.
  • Grounding is metered separately. Google Search as a tool gives you 5,000 free requests a month, shared across every Gemini 3.x model you run, then $14 per 1,000. Google Maps grounding uses the same allowance.
  • There is a real free tier. Input, output and caching are free of charge for standard use, which is enough to prototype before you spend anything.

Try the numbers on your own workload

The headline rate only tells you so much. What matters is your monthly volume, your input-to-output ratio, and which tier you run on. Plug your own numbers in below to see the monthly cost today versus what the same workload costs after the January 2027 increase.

The hidden cost: 3.8 Flash is verbose on purpose

The number that changes the buying decision is not on the pricing page. It is in how many tokens the model spends to do a task.

Google states the tradeoff openly in the launch post: 3.8 Flash "works harder," executing extra reasoning steps and calling tools iteratively, and "at times, the model might use more tokens to maximize performance." That is a feature for quality, and a cost for your invoice, because thinking tokens bill at the $3.75 output rate.

The independent measurement backs it up. Artificial Analysis needed 120M output tokens to run its whole Intelligence Index against 3.8 Flash, against a 71M median for the class. AA's own one-word verdict was "very verbose."

Bar chart comparing Gemini 3.8 Flash at 120M output tokens against a 71M median, with the note that thinking tokens bill as output
Bar chart comparing Gemini 3.8 Flash at 120M output tokens against a 71M median, with the note that thinking tokens bill as output

On Hacker News, one commenter put a sharper number on it than the marketing did:

Hacker News

"On artificial analysis it's only equal to opus 5 medium effort. Opus 5 max scores 63. Further, opus 5 medium outputs 4x fewer tokens to achieve the same result, negating a lot of the speed difference."

So the real cost per unit of work is higher than $3.75 per 1M suggests, because 3.8 Flash spends more units per task. Artificial Analysis put the blended cost at $0.58 per task on its index, which lands it 59th of 195 models on cost efficiency, not near the top where the raw token rate would imply. If you want to keep an eye on this on your own traffic, that is exactly what LLM tracking tools are for.

Same price as 3.7 Flash, so which do you buy?

This is the actual decision, and Google answers it for you. Since 3.8 and 3.7 Flash cost the same to the cent, price cannot break the tie. Google's own guidance does: for workloads "where compute efficiency is the primary constraint," developers should "continue to rely on Gemini 3.7 Flash, which remains fully supported for efficiency-first workloads."

Put plainly:

  • Long agentic coding, quality over cost: 3.8 Flash. It beats 3.7 on every benchmark Google published, and switching is free. If you are shopping the wider category, my AI coding assistant roundup has the field.
  • High-volume, latency-sensitive, cost-first: stay on 3.7 Flash. It burns fewer tokens for the same rate, and a Googler on the launch thread confirmed 3.8's extra token usage is the intended behaviour, not a bug.

There is a latency wrinkle too. 3.8 Flash ranks 3rd of 195 for raw output speed, but its time to first token is 13.30 seconds against a 2.99 second median. For anything a human waits on, the "fast" model is the slow one to start responding. That is throughput, not responsiveness, and it is easy to misread from the speed headline.

How the rate compares to other frontier models

If you are cross-shopping, here is where Gemini 3.8 Flash sits against the models it actually competes with on price, per 1M tokens, standard tier, today.

ModelInputOutputNote
Gemini 3.8 Flash$0.75$3.75Doubles Jan 2027
Gemini 3.7 Flash$0.75$3.75Identical, fewer tokens per task
Claude Opus 5~$5.00~$25.004x fewer output tokens per task
Claude Sonnet 5lower than Opuslower than OpusCheaper Claude tier

The honest framing from the same HN thread: one developer noted that Gemini Flash, alongside a couple of other cheap models, delivers "90% of the performance for a small fraction of the cost" of the frontier tier. That is the real pitch, and it is a fair one. Just remember the fraction is measured on the raw rate, before you add the verbosity tax and the January increase. My roundup of Gemini alternatives covers the rest of the field, and if you are weighing Google against OpenAI, the ChatGPT vs Gemini comparison is the place to start, with the full OpenAI models list for the other side of that.

Where you can actually run it

  • Free tier: input, output and caching free through Google AI Studio. Enough to prototype.
  • Default in Antigravity: 3.8 Flash is the default model in Google Antigravity, with generous quotas, which several developers on HN flagged as the cheapest way to put it through real work.
  • Consumer apps: the Gemini app, AI Mode in Search and Gemini in Sheets, but only on Google AI Pro or Ultra. If you are deciding between those, see Google AI plans.
  • How to turn it on or off: if you are managing Gemini access across a Workspace, here is how to enable and disable Gemini.

What this means if you are pricing a support bot

I spend most of my time thinking about how buyers search for and price AI for the helpdesk, so here is the part that matters if the reason you are reading a model pricing page is to automate customer support.

A per-token rate of $0.75 looks irresistible next to a per-resolution helpdesk add-on. But three things break that comparison. A support model is answering short, high-volume tickets where the 13.30 second time to first token is felt on every reply. The verbosity that makes 3.8 Flash good at long coding jobs is the wrong shape for ticket deflection, and it bills at the output rate. And the token rate is a small slice of the real cost of running support automation, which is dominated by the retrieval and testing work around the model, not the model call itself.

I have spent the last three-plus years watching AI agents run on live support queues, and the pattern is consistent: a confident-sounding model can quietly give wrong answers, and no per-token price fixes that. It is why every eesel rollout is simulated against your real past tickets before it ever replies to a customer, so you see the resolution rate and the cost per ticket up front instead of discovering both on the invoice.

If you are picking a model to build support automation yourself, price it honestly with the verbosity tax and the 2027 rate baked in. If you would rather not hand-wire a model to your helpdesk at all, eesel packages an AI helpdesk teammate that plugs into your existing tools, learns from your past tickets, and is free to try, no per-token math required. It is the difference between buying an engine and hiring someone who can already drive.

Frequently Asked Questions

How much does Gemini 3.8 Flash cost?
Gemini 3.8 Flash pricing is $0.75 per 1M input tokens and $3.75 per 1M output tokens, on the standard paid tier, through 31 December 2026. On 1 January 2027 both rates double, to $1.50 and $7.50. Output includes thinking tokens, so a model that reasons more bills you more. There is also a free tier. For the wider Gemini range, see my Google Gemini 3 pricing breakdown.
Is Gemini 3.8 Flash pricing the same as 3.7 Flash?
Yes, line for line. Input, output, caching, batch, flex and priority rates are identical between the two models, and Google's own model card says 3.8 Flash is built on 3.7 Flash rather than a new base model. Since the price is the same, the buying question is whether 3.8's extra token burn is worth it for your workload. For latency-sensitive jobs like ticket classification, 3.7 Flash is often the better buy.
Why does the Gemini 3.8 Flash price double in January 2027?
The $0.75 and $3.75 rates are introductory, the same promotional pricing Google first applied to 3.7 Flash. From 1 January 2027, standard input goes to $1.50 per 1M and output to $7.50 per 1M. If you are sizing a 2027 budget on today's numbers, double them. The batch, flex and priority multipliers stay the same, so those rates double too.
What are the batch and priority prices for Gemini 3.8 Flash?
The batch and flex tiers are exactly 50% off standard, so $0.375 input and $1.875 output per 1M today. The priority tier is 1.8x standard, at $1.35 input and $6.75 output. Context cache reads are $0.075 per 1M plus $0.50 per 1M per hour of storage. Every one of these doubles on 1 January 2027 alongside the standard rate.
Is Gemini 3.8 Flash cheaper than Claude Opus 5?
On the sticker price, yes, by a wide margin. Gemini 3.8 Flash is $0.75/$3.75 versus Claude Opus 5 at roughly $5/$25 per 1M. But Opus 5 medium hits the same 59 intelligence score using about 4x fewer output tokens per task, which closes some of that gap in practice. My Gemini vs Claude Opus comparison goes deeper on the tradeoff.
Is there a free tier for Gemini 3.8 Flash?
Yes. The Gemini API free tier covers input, output and context caching at no charge, accessible through Google AI Studio, and 3.8 Flash is the default model in Google Antigravity. Consumer access in the Gemini app, AI Mode and Google Sheets needs a Google AI Pro or Ultra subscription, so check Google AI plans first.
Is Gemini 3.8 Flash pricing good for customer support?
The per-token rate is cheap, but the total cost of a support model is not just the rate. Support work is short, high-volume and latency-sensitive, and 3.8 Flash is tuned to think longer and burn more output tokens, with a 13.30 second time to first token. That token burn shows up on the bill. The bigger point is that model price is rarely what makes AI for customer service work, retrieval and testing matter more, which is what an AI helpdesk agent handles.

Share this article

Kurnia Kharisma Agung Samiadjie

Article by

Kurnia Kharisma Agung Samiadjie

Kurnia is a software engineer and writer at eesel AI with two years of SEO experience, writing about AI tools, helpdesk software, and customer support. He pairs a developer's understanding of how these products are built with search-driven research into what actually ranks and resonates with the people searching for them.

Related Posts

All posts →
Illustration of a team reviewing Gemini 3.8 Flash, with a speed gauge, a rocket, and a verdict checkmark
Trending

Gemini 3.8 Flash review: fast, verbose, and not the upgrade the number implies

A hands-on Gemini 3.8 Flash review: what it's good at, where it falls down, the 13-second catch nobody quoted, and whether to switch from 3.7 Flash.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieSep 8, 2026
Illustration of a fast-moving robot coding on a laptop while a person watches, representing Gemini 3.8 Flash
Trending

Gemini 3.8 Flash: what it is, honest benchmarks, and my review

Google shipped Gemini 3.8 Flash on September 2, 2026, three weeks after 3.7. Same price, better scores, and one line of fine print that changes the answer.

Alicia Kirana UtomoAlicia Kirana UtomoSep 3, 2026
Illustration of a multimodal AI model turning inputs into tokens that funnel down to a dollar sign, for a GLM-5.3-Flash pricing breakdown
Trending

GLM-5.3-Flash pricing: every rate, the promo cliff, and the real cost

GLM-5.3-Flash pricing in full: the $0.075/$0.25 promo rates, the September cliff, the coding plan, and the throughput gap that changes your real cost.

Rama Adi NugrahaRama Adi NugrahaAug 29, 2026
A runner carrying a lightning bolt sprinting past a piggy bank, illustrating GLM-5.3 Flash speed and low cost
Trending

GLM-5.3 Flash review: frontier scores at flash cost

A hands-on GLM-5.3 Flash review: the benchmarks it actually posts, what its 4.5-cent-a-task price hides, where it breaks, and who should run it.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieAug 29, 2026
A reviewer looking at a verdict scorecard with two effort dials labelled low and max, beside the DeepSeek whale
Trending

DeepSeek V4 Flash review: one model, two personalities

A DeepSeek V4 Flash review built on the numbers both scoreboards publish. The cheap run and the smart run are the same weights, and that changes the verdict.

Riellvriany IndriawanRiellvriany IndriawanAug 4, 2026
DeepSeek V4 Flash pricing: what you'll actually be billed
Trending

DeepSeek V4 Flash pricing: what you'll actually be billed

DeepSeek V4 Flash lists at $0.14 in and $0.28 out per million tokens. Real users have posted blended rates under a cent. Here is what decides which one you get.

Alicia Kirana UtomoAlicia Kirana UtomoAug 4, 2026
DeepSeek V4 Flash: specs, pricing, and what it's really for
Trending

DeepSeek V4 Flash: specs, pricing, and what it's really for

DeepSeek V4 Flash costs $0.14 in and $0.28 out per million tokens, and it outscores DeepSeek's own expensive tier. Here's what the price card doesn't tell you.

Rama Adi NugrahaRama Adi NugrahaAug 4, 2026
Two people arm wrestling across a table while a third watches, illustrating a head-to-head model comparison
Trending

DeepSeek V4 Flash vs GPT-5.6: which one do you build on?

DeepSeek V4 Flash vs GPT-5.6 on August 2026 numbers. The real fight is Flash against Luna, intelligence is a tie, and the deciding factors are speed, vision and data.

Rama Adi NugrahaRama Adi NugrahaAug 4, 2026
Illustration comparing the DeepSeek V4 Flash and V4 Pro model tiers
Trending

DeepSeek V4 Flash vs V4 Pro: which tier should you use?

DeepSeek's cheap tier now scores higher than its expensive one on the independent board. Here is exactly where that holds, and the two places it does not.

Rama Adi NugrahaRama Adi NugrahaAug 3, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free