Preloader
Others
  • Estimated reading time: 13 Minutes

Google Gemini API Pricing 2026: Flash, Pro, and the Costs

Google Gemini API Pricing 2026: Flash, Pro, and the Costs

Prices and quotas below were checked against Google's official pricing page on August 20, 2026. Google moves fast here. Gemini 3.7 Flash landed on August 13 and reset the Flash tier's pricing within days, so confirm current rates before making a budget decision.

Google markets the Gemini API as the budget-friendly door into frontier AI, and for a lot of use cases, it is. But open the actual pricing page and you'll find over a dozen line items: four separate service tiers, cache storage fees, tool charges, introductory rates with expiry dates, and a deprecation schedule that can quietly retire the model you built on. Google Gemini API pricing in 2026 is not one number. It's a set of tradeoffs you make per model, per feature, and per how you're calling it. That matters whether you're comparing Gemini against other APIs for a new project, or you already have it in production and want to know exactly what you're paying for. Here's what each piece actually costs, and where the real spend tends to hide.

Google Gemini API Pricing 2026 at a Glance

Gemini's API pricing splits across three account tiers, and which one you're on decides both your rate and what you can even access.

Tier Who it's for What you get
Free Testing and low-volume prototyping Free input and output tokens on Flash, Flash-Lite, and the current Gemini 2.5 Pro; newer Pro preview models and image/video generation are excluded, and your content is used to improve Google's products
Paid Production apps Higher rate limits, access to the newest preview models, Batch and Flex discounts, and content not used to improve Google's products
Enterprise Custom deployments Delivered through the Gemini Enterprise Agent Platform rather than the standard developer API, with provisioned throughput and volume discounts

Pricing itself is quoted per token, a token being roughly three-quarters of a word, the unit models actually process. Google prices input and output tokens separately for every model, and output is always the more expensive half.

Underneath the account tier sits a second choice most teams miss: how you call the model. Google publishes four service lanes with four different rates for the same model. Standard is the default. Batch and Flex both cut the rate in half in exchange for delayed processing. Priority costs 80% more for guaranteed low latency. Picking a lane is worth as much as picking a model.

Model-by-Model Pricing: Flash, Flash-Lite, and the Pro Line

This is Standard, pay-as-you-go pricing per 1 million tokens on the paid tier. Batch, Flex, and caching rates are covered separately below, since they change the math in ways worth handling on their own.

Model Input ($/1M tokens) Output ($/1M tokens) Free tier
Gemini 2.5 Flash-Lite $0.10 (text/image/video), $0.30 (audio) $0.40 Eligible
Gemini 3.1 Flash-Lite $0.25 (text/image/video), $0.50 (audio) $1.50 Eligible
Gemini 2.5 Flash $0.30 (text/image/video), $1.00 (audio) $2.50 Eligible
Gemini 3.5 Flash-Lite $0.30 (text/image/video/audio) $2.50 Eligible
Gemini 3 Flash Preview $0.50 (text/image/video), $1.00 (audio) $3.00 Eligible
Gemini 3.7 Flash $0.75 through Dec 31, 2026, then $1.50 $3.75 through Dec 31, 2026, then $7.50 Eligible
Gemini 3.6 Flash $0.75 through Dec 31, 2026, then $1.50 $3.75 through Dec 31, 2026, then $7.50 Eligible
Gemini 2.5 Pro $1.25 (prompts up to 200k), $2.50 (longer) $10.00 (prompts up to 200k), $15.00 (longer) Eligible
Gemini 3.5 Flash $1.50 $9.00 Eligible
Gemini 3.1 Pro Preview $2.00 (prompts up to 200k), $4.00 (longer) $12.00 (prompts up to 200k), $18.00 (longer) Paid tier only

Three things are worth reading into that table.

The price gap between tiers is intentional, not incidental. Flash-Lite models are built for high-volume, low-complexity calls, the kind where a slightly worse answer costs you nothing, like classifying a support ticket or pulling a date out of an email. Pro-tier models cost noticeably more per output token because they're built for work where getting the answer right matters more than getting it cheap. Picking the wrong tier for the job, not the base rates themselves, is usually where teams end up overspending.

Newer does not mean pricier. Gemini 3.7 Flash shipped on August 13, three weeks after 3.6 Flash, at half of what 3.6 Flash originally cost. Google then moved 3.6 Flash onto the same introductory rate, so the two currently cost exactly the same and the upgrade decision comes down to behaviour and latency rather than price. Meanwhile Gemini 3.5 Flash, an older model, still bills at $1.50 and $9.00. Being on the newest Flash model is currently the cheap option, which is the reverse of how most API lineups work.

That introductory rate has a date on it. Both 3.7 Flash and 3.6 Flash revert to $1.50 input and $7.50 output on January 1, 2027, and their cache rates double at the same time. If you're modelling annual spend on today's Flash pricing, you're modelling a number that doubles four months in. That single detail is worth more than most of the rate card.

Worth flagging on the other end: Gemini 2.0 Flash and 2.0 Flash-Lite are still listed on the pricing page but were deprecated and shut down on June 1, 2026. Imagen 4 followed on August 17, 2026. If your integration was still calling any of them, that migration is already overdue, check now rather than waiting for a failed request to tell you.

What Gemini's Free Tier Actually Covers

The free tier is real, and it's more generous than the headline "free" suggests. Google's pricing page lists free input and output tokens across Flash, Flash-Lite, and the current Gemini 2.5 Pro. Context caching is free too on the Gemini 3.x Flash models, though not on the Flash-Lite line, 2.5 Flash, or either Pro model, where the free tier column reads "not available" for caching.

What it doesn't give you: the newer Gemini 3.1 Pro Preview, image generation, or video generation. Those need the paid tier regardless of your volume. If your prototype depends on Gemini's newest preview-tier reasoning model, you'll hit that wall the moment you move past a basic demo.

The other free-tier cost isn't billed in dollars. Google states plainly that free-tier content is used to improve its products and paid-tier content is not. For a hobby project that's fine. For anything touching customer data, it's the reason to upgrade regardless of what your token volume looks like.

Tool usage has its own free allowance, and it's uneven across models. Google Search grounding, the feature that lets a model pull in live search results, gives Gemini 2.5 Flash and 2.5 Flash-Lite a shared 500 requests per day on the free tier. Gemini 2.5 Pro gets none: grounding on Pro requires the paid tier from the first call. Move to the paid tier and the numbers shift again. Flash and Flash-Lite jump to a shared 1,500 requests per day, Pro gets its own separate 1,500, and both bill at $35 per 1,000 after that. Gemini 3.x models run on a different allowance entirely: 5,000 free requests per month shared across the whole 3.x lineup, then $14 per 1,000. Google Maps grounding runs its own third set of quotas on top of that, at $25 per 1,000 on the 2.5 line and $14 per 1,000 on 3.x. Worth checking which meter actually applies before you budget for grounding.

What Google Gemini API Pricing 2026 Doesn't Show You

A few charges live outside the main rate table, and they're the ones that catch teams off guard once they're past the prototype stage.

Thinking tokens are the big one. Look closely at Google's pricing page and every output row is labelled "output price, including thinking tokens." Reasoning models spend tokens working through a problem before they answer, you never see those tokens, and they bill at the full output rate. A three-sentence reply to a hard prompt can carry thousands of billed tokens behind it. This is the single widest gap between what the rate card implies and what the invoice says, and it gets wider the harder your prompts are. Models that expose a thinking level setting let you turn it down for simple tasks, which cuts output tokens directly.

Context caching cuts both ways. Caching lets you reuse previously processed input instead of paying to reprocess it on every call, and the saving is substantial: cached input runs about 90% below the standard input rate, so Gemini 3.1 Pro Preview drops from $2.00 to $0.20 per million tokens. The catch is the storage meter. Cached content is billed by the hour whether or not you're actively calling the model, at roughly $1.00 per million tokens per hour on most Flash and Flash-Lite models and $4.50 on Pro-tier models like Gemini 2.5 Pro and the 3.1 Pro Preview. Caching pays off when you repeatedly reuse a large, stable block of context, like a document assistant referencing the same manual all day. Cache something large against a Pro model for a chatbot nobody uses overnight and the storage fee outruns the token savings.

The 200K cliff is per prompt, not per month. On both Gemini 2.5 Pro and Gemini 3.1 Pro Preview, input and output rates step up once a single prompt crosses 200,000 tokens. It isn't an average across your traffic, so a handful of very large requests can bill at the higher rate while your dashboard still shows a comfortable mean. RAG pipelines that pack context aggressively, long conversation histories, and document-analysis jobs cross that line constantly.

Batch and Flex halve the rate, in exchange for waiting. Both process requests within a rolling window rather than instantly. Good fit for nightly summarisation, poor fit for anything a user is waiting on.

Priority costs 80% more. It's there for teams that need guaranteed low latency, and it's easy to opt into for a production app without registering what the multiplier does to a monthly total.

Tool costs behave the same way. Search grounding is free up to a quota, but once an agentic workflow calls it on every turn of a conversation, that quota disappears fast, and the overage rate kicks in per request rather than per token. It's easy to size a budget around a model's base rate and forget that every grounded call, every cached read, and every batch job carries its own separate meter.

Where Gemini Sits on the Broader Pricing Map

Every major provider now offers a lightweight, high-volume model and a slower, more capable one, and Gemini's lineup follows that same pattern. Lightweight models across the market tend to sit in a low, sub-dollar range per million input tokens, with output priced several times higher than input on nearly every provider. Frontier-class reasoning models, regardless of who makes them, cost several times more on output than the lightweight tier, and that gap widens fast at the top: output pricing commonly runs from high-single into double-digit dollars per million tokens for a provider's mainline flagship, and climbs well past that for the most premium reasoning variants.

Where this gets complicated isn't the price gap between providers, it's that the cheapest model for your use case today might not be the cheapest one in three months. Gemini itself is the proof: three Flash releases between May and August 2026, a price cut that halved the Flash tier, and a reversion date already on the calendar. Model lineups shift, older versions get deprecated, and pricing tiers get restructured, which is exactly the kind of thing that makes committing to a single provider's pricing sheet a moving target.

That's also why more teams evaluate pricing at the workload level instead of the provider level. A support-ticket classifier and a coding agent rarely belong on the same model, regardless of which company built it, and the model that's cheapest for one workload is often not the cheapest for the other, even inside a single provider's own lineup.

How Infron Handles Gemini API Pricing

If you're the one deciding whether to build on Gemini, or you're already running it in production and watching the bill split across a dozen line items, the tables above are only half the picture. That same pricing sheet can turn into a single line on a single invoice instead, whether you're still evaluating Gemini or already depend on it.

Unified billing. Gemini's own pricing page separates standard rates, batch and flex discounts, cache storage, and tool charges into different line items you have to reconcile yourself. Through Infron, all of that, plus usage across 400+ AI models, lands on one consolidated invoice, so you're not cross-referencing a dozen SKUs to answer what a single feature actually cost.

Pass-through pricing. The markup risk with any gateway is real, so the underlying model cost passes through at zero markup, you pay the same per-token rate Google charges directly. Infron's own transaction fee is layered on top separately and disclosed upfront, not hidden inside inflated per-token pricing. The exact fee structure is laid out in the API documentation before you commit to anything.

Automatic failover. Gemini retires model variants on a schedule, like the June 2026 shutdown of Gemini 2.0 Flash and the August 2026 shutdown of Imagen 4. If a model you depend on gets deprecated or rate-limited, Infron can fail over to another model without you rewriting your integration first.

If that matters for your stack, Infron is worth a look.

Conclusion

Gemini's pricing rewards teams who read past the headline numbers. The per-token rates are only the start. What actually determines your bill is which service lane you're calling, how many thinking tokens your prompts burn, whether you're caching or batching, whether today's rate is an introductory one, and whether the model you built on is still around next quarter. None of that shows up in a single line on Google's pricing page. Treat the free tier as a prototyping tool rather than a production plan, and treat the standard rate table as a starting point rather than the full cost, and the rest of the numbers stop being surprises. None of this is a reason to avoid Gemini, it's a reason to build in enough flexibility that a pricing change or a deprecation notice doesn't force a rewrite.

FAQ

What is Google Gemini API pricing in 2026?

Gemini API pricing is split across a free tier for prototyping, a paid tier for production use, and an enterprise tier for custom deployments. Paid-tier rates are charged per million input and output tokens and vary by model, from $0.10 per million input tokens on Gemini 2.5 Flash-Lite to $4.00 on the Pro preview line at long context. The enterprise tier sits outside this per-token structure entirely, priced through custom agreements via the Gemini Enterprise Agent Platform.

How much does Gemini 3.7 Flash cost?

Gemini 3.7 Flash launched on August 13, 2026 at an introductory $0.75 per million input tokens and $3.75 per million output tokens, half of what Gemini 3.6 Flash originally cost. That rate runs through December 31, 2026 and doubles to $1.50 and $7.50 on January 1, 2027. Google moved 3.6 Flash onto the same introductory rate, so the two models currently cost the same.

Does the Gemini API have a free tier?

Yes. Flash, Flash-Lite, and the current Gemini 2.5 Pro are all listed with free input and output tokens. The newer Gemini 3.1 Pro Preview, image generation, and video generation all require the paid tier regardless of how little you're using them. Free-tier content is also used to improve Google's products, which paid-tier content is not.

Why is my Gemini bill higher than the rate card suggests?

Usually thinking tokens. Google bills a model's internal reasoning at the full output rate, and those tokens never appear in the response you see. Output pricing on every model is labelled "including thinking tokens" for that reason. Grounded search calls, cache storage, and prompts crossing the 200K context threshold are the other three common gaps between the headline rate and the invoice.

What's the difference between Gemini's standard, batch, and flex pricing?

Standard charges the full rate for an immediate response. Batch and Flex both cut that rate in half but process requests within a rolling window instead of instantly, making them suited to offline jobs rather than live user requests. Priority runs in the other direction, costing 80% above standard for guaranteed low latency.

How does context caching affect Gemini API costs?

Context caching lets you reuse previously processed input at roughly 90% below the standard input rate. But cached content is also billed for storage by the hour, at around $1.00 per million tokens per hour on Flash models and $4.50 on Pro models, so caching something large that sits idle can add cost rather than save it. It pays off most for workloads that repeatedly reuse large, stable context, like a document assistant referencing the same manual, rather than short one-off exchanges.

Weekly trending
Our Sponsors

Our blog is proudly supported by industry-leading sponsors.