Bottom line first
- Do not hunt for a single “Gemini API price.” Pick a billing track: free-tier quota, paid Standard, Batch/Flex, or Vertex / Enterprise.
- Token cost is dominated by output + thinking tokens, not the sticker on input.
- Official tables still list Flash / Flash-Lite (and 2.5 Pro) as free of charge at low RPM/RPD; Gemini 3.1 Pro Preview has no free tier.
- Through 2026-12-31, 3.7 Flash and 3.6 Flash use intro rates (~$0.75 / $3.75 per 1M). On 2027-01-01 they double. Do not budget the intro price as forever.
- Retries after a closed laptop lid repurchase the same thinking tokens. The watershed is track × output mix × execution host, not the model nickname.
Model quality is not the watershed. Billing track and thinking-token share are.
0. Bottom line
As of 2026-08-19, the Gemini Developer API pricing page is a matrix: model × Standard/Batch/Flex/Priority × tokens per million USD. Google AI Studio chat is still free in many regions. Production Gemini API pricing depends on billing enablement, whether the model still has a free tier, input/output mix, thinking budgets, Batch usage, and context caching.
Working numbers (USD / 1M tokens, paid Standard, excl. tax; verify live):
- Cheapest text: 2.5 Flash-Lite ~$0.10 in / $0.40 out, free tier still listed.
- Balanced: 2.5 Flash ~$0.30 / $2.50; 3.5 Flash-Lite ~$0.30 / $2.50.
- Newer Flash: 3 Flash Preview ~$0.50 / $3.00; 3.5 Flash ~$1.50 / $9.00.
- Intro window to 2026-12-31: 3.7 / 3.6 Flash ~$0.75 / $3.75, then ~$1.50 / $7.50.
- Flagship: 3.1 Pro Preview no free tier; ≤200k ~$2 / $12, >200k ~$4 / $18. 2.5 Pro ≤200k ~$1.25 / $10.
Output price includes thinking tokens. Agent loops are not chat completions. How Gemini fits Function Calling and MCP: 2026 AI Agent stack.
1. Why “what does Gemini API cost?” has no one number
Teams treat Studio chat as infinite free API, quote Pro as if it were Flash, and price only input. PoC week is $0; week one in production fills output tokens, then a closed laptop retries the same thinking budget.
The 2026 table adds Standard/Batch/Flex/Priority, intro expiry dates, and 200k prompt tiers. Posts frozen in May 2026 miss 3.7 Flash intro pricing and “thinking counts as output.”
- Free is not unlimited: RPM/RPD, and unused-to-improve is Yes on free.
- Batch/Flex ≈ half price for delay; Priority costs more for peak access.
- Google Search grounding: ~5,000 free/month shared on Gemini 3.x, then ~$14 / 1k; 2.5 often 1,500 RPD then $35 / 1k.
- Audio/image/video, TTS, Imagen, Veo, Lyria are separate rows.
Claude Code is often a seat subscription; Gemini Developer API is usually token cost. See Claude Code 2026 pricing.
2. Five billing tracks
2.1 Track A: AI Studio UI
Try prompts. Does not give production RPM, SLA, or “not used to train” guarantees. Do not put a screenshot in the finance model.
2.2 Track B: Gemini API free tier
API key from AI Studio. Flash lines list Free of charge; 3.1 Pro Preview is Not available. Limits: rate limits. Fine for schema tests, not 24/7 agents.
2.3 Track C: Paid Developer API
Higher limits, caching, Batch/Flex/Priority; paid tier usually not used to improve products. Default production track. See Billing.
2.4 Track D: Vertex / Enterprise
IAM, VPC-SC, contracts, regions. Unit prices may differ. See Vertex pricing. Choose for compliance, not a 3-cent blog delta.
2.5 Track E: Execution host (hidden)
The API bill does not include Macs or MCP. Failed iOS builds and lid-close timeouts buy thinking tokens twice. Where MCP should live: MCP on Cloud Mac vs VPS vs local.
3. Comparison with one header set
| Layer / option | Entry | Execution | Context | Cost | Permission boundary |
|---|---|---|---|---|---|
| API free tier | API key / Studio | Reasoning + functionCall inside quota | May be used to improve products | Sticker $0; real cap is RPM/RPD | Project quotas; never ship keys to the browser |
| Paid Standard | Billed Developer API | Prod RPM; paid Grounding/cache | Paid row usually No on training | Per 1M tokens; thinking in output | Keys on the backend |
| Batch / Flex | Async / delay-tolerant | Same model, discount for latency | Evals, night jobs | ~50% off; often no free Batch | Still your GCP invoice |
| Priority | Peak access | Same model, higher USD | Online spike agents | e.g. 3.7 Flash intro ~$1.35 / $6.75 | When SLO beats unit price |
| Vertex / enterprise | GCP / procurement | Regions, private network, support | Enterprise boundary | Contract ≠ blog table | IAM / VPC-SC |
| Execution host | Laptop / VPS / Cloud Mac | MCP, tests, signing — not inference | Repos, certs, daemons | Machine lease + wasted retry tokens | Process user and egress |
A cheaper Flash does not fix “every timeout rebills thinking.” Vertex does not fix a forked schema. Price is a tick mark on the track.
3.1 Standard cheat sheet (USD / 1M, 2026-08 official)
| Model | Free in/out | Paid input | Paid output (incl. thinking) | Notes |
|---|---|---|---|---|
| 2.5 Flash-Lite | Yes | $0.10 | $0.40 | Cheapest text at scale |
| 2.5 Flash | Yes | $0.30 | $2.50 | Audio in $1.00 |
| 2.5 Pro | Yes | $1.25 / $2.50 (>200k) | $10 / $15 | Long-context tiers |
| 3 Flash Preview | Yes | $0.50 | $3.00 | Audio in $1.00 |
| 3.1 Flash-Lite | Yes | $0.25 | $1.50 | Audio in $0.50 |
| 3.5 Flash-Lite | Yes | $0.30 | $2.50 | High-volume agents |
| 3.5 Flash | Yes | $1.50 | $9.00 | Cached input $0.15 |
| 3.6 / 3.7 Flash | Yes | $0.75→$1.50 | $3.75→$7.50 | Intro until 2026-12-31 |
| 3.1 Pro Preview | No | $2 / $4 (>200k) | $12 / $18 | Do not hammer with free keys |
4. Scenario matrix
| Layer / option | Entry | Execution | Context | Cost | Permission boundary |
|---|---|---|---|---|---|
| Learn the API / demo | Studio + free key | Flash-Lite / 3 Flash, thinking off | One notebook | Target $0; stop on 429 | Test keys only |
| Weekly product, low DAU | Paid Standard + Flash-Lite | Cap tool steps | Short context + cache system | Budget output 3–8× input | Backend proxy; no browser keys |
| Coding agent / MCP hops | Paid Standard; optional 3.7 intro | Small thinking budget; fuse on fail | Repo summaries, not the whole tree | Output+thinking dominate | Always-on MCP so retries die |
| Night evals / ETL | Batch / Flex | Same model, hour-scale OK | Bulk files | Discount for SLA | Queue and retry policy written down |
| Compliance / region / private net | Vertex / contract | Same family, different invoice | Data in VPC | Procurement, not this blog | Least-privilege IAM |
| iOS / Xcode colocated agent | Cloud Mac Host + MCP | Gemini infers only | Certs and DerivedData on Mac | Day-lease machine + API tokens | Keychain never leaves the box |
5. Recommended stacks
A | Personal PoC Studio → free 2.5 Flash-Lite → thinking off / tiny budget → no production data B | Indie product (default) Paid Developer API + 3.5 Flash-Lite or 2.5 Flash → cache the system prompt → max agent steps + daily USD cap → Batch for overnight evals (~50% off) → 3.1 Pro only as a hard-case router C | 2026 Q4 intro window Coding path on 3.7 Flash Standard ($0.75/$3.75) → stress the 2027-01-01 double → never burn Standard on jobs that can wait for Batch D | Apple delivery Gemini = inference only → MCP / xcodebuild / signing on Cloud Mac → lid-close retries hit token cost directly
6. Pitfalls
- Pitfall 1: Studio free ≠ infinite API. Free is a rate limit; 429 is the product.
- Pitfall 2: Quoting input only. Output includes thinking; agents often emit ≥ input.
- Pitfall 3: Hitting 3.1 Pro Preview with a free key. Official free column is Not available.
- Pitfall 4: Writing 3.7 intro rates into an annual budget with no expiry.
- Pitfall 5: Keys in the browser. Scrapers bill your paid tier.
- Pitfall 6: Unbounded retries. Lid close and MCP timeouts repurchase thinking.
- Pitfall 7: Mixing Grounding, image, and TTS into the text-token cell.
7. Seven steps
- Open the official table; record model id, free-tier yes/no, intro expiry.
- Run the real production prompt in AI Studio; read input/output/thinking. Never estimate from “hello.”
- Cap tool steps and daily USD; on 429 fall back to Flash-Lite, do not hammer Pro.
- Cache repeated prefixes (system prompt, tool JSON Schema); check cached-read price.
- Offline evals on Batch/Flex; live traffic on Standard or Priority.
- Move MCP and builds off lid-closing laptops. See MCP and Cloud Mac.
- Internal rate card: default model, upgrade model, 2027-01-01 stress price; diff the invoice monthly.
8. FAQ
Is there still a Gemini API free tier?
Yes, per model. Many Flash / Flash-Lite rows and 2.5 Pro still list Free of charge on Standard. 3.1 Pro Preview does not. Free means low RPM/RPD and possible use to improve products—not a production quota.
How are tokens billed? Where does thinking go?
Input and output are priced per million tokens. Official copy says output includes thinking tokens. Context cache is a third column: storage per hour, hits far cheaper than fresh input.
Is Batch always cheaper?
Paid Batch/Flex is often ~50% off, but Batch frequently has no free tier. You trade latency. Unsuitable for user-facing chat.
Developer API vs Vertex—which is cheaper?
Do not cross-shop sticker tables. Vertex sells region, IAM, and contracts. Startups move faster on Developer API; use Vertex when the compliance checklist says so—not to save three cents and fork tool schemas.
Why put the agent on a cloud Mac?
Gemini sells inference. If MCP, Xcode, or Keychain die on a closed laptop, paid tier rebills thinking. Always-on Cloud Mac cuts wasted token cost, not the model’s unit price.
9. Summary
2026 Gemini API pricing is an expiring matrix: free tier for trials, paid Standard for live, Batch for offline, Priority for spikes, Vertex for compliance. Monthly spend is output + thinking share plus whether you pay twice for retries.
Order of work: lock model ids on the official table → measure tokens on a real prompt → default Flash-Lite/Flash → route Pro → cache and Batch to cut input → keep the execution host up. Stickers move; tracks and circuit breakers should not be rewritten weekly.
Spend tokens on inference, not on lid-close retries
Token cost is already printed per million. What blows the budget is the agent timing out on a laptop, MCP dropping, and Xcode signing failing—then thinking runs again. Dedicated M4 Cloud Mac keeps Host, repo, and MCP on one always-on path. Low idle power and stable SSH make day-lease acceptance then monthly lock-in a sane way to make every million output tokens hit a real step.