Limited offer

What Is Gemini 3.5 Pro? Features, Pricing, and Release Date (2026)

AI EngineeringGemini 3.5 · Model guide
2026-07-1312 min read

Google-confirmed vs reported: 2M context, Deep Think, API pricing, mid-July GA timeline; five-axis comparison with Flash and GPT-5.6 plus a 7-step rollout.

As of July 13, 2026, Google has shipped Gemini 3.5 Flash at I/O but Gemini 3.5 Pro is still not publicly GA. This guide separates what Google confirmed from what outlets report so you can decide whether to run on Flash today or reserve budget for Pro.

Key takeaways

  1. Officially confirmed: Google announced the Gemini 3.5 family at I/O on May 19, 2026; Gemini 3.5 Flash reached GA that day; Pro is in internal use with a planned public launch "next month"—that timeline has slipped.
  2. As of July 13: Pro remains in Vertex AI enterprise preview; the public API has no GA entry for gemini-3.5-pro; multiple outlets point to July 17 as a target date, but Google has not confirmed it.
  3. Not officially confirmed but widely reported: 2M-token context, a Deep Think reasoning layer, API pricing around $15/$60 per million tokens, and an Ultra-tier subscription at roughly $250/month.
  4. Available today: Gemini 3.5 Flash (1M context, $1.50/$9 per million tokens) already covers most agent and coding workloads.
  5. The real decision is not "which model is smarter" but context length × reasoning cost × execution environment—for long-running agents, pair with a Cloud Mac running your toolchain 24/7.
Neural network and data-flow visualization symbolizing Gemini 3.5 Pro long context and reasoning
The open question for Gemini 3.5 Pro is not whether it exists, but whether GA-day specs match the preview.

Conclusion first: Flash works today; Pro is worth waiting for, but do not bet production on it

Until Google publishes an official model card, every number you hear about Pro is a bet, not a contract term.

If you remember only three things: ① Gemini 3.5 Flash is already Google's default workhorse for developers in 2026; ② Gemini 3.5 Pro is the family's heavy artillery, but as of mid-July it still has not reached public GA; ③ tying production systems to a rumored "July 17 launch" carries more risk than reward.

Google's official blog post is deliberately measured: the 3.5 family emphasizes "frontier intelligence with action"—aimed at agents and coding; Flash went fully available that day; Pro is "already in internal use, with a public launch expected next month." Sundar Pichai used the same framing on the I/O stage. But the official post does not mention 2M context, Deep Think, or API unit pricing—those come from later press coverage and enterprise-preview leaks, and should be treated with different confidence levels.

For engineers, the pragmatic move is: run your Antigravity / Gemini API / Vertex agent pipelines on Flash, keep tasks that need "the whole repo in one 2M-token shot" or "multi-step hard reasoning" on your backlog, and switch your default model only after a GA entry appears in the Gemini API changelog. That matches what we concluded in our GPT-5.6 and Gemini model-selection guide: model launches are coming faster than ever; fix your workflow entry points and execution nodes first.

1. Why is Gemini 3.5 Pro harder to ship than Flash?

Flash reaching GA on I/O day was not accidental. Google positioned 3.5 Flash as the high-throughput, agent-first default: 1M context, dynamic thinking on by default, and benchmarks such as Terminal-Bench 2.1 at 76.2% were published in the I/O announcement roundup. The goal was to prove that "Flash can compete with flagship models"—critical for Search, the Gemini app, and Antigravity's default experience.

Pro plays a different role. Multiple tech outlets (including Business Insider and Geeky Gadgets) report that Google found token efficiency and quality on long-horizon agent tasks in enterprise preview below its internal bar, triggering an architecture-level rebuild that pushed GA from June into July, with July 17 cited as a target date. Google declined to comment on specific dates—meaning "July 17" is an industry-consensus target window, not a promise you can write into an SLA.

Where did the old plan break? Many teams rewrote their roadmaps right after I/O as "full Pro cutover in June":

  • Budgets were locked to "rumored $15/$60" pricing, but only Flash was billable in June, distorting financial models.
  • Prompts and toolchains built in Vertex preview were not abstracted behind model IDs, forcing rework at GA switchover.
  • "2M tokens, whole repo, one shot" became an architectural assumption, ignoring that Flash at 1M plus retrieval/chunking handles 80% of tasks.
  • Long agents ran on local laptops that disconnect when the lid closes—a tooling problem misattributed to Pro not shipping yet.

The core lesson from Pro's delay: Google would rather slip a month than ship an underperforming "Ultra successor" and damage the brand. For users, that is a signal—Flash is already strong enough that Pro must be clearly better to justify its own price tier.

2. What is Gemini 3.5 Pro? How it splits from Flash

2.1 Family positioning (official framing)

Per Google's description, Gemini 3.5 is the 2026 next-generation model family built around agents, coding, and long-horizon tasks. Flash handles "default, fast, broad coverage"; Pro handles "hardest reasoning, longest context, highest quality"—carrying the narrative once held by Gemini Ultra / 3.1 Pro, though Google's I/O copy does not use "Ultra renamed" language and refers only to "3.5 Pro."

2.2 Gemini 3.5 Flash (GA, the control group)

You cannot evaluate Pro without Flash as a baseline:

  • API ID: gemini-3.5-flash (GA, no preview suffix)
  • Context: 1,048,576 input tokens / 65,536 output tokens (official docs)
  • Pricing: $1.50 / $9.00 per million input/output tokens; cached input $0.15
  • Distribution: Gemini app, AI Mode in Search, Gemini API, AI Studio, Android Studio, Google Antigravity, Vertex AI, Gemini Enterprise
  • Capabilities: multimodal input, function calling, code execution, search tools, dynamic thinking

2.3 Gemini 3.5 Pro (preview / pending GA)

As of 2026-07-13, the publicly confirmable status is:

  • Google acknowledges Pro exists and is used internally; the public rollout planned for "next month (June)" has slipped.
  • Vertex AI enterprise customers can trial it in limited preview (not all developers can self-serve access).
  • The public Gemini API model list centers on gemini-3.5-flash and gemini-3.1-pro-preview, with no GA model card for Pro yet.

Reported Pro differentiators (to be validated at GA):

  • 2M-token context—roughly double Flash, aimed at whole-repo code, long compliance documents, and extended agent trajectories.
  • Deep Think reasoning layer—multi-step logic, complex math, and planning; reportedly tied to subscription tiers (see pricing section).
  • Stronger multimodal and UI generation—continuing Gemini 3's interactive web UI capabilities, positioned for hardest reasoning.
  • Lower latency is not the pitch—Pro prioritizes quality, complementing Flash rather than replacing it.

3. Release timeline: from I/O to mid-July

Date Event Confidence
2026-05-19 Google I/O: Gemini 3.5 announced; Flash GA; Pro in internal use, public launch planned for next month Official
2026-05 ~ 06 Flash becomes default in Gemini app / Search AI Mode; Antigravity and Enterprise follow Official
Late June 2026 Press reports Pro GA slipping from June to July; Vertex preview continues; API changelog shows no Pro GA Press + API status
2026-07-08 ~ 17 Multiple outlets cite Google target of July 17 GA; architecture rebuild and preview feedback cited Press (unconfirmed)
2026-07-13 (article date) Pro still not publicly GA; developers should run production on Flash and watch the changelog Current state

How to watch for the official announcement? Subscribe to two sources: the Google DeepMind / Gemini models blog and the Gemini API Changelog. GA day usually brings all three at once: model card, pricing page update, and AI Studio default option. A headline alone without those three still means preview.

4. Feature checklist: officially confirmed vs. pending announcement

Capability Flash (GA) Pro (expected / reported) Status
Long-horizon agent tasks Officially promoted Stronger reasoning and planning Flash confirmed; Pro pending GA evaluation
Context window 1M tokens 2M tokens (reported) Flash official; Pro unannounced
Deep Think / deep reasoning Dynamic thinking on by default Separate Deep Think layer (reported) Pro unannounced
Coding / Terminal-Bench 76.2% (I/O figure) Expected above Flash Flash has numbers; Pro TBD
Multimodal Image, text, audio, video input Stronger interactive UI (reported) Flash confirmed
Antigravity / Vertex Integrated In preview Flash GA; Pro enterprise preview

For developers, whether a feature is "worth it" depends on task shape:

  • Single-pass code review + patch generation under 800K tokens: Flash is usually enough, with controllable cost.
  • Whole-repo dependency graph + cross-service tracing in one output: 2M is theoretically advantageous, but wait for Pro's measured latency and hallucination rate.
  • Multi-step math proofs, complex compliance reasoning: if Deep Think is a separate tier, it may land near OpenAI reasoning.effort=max or Claude extended thinking.

5. Gemini 3.5 Pro pricing: official blank space and reasonable estimates

Google has not published official API pricing for Pro. The breakdown below avoids treating rumors as invoices.

5.1 Confirmed reference pricing: Gemini 3.5 Flash

Item Price (USD)
Input tokens $1.50 / million (~$1.65 in non-global regions)
Output tokens $9.00 / million (~$9.90 in non-global regions)
Cached input $0.15 / million

5.2 Reported Pro price points (unconfirmed)

Multiple outlets converge on the same order of magnitude—use only for budget sandboxes:

  • API: roughly $15 / million input and $60 / million output—about 10× Flash.
  • Consumer subscription: Deep Think or full Pro capabilities may require Gemini Advanced / Ultra tier (~$250/month) (press framing, not a Google pricing page).
  • Some outlets suggest milder API guesses (e.g., $1.25 / $10), conflicting with the 10× hypothesis—again, GA-day model card wins.

5.3 Rough competitor math (decision aid, not live quotes)

Place rumored Pro pricing alongside Flash and GPT-5.6 / Codex quota logic in the same framework:

  • High-frequency agents, millions of tokens per day: Flash or GPT-5.6 Luna/Terra is more realistic.
  • Low-frequency, ultra-long context, one-shot success matters: Pro at 2M may have TCO advantages (fewer calls for higher quality)—pending official unit pricing.
  • Terminal coding agents running 24/7: beyond model fees, execution environment (always-on Cloud Mac) often matters more than a 5× token price delta.

6. Five-dimension comparison: Pro (expected) vs. Flash vs. GPT-5.6 vs. Claude

Dimension Gemini 3.5 Flash Gemini 3.5 Pro (expected) GPT-5.6 Sol Claude Mythos 5
Availability GA, all channels Preview / pending GA GA (2026-07) GA
Context 1M (official) 2M (reported) Varies by plan / API ~200K+
Coding agent Terminal-Bench 76.2% TBD Terminal-Bench 88.8% SWE-bench Pro leader
API price (order of magnitude) $1.5 / $9 Rumored ~$15 / $60 Sol / Terra / Luna tiers Mid-to-high tier
Best entry point Antigravity / Gemini API Vertex + AI Studio (pending) Codex / Cursor Claude Code CLI

This table is not about "who wins" but that there is no single default winner. Flash is already a "fast enough, cheap enough" agent base inside Google's ecosystem; if Pro delivers 2M + Deep Think, it fills the gap for ultra-long documents and hard reasoning; OpenAI and Anthropic still lead on terminal coding and IDE entry points. See our GPT-5.6 coding model-selection guide for more.

7. Scenario decision matrix

Scenario Recommend now Re-evaluate after Pro GA Avoid
Antigravity default agent 3.5 Flash Switch to Pro per task Blocking work waiting for Pro
Single-pass code review <1M tokens 3.5 Flash Forcing Pro preview
Whole monorepo analysis in one pass Flash + chunking / retrieval 3.5 Pro 2M Hard-running on a local laptop
Complex finance / compliance multi-step reasoning Flash + human review Deep Think (if announced) Long agents on free tier
iOS / Xcode CI agent Flash API + Cloud Mac Pro long context Web UI only, no toolchain
Production SLA systems Flash GA + pinned version Pro GA + gradual rollout Binding to rumored launch date

8. Recommended stacks

[Stack A — individual developer / exploration]
Model: Gemini 3.5 Flash (AI Studio or API)
Entry: Antigravity or custom scripts + function calling
Execution: local for short tasks; move long tasks to Cloud Mac
Budget: estimate tokens at Flash $1.5/$9; do not reserve for Pro 10×

[Stack B — enterprise Vertex]
Model: Flash in production + Pro in separate preview project
Governance: different service accounts for preview vs. production; model ID via env vars
Monitoring: changelog RSS + billing alerts (BigQuery billing export)
Cutover: 10% traffic canary after Pro GA → benchmark regression → full rollout

[Stack C — multi-model coding team]
Retrieval / large-repo scan: Gemini 3.5 Flash
Hard reasoning / architecture review: wait for Pro or Claude extended thinking
Terminal edits: Claude Code or Codex on Cloud Mac worktree
Daily IDE: Cursor + GPT-5.6 Terra
Principle: one primary model per task—avoid double token spend

9. Common pitfalls

  • Pitfall 1: Treating press coverage as a model card. 2M context, Deep Think, and $15/$60 never appeared in Google's official I/O post; label them "unconfirmed" before writing architecture docs.
  • Pitfall 2: Flash is just a stopgap. I/O data shows Flash beating the previous 3.1 Pro on multiple agent benchmarks—it is the primary model, not a "/lite" SKU.
  • Pitfall 3: Everyone must switch the day Pro ships. Longer context ≠ lower total cost; a failed 2M single-shot call with retries can cost more.
  • Pitfall 4: Ignoring differences between Antigravity and the API. The same model has different tool lists, permissions, and latency per harness—test on your real entry point.
  • Pitfall 5: Deploying nothing while waiting for Pro. Competitors will not wait; running your agent pipeline on Flash now beats idle weeks.

10. Rollout steps (7 steps)

  1. Inventory tasks: list the share that need >1M context or hard reasoning; if under 20%, Flash can be your primary model.
  2. Stand up Flash production path: create a project in AI Studio or Vertex; manage API keys / workload identity per environment.
  3. Abstract model IDs: use GEMINI_MODEL=gemini-3.5-flash in code—no scattered hardcoded strings—so Pro GA is a config change.
  4. Move long agents to execution nodes: multi-hour toolchain jobs belong on a Cloud Mac to avoid sleep disconnects (orthogonal to Flash vs. Pro, but critical).
  5. Subscribe to the changelog: Gemini API + Google AI blog; set a calendar reminder to check mid-to-late July for GA.
  6. Pro GA-day workflow: read model card → run internal golden set → compare Flash cost and quality → 10% canary.
  7. Two-week retrospective: log token volume, task success rate, and human interventions—let data decide whether Pro is worth ~10× unit pricing.

11. FAQ

Has Gemini 3.5 Pro launched?

As of July 13, 2026, it has not reached public GA. Google announced at I/O on May 19 that Pro was in internal use with a planned launch the following month; that timeline has slipped. Enterprise customers can trial it in limited Vertex AI preview. Multiple outlets cite July 17 as a target date, but Google has not officially confirmed it.

What is the difference between Gemini 3.5 Pro and Flash?

Flash is fully available, optimized for fast agents and coding, with 1M context and pricing at $1.50/$9 per million tokens. Pro is expected as the family flagship: press reports cite 2M context, Deep Think deep reasoning, higher API pricing, and stronger multi-step task ability—final specs depend on the GA model card.

How much does Gemini 3.5 Pro cost?

Google has not published official Pro pricing. Common press guesses are API rates around $15/$60 per million input/output tokens—roughly 10× Flash—with Deep Think possibly tied to an Ultra-tier subscription near $250/month. Until the official pricing page updates, put only Flash prices in contracts and budgets.

Which Gemini model should I use now?

For production, prioritize Gemini 3.5 Flash (gemini-3.5-flash). For agents, Antigravity, and Enterprise orchestration inside Google's ecosystem, Flash is the 2026 default. Evaluate Pro only within preview access, isolated from production traffic.

Is 2M context worth waiting for Pro?

It depends on the task: if 1M plus RAG/chunking covers 80% of needs, do not wait. If you routinely need a single call with the whole repo and retry cost is extreme, run an A/B after Pro GA. Do not architect around rumored 2M before GA.

12. Summary

What is Gemini 3.5 Pro?—Google's 3.5 family flagship announced at I/O 2026, positioned for hardest reasoning and longest context, but still not publicly GA as of mid-July. On features and pricing, Flash has a complete model card; Pro's 2M context, Deep Think, and $15/$60 figures remain "credible reports, not announced." On release timing, the window slid from "June" to "mid-to-late July"—trust the API changelog, not media countdowns.

The pragmatic path is one line: run agents on Flash today; treat Pro as an incremental upgrade tomorrow. Model wars make headlines every week; engineering teams should fix execution environment, model ID abstraction, and cost dashboards—the rest can wait until Sundar signs the next blog post before you change default configs.

While you wait for Pro, stabilize agents with Flash + Cloud Mac

Call Gemini 3.5 Flash via API for coding agents and run the execution layer on a Cloud Mac mini M4xcodebuild, Git worktrees, and long SSH sessions without disconnects. When Pro reaches GA, change the model ID in an environment variable; the pipeline stays intact. Apple Silicon unified memory suits million-token context workloads; macOS isolation reduces the risk of agents touching files on your local machine.

Start with daily rental for Antigravity / Gemini API smoke tests, then upgrade to monthly for always-on use. kvmboot Cloud Mac mini M4 is the pragmatic starting point for "Flash today, Pro tomorrow"see plan options and turn the GA wait into a deliverable agent pipeline.