Gemini

Gemini 3.5 Flash launches at Google I/O 2026: the Flash model that beats last-gen Pro at coding and AI agents

Abstract illustration of artificial intelligence

On May 19, 2026 at Google I/O 2026, Google announced Gemini 3.5 Flash — the flagship of the 3.5 generation, positioned as the "default" model for coding and AI agents. The most striking part: this is only the Flash tier (traditionally the lightweight, fast, cheap line), yet according to the DeepMind site it already beats last generation's Gemini 3.1 Pro across a host of coding and agentic benchmarks — while being noticeably faster and cheaper.

Update 24/09/2026: Gemini 3.5 Flash is generally available and Google's official API pricing now lists $1.50 input / $9.00 output per 1M tokens, matching the press figures below. Gemini 3.5 Pro has still not been released broadly: on 21/07/2026 Google said it "is currently testing with partners". Google has since shipped newer Flash models — Gemini 3.6 Flash (21/07), 3.7 Flash (13/08) and 3.8 Flash (02/09/2026) per the Gemini API changelog — so 3.5 Flash is no longer Google's newest Flash model.

Quick summary

  • When: May 19, 2026, at Google I/O 2026.
  • What: Gemini 3.5 Flash — the "default" model for coding & AI agents; status "available now" (per DeepMind). The 3.5 Pro version was announced as "coming soon" and, as of 24/09/2026, is still testing with partners.
  • Highlight: The Flash model beats last generation's Gemini 3.1 Pro on coding/agentic benchmarks.
  • Speed: ~4x other frontier models; a 1 million (1M) token context window.
  • API pricing: $1.50 / $9 per 1M tokens (input/output) — first reported by the press, now confirmed on Google's pricing page.

What's new at Google I/O 2026?

At its annual event, Google put Gemini 3.5 Flash front and center, positioning it as the default model for coding and building AI agents. Per the DeepMind site, the model is "available now", while the more powerful version — Gemini 3.5 Pro — was introduced as "coming soon".

Beyond the model, Google also updated the Gemini app with a slate of new features: Daily Brief (a daily summary), Gemini Spark (an agent running 24/7) and Gemini Omni (AI video). The company also announced that Gemini has reached more than 900 million monthly users, serving 230+ countries and 70+ languages.

Laptop displaying lines of programming code
Gemini 3.5 Flash is positioned as the default model for coding and AI agents. Photo: Pexels

Benchmarks: the Flash model beats last-gen Pro

According to the figures on the DeepMind site, Gemini 3.5 Flash posts impressive results across coding and agentic benchmark suites:

  • Terminal-Bench 2.1: 76.2%
  • MCP Atlas: 83.6%
  • GDPval-AA: 1656 Elo
  • Finance Agent v2: 57.9% — versus 43.0% for last generation's Gemini 3.1 Pro
  • CharXiv: 84.2%
Table — Gemini 3.5 Flash benchmarks (DeepMind figures)
BenchmarkGemini 3.5 FlashGemini 3.1 Pro
Terminal-Bench 2.176.2%—
MCP Atlas83.6%—
GDPval-AA1656 Elo—
Finance Agent v257.9%43.0%
CharXiv84.2%—

Worth noting: the Finance Agent v2 figure shows the new Flash model clearly edging out the previous generation's Pro on a complex agentic task — evidence for the "faster, cheaper, but no weaker" message.

Speed, context and pricing

Per the announcement, Gemini 3.5 Flash runs roughly 4x faster than other frontier models, together with a 1 million (1M) token context window — enough to handle a large codebase or a long document in a single pass.

On cost, the API price first reported by the press — and since confirmed on Google's official pricing page (paid tier) — is $1.50 per 1M input tokens and $9 per 1M output tokens (output includes thinking tokens). That is a very competitive level for a model with coding/agentic capability of this caliber — and part of why Google positions it as the "default" choice.

Table — Gemini 3.5 Flash API pricing (Google official pricing, paid tier)
Token typePrice / 1M tokens
Input$1.50
Output (incl. thinking tokens)$9.00

What it means for product builders

The fact that a Flash tier — traditionally the cost-optimized line — can beat the previous generation's Pro shows that the cost per unit of AI capability keeps falling fast. For teams building AI agents, coding tools, and internal assistants, this opens up the ability to run complex tasks at a price point that was previously reserved for premium models.

The ~4x speed and 1M-token context also have practical implications: agents respond faster, read more context in a single call, and reduce the number of loops and the overall cost.

The Namtech perspective

A model that is both powerful and cheap is good news for any development team. However, keep one important point in mind: whether you use Gemini or any other cloud AI, your data is still sent to a foreign provider's infrastructure. For workloads involving sensitive data — customer information, internal records, core source code — this is a compliance risk (Vietnam's Personal Data Protection Law No. 91/2025/QH15, cross-border data transfers) and a question of control.

A reasonable balance: use powerful cloud models like Gemini 3.5 Flash for general tasks, but for core data consider in-house AI running on your own company's infrastructure so the data never leaves the organization.

AI robot interacting with a digital interface
A faster and cheaper model for AI agents and coding. Photo: Tara Winstead / Pexels

FAQ

Is Gemini 3.5 Flash usable yet?

Per the DeepMind site, Gemini 3.5 Flash is "available now" as of launch (May 19, 2026). The more powerful version — Gemini 3.5 Pro — was introduced as "coming soon"; as of 24/09/2026 Google says it is still testing with partners.

How much does the Gemini 3.5 Flash API cost?

$1.50 per 1M input tokens and $9.00 per 1M output tokens (paid tier, output includes thinking tokens), per Google's official Gemini API pricing page.

Why does the Flash model beat last generation's Pro?

This is a newer-generation model (3.5 versus 3.1). Per DeepMind's figures, the clearest example is Finance Agent v2: Gemini 3.5 Flash scores 57.9% versus 43.0% for Gemini 3.1 Pro — reflecting advances in architecture and training between the two generations.

Powerful and cheap is good — but the data still goes to the cloud

For sensitive data, Namtech deploys a private in-house AI platform — the model runs on your infrastructure, the data stays on-premises and never leaves the organization.

Book a free consultation

Note: This article was first compiled from public sources on 22/06/2026. Last updated 24/09/2026: API pricing confirmed against Google's official pricing page, Gemini 3.5 Pro status and newer Flash releases added. For reference only.

Get started

Start with a free assessment

To define the right package and detailed scope, Namtech offers a short, no-cost assessment.

We reply within 1 business day. No spam, we never share your info.