GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Google Enters I/O Trailing Mythos by 13 Points and GPT-5.5 by 8 on SWE-Bench — Analysts Expect a Point Release, Not Gemini 4.0

Google I/O 2026 opens Tuesday May 19 at Shoreline Amphitheatre. The keynote slot carries Gemini model expectations that, for the first time in I/O history, Google enters defending third place.

On SWE-Bench Verified — the coding benchmark that most closely tracks production agentic performance — the current leaderboard reads: Claude Mythos Preview at 93.9%, GPT-5.5 at 88.7%, Gemini 3.1 Pro at 80.6%. That 13-point gap to Mythos and 8-point gap to GPT-5.5 is not noise. At 500 evaluated instances, it translates to roughly 65 additional correct patches per run. On Artificial Analysis’ Intelligence Index, GPT-5.5 xhigh and high hold the top two positions; Claude Opus 4.7 follows in third. Gemini 3.1 Pro Preview is further back.

What the Keynote Will Not Deliver

Analysts tracking Google’s release cadence now expect Gemini 3.5 Pro rather than the 4.0 full-generation model that earlier previews suggested. The pattern supports that read: Gemini 1 → 1.5 at I/O 2024; Gemini 2 → 2.5 at I/O 2025; Gemini 3 → 3.5 at I/O 2026. A full-generation increment typically requires 12-18 months between deployments; Google released Gemini 3 Pro in late 2025.

Gemini 3.5 Pro, as described in pre-I/O signal, targets software engineering and advanced multi-step reasoning — the exact benchmark categories where the current deficit is widest. Expected improvements in long-form output accuracy and context handling are consistent with narrowing the SWE-Bench gap, but leapfrogging Mythos at 93.9% requires more than a point release.

The Real Strategy: Behavior, Not Benchmarks

Google’s strategic answer to trailing on model quality is a platform move. “Gemini Intelligence” — confirmed ahead of I/O at the Android Show on May 12 — is a proactive agentic layer that operates across phones, watches, cars, and glasses without requiring a user prompt. It builds shopping carts, books reservations, reads screen context across apps, and autofills forms using data pulled from Gmail and Calendar.

The framing is deliberate. Where Anthropic and OpenAI compete model-to-model on benchmark tables, Google is arguing for device-level ubiquity: the AI that’s already on 3 billion Android devices and embedded in every Chrome instance doesn’t need to win on SWE-Bench to win the market.

The Android Show confirmed specific capabilities: Gemini in Chrome on Android for cross-site research and form completion; Rambler for filler-word-filtered voice-to-text; Create My Widget for natural-language widget generation; and Gemini Intelligence-powered multi-app task automation starting with Samsung Galaxy and Google Pixel this summer.

Project Astra and the API Developers Are Watching

Project Astra — Google’s real-time multimodal agent capable of processing live audio and video simultaneously — is expected to move from limited preview into production API access at I/O. This would be the first time developers can build against a streaming video + reasoning + conversation API from a major lab. OpenAI and Anthropic both offer computer use, but neither provides live video input as a first-class API input type at production scale.

If Astra ships as a production Vertex AI API at I/O, it opens a category that does not currently exist in the market — real-time visual reasoning agents — before any competitor has a comparable offering. That outcome would shift the competitive narrative away from benchmark tables.

A separate signal from this week: leaked video clips in the Gemini app showed content that no current Gemini product can generate — 4K footage with synchronised audio, object swapping via chat instructions. This is consistent with a “Gemini Omni” unified multimodal API that collapses Gemini 3.1 text, Veo 3.1 video, and Imagen 4.0 image into a single API call. If that ships, it directly addresses one of the most cited developer pain points: maintaining three separate integrations for a multimodal application.

The Numbers That Define the Keynote

BenchmarkMythosGPT-5.5Gemini 3.1 Pro
SWE-Bench Verified93.9%88.7%80.6%
Intelligence Index (AA)—#1 (xhigh)#4
Terminal-Bench 2.0—82.0% (Codex)—

To matter on benchmarks Tuesday, Google needs to post above 88.7% on SWE-Bench Verified for Gemini 3.5 Pro. Anything below that sustains the current gap narrative. Anything above it forces a reframe.

The I/O keynote starts 10 AM PT May 19. Developer sessions run through May 20. The Astra API session — if it appears — is the one session that changes what developers can build, regardless of what the leaderboard says.