GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Google DeepMind WeatherNext 3: Hourly Satellite Forecasts Cut Precipitation Error 60%

Google DeepMind has launched WeatherNext 3, a successor to the WeatherNext models that made headlines in August for matching two-day cyclone accuracy on a three-day horizon. WeatherNext 3 is a different kind of upgrade: instead of a research advance, it is a production deployment with a materially changed architecture. The model is live today in Google Search, Maps, and Gemini.

The central change is the data source. WeatherNext 2 operated on processed atmospheric grids — conventional numerical weather prediction output. WeatherNext 3 ingests raw satellite imagery directly, specifically NASA’s Integrated Multi-satellite Retrievals for GPM (IMERG) dataset alongside Google’s own satellite-radar precipitation reanalysis. The shift from processed grids to raw satellite observations closes a latency gap that had limited how frequently the model could update.

Hourly Cadence

Previous global weather AI models, including WeatherNext’s prior generations, produced forecasts at six-hour or longer intervals. WeatherNext 3 generates new predictions every hour. For rapidly developing events — convective rain, coastal fog, flash-flood precursors — the difference between a six-hour-old forecast and a current one is material. Hourly cycling on a global model was previously confined to the most expensive numerical prediction systems.

Benchmark Numbers

Against established precipitation evaluation datasets:

  • -60% Continuous Ranked Probability Score (CRPS) vs IMERG benchmarks
  • -30% CRPS vs MRMS (Multi-Radar Multi-Sensor) data
  • -10% CRPS vs rain-gauge observations

CRPS measures probabilistic forecast sharpness and reliability simultaneously — a lower score is better. A 60% reduction against IMERG is not marginal. IMERG is the reference standard for global satellite precipitation; beating it at that margin with a model that also runs hourly is the combination that makes this notable.

Resolution

WeatherNext 3 predicts surface temperature and humidity at 5km resolution globally, and wind at 10km resolution. The model generalises to locations not encountered in training — a meaningful practical claim given the sparse observational record in the Southern Hemisphere, open ocean areas, and developing-world regions.

Industry Deployment

DeepMind has positioned WeatherNext 3 for direct industry application, specifically renewable energy operations. Wind and solar farm operators need accurate forecasts of radiation and cloud cover to manage dispatch and grid bidding. At 5km surface resolution with hourly updates, WeatherNext 3 provides forecast granularity that previously required specialised mesoscale models or commercial subscriptions.

The model feeds directly into Google products — Search weather panels, Maps routing decisions, and Gemini assistant responses. Scale of deployment: billions of daily forecasts.

What Changes for Meteorology

WeatherNext 2 established that AI models could match or exceed numerical weather prediction on most headline metrics. WeatherNext 3 demonstrates that the architecture advantage compounds once raw observations replace processed grids as the primary input. The bottleneck had been data normalisation: converting satellite imagery into the grid formats expected by AI models added complexity and latency. Training directly on imagery removes that layer.

The tradeoff is model complexity and training cost. Training on raw sensor data at global scale with NASA GPM and radar reanalysis as joint inputs is a significant compute undertaking. The result is a model that updates hourly rather than every six hours — a shift that matters more for end-user applications than for benchmark leaderboards.