GeoscienceAnalysis

AI Weather Models Are Coming for Traditional Forecasting

GraphCast, Pangu-Weather, and FourCastNet showed that learned models can match or beat physics-based forecasts on reanalysis data. ECMWF now runs one of them operationally. The limits matter as much as the results.

Key facts

  • GraphCast predicts hundreds of variables 10 days ahead at 0.25° resolution in under a minute, and outperformed operational deterministic systems on 90% of 1,380 verification targets (Science, 2023).
  • Pangu-Weather, trained on 39 years of data with 3D networks carrying Earth-specific priors, reported stronger deterministic results than ECMWF's IFS on reanalysis data (Nature, 2023).
  • These comparisons use reanalysis data, not independent observations, and the models learn from archives that physics-based systems help produce.
  • ECMWF has run its own machine-learning model, AIFS, operationally since February 2025, with an ensemble version since July 2025.
On this page

The end of the supercomputer monopoly — partly

For sixty years, weather forecasting has been dominated by numerical weather prediction (NWP): systems of partial differential equations solved on some of the largest computers on Earth. ECMWF's IFS, NOAA's GFS, and Météo-France's Arpège are the reference systems in that tradition.

Since 2022, a second family of models has entered the conversation. Instead of solving physics, they learn statistical relationships between atmospheric states from decades of reanalysis data — the archives produced by data assimilation systems that combine observations with short-range model forecasts. Three models established the case.

The AI challengers

GraphCast (Google DeepMind)

GraphCast is a graph neural network trained directly on reanalysis data. According to the paper published in Science, it predicts hundreds of weather variables for the next 10 days at 0.25° resolution globally in under one minute, and significantly outperformed the most accurate operational deterministic systems on 90% of 1,380 verification targets. The same paper reports better severe-event prediction for tropical cyclone tracking, atmospheric rivers, and extreme temperatures (Lam et al., 2023).

Pangu-Weather (Huawei)

Pangu-Weather uses 3D deep networks with Earth-specific priors and a hierarchical temporal aggregation strategy to reduce error accumulation over long lead times. Trained on 39 years of global data, it reported stronger deterministic forecast results than ECMWF's operational IFS on reanalysis data in every tested variable, and higher tropical-cyclone tracking accuracy than ECMWF-HRES when initialized with reanalysis data (Bi et al., 2023).

One detail worth noting for anyone reading the literature: the Pangu-Weather paper carries a published correction, and its headline comparison is against a single deterministic system in a specific evaluation setting — not against an ensemble, and not on independent observations.

FourCastNet (NVIDIA)

FourCastNet approaches the problem spectrally, using Adaptive Fourier Neural Operators. It is the earliest of the three and remains an arXiv preprint rather than a peer-reviewed paper, which is worth stating plainly when citing its results (Pathak et al., 2022).

What these comparisons actually measure

The distinction between reanalysis and observation is the part most summaries skip.

Verification against reanalysis asks how well a model predicts the state estimate that data assimilation produced. That estimate is built partly from short-range NWP forecasts, so it is not a neutral referee. It is also the data both families of models are trained and tuned on. Beating a physics-based model on reanalysis is a real result, but it is not the same as beating it on independent radiosonde, aircraft, and satellite observations, and it says little about events with no analogue in the archive.

Three further caveats:

  1. Training depends on the physics-based stack. Learned models are trained on reanalysis archives; producing those archives requires the NWP systems being compared against, plus the observing network that feeds them.
  2. Interpolation is not simulation. A learned model has no explicit representation of the atmosphere. It cannot guarantee conservation laws, and it has no physical mechanism for extrapolating into regimes absent from its training period.
  3. Operational use is a higher bar than a verification table. A forecast that is occasionally excellent and occasionally broken is harder to depend on than one with a slightly worse average and a stable failure mode.

Why the shift still matters

Inference cost, not accuracy, is what changed the operational conversation. Where a physics-based forecast requires a large HPC allocation per run, a trained model can be evaluated on a single accelerator in seconds to minutes. That changes what is affordable:

  • Ensembles. Large ensembles become cheaper to generate, which matters for probabilistic forecasting more than deterministic skill does.
  • Tropical cyclones. Track prediction is the strongest reported result for learned models; intensity and rapid intensification remain harder.
  • Rapid refresh. Running a model hourly, or on demand for a specific region, becomes feasible in a way it is not with a full NWP cycle.

What comes next

This is no longer a research curiosity. ECMWF has been running its own machine-learning model, the Artificial Intelligence Forecasting System (AIFS), operationally since 25 February 2025, added an ensemble version in July 2025, and upgraded both in May 2026. AIFS produces four runs per day, 6-hourly steps out to 15 days, on an approximately 0.25° grid, and its output is published openly (ECMWF).

The realistic near-term picture is not replacement. It is a mixed system: physics-based ensembles as the backbone, learned models for speed, specific variables, and ensemble size — with evaluation that has to keep up. WeatherBench 2 exists for exactly that purpose, and it is the right place to look before trusting any single verification number, including the ones in this article.

Getting started

Two low-friction entry points if you want to work with this data yourself:

# ECMWF's open data client for ERA5 and other Copernicus datasets
pip install cdsapi

# WeatherBench 2 evaluation framework (clone the repository)
git clone https://github.com/google-research/weatherbench2

The models are cheap to run and the data is open. The evaluation is the hard part.

Sources

  1. Learning skillful medium-range global weather forecasting — Science (Lam et al., 2023)paperDOI: 10.1126/science.adi2336Retrieved Sep 16, 2026
  2. Accurate medium-range global weather forecasting with 3D neural networks — Nature (Bi et al., 2023)paperDOI: 10.1038/s41586-023-06185-3Retrieved Sep 16, 2026
  3. FourCastNet: A Global Data-driven High-resolution Weather Model using Adaptive Fourier Neural Operators — arXiv preprint (Pathak et al., 2022)paperRetrieved Sep 16, 2026
  4. AIFS Machine Learning data — ECMWFdocsRetrieved Sep 16, 2026
  5. WeatherBench 2: evaluation framework for data-driven weather models — Google ResearchcodeRetrieved Sep 16, 2026

Masood Sultan

Computational geoscientist and AI engineer. Writes about machine learning systems, geospatial data, and the methods behind scientific computing. Published in Journal of Applied Geophysics (2020).