The question of where Hurricane Beryl might eventually make landfall was not insignificant when it was traveling over the Atlantic in July 2024, heading roughly southwest. Conventional forecast models were pointing to southern Mexico by executing intricate fluid dynamics equations on large computing clusters. Google DeepMind‘s AI weather model, GraphCast, was pointing in a different direction. GraphCast had predicted the steep northward shift that would send Beryl ashore in southern Texas almost a week before the storm’s actual behavior verified it. Mexico was still receiving calls from the world’s most prestigious and costly forecasting system, the top European physics-based model. The AI was correct. It wasn’t the supercomputer.

In the current discussion regarding what AI weather models can truly accomplish and whether the meteorological community’s long-standing investment in conventional numerical weather prediction is truly under threat, that instance has become one of the most frequently referenced examples. Based on the Atlantic hurricane season of 2024 and the years of benchmark testing that preceded it, the simple answer is definitely yes. Major organizations like NOAA and the European Center for Medium-Range Weather Forecasts have shifted from closely monitoring these systems to integrating them operationally, though not in every situation and not without significant warnings.
Understanding the fundamental causes of AI models’ performance is important since it clarifies both their advantages and disadvantages. The physical equations governing atmospheric fluid dynamics are solved in traditional numerical weather prediction. ECMWF’s flagship method requires supercomputing hardware that costs hundreds of millions of dollars to maintain and takes hours to compile a 10-day global prediction, demonstrating how computationally expensive this is.
GraphCast operates in a different way. Instead of calculating how atmospheric conditions change over time from physical fundamental principles, it was trained on ERA5, a reanalysis dataset that covers more than 40 years of global atmospheric records. It takes less than a minute to produce a 10-day worldwide forecast using just one contemporary GPU. There is a noticeable difference in speed. According to benchmarks, Huawei’s rival AI model, Pangu-Weather, operates 10,000 times faster than traditional ensemble methods.
The edge in track prediction during Beryl was not accidental. GraphCast outperformed traditional models in both Atlantic and Pacific hurricane basins during the first five days of storm formation, according to a report given at a hurricane prediction improvement conference in Miami in late 2024 that examined the program’s performance from 2021 to 2024. Emergency managers employing GraphCast guidance would have had significantly more warning time because it met accuracy targets 12 hours earlier than the US Global Forecast System.
Another example was Hurricane Francine in 2024, when ECMWF’s AI-based AIFS model predicted the Louisiana landfall ten days ahead of most competing forecasts. Additionally, DeepMind later showed that GraphCast would have given National Hurricane Center forecasters earlier warning signs during Hurricane Otis in 2023, which quickly developed before hitting Mexico in a way that caught standard models completely off guard. When DeepMind presented them with the retrospective analysis, they made this clear.
The institutional meteorology community has responded in a measured but tangible way. AIFS became the first significant national or international meteorological service to formally implement an AI model in addition to its conventional methods when ECMWF upgraded it to operational status in 2024. Using the agency’s own Global Data Assimilation System data, NOAA refined GraphCast to create its own version, AIGFS, which it subsequently operationalized. During the 2025 hurricane season, the National Hurricane Center ran live AI forecasts alongside conventional models for the first time. These are no longer pilot programs. They are operational tools used in actual forecast rooms where actual decisions on the deployment of emergency resources and evacuations are made.
The researchers who work on these systems are open about their shortcomings. AI models can track a hurricane’s center path with remarkable accuracy, but they can smooth over the intensity gradients that are important for damage prediction. However, they struggle with the fine-scale wind field structure inside a storm. Another area where AI systems frequently do poorly is extreme localized rainfall, the type that causes catastrophic flooding independent of the storm track question. Additionally, the “black swan” issue is real: storms that emerge in ways that deviate from past observational patterns—a category that is anticipated to expand as climate change pushes weather systems into previously unheard-of territory—remain a known weakness because these models learn on historical data.
