Atmospheric rivers (ARs) significantly impact the Arctic climate system by enhancing atmospheric heat and moisture transport and altering the local energy budget. Developing AR detection tools (ARDTs) is critical yet challenging. This study evaluates 12 ARDTs in the Arctic to assess their performance in representing atmospheric heat (represented by moist static energy) and moisture transport, as well as surface downward longwave radiation (LWD) and precipitation impacts, spanning 2000 to 2019 using ERA5 reanalysis. We find that AR occurrence frequency in the Arctic varies widely, from less than 1% to over 13%, depending on the ARDT. This variability leads to differences in contributions to poleward atmospheric heat (<1%-33%) and moisture (<1%-49%) transport. The highest AR frequency, and corresponding contributions to atmospheric heat and moisture transport, occurs over the Atlantic sector during non-summer seasons for most ARDTs. This region aligns with the primary poleward moisture pathway and the end of climatological mid-latitude storm tracks, highlighting strong connections between Arctic ARs and mid-latitude cyclones. ARs induce significant LWD anomalies, largest in winter, smallest in summer, and also substantially contribute to the seasonal precipitation. Global ARDTs detect fewer ARs with larger anomalies (>100 W m-2 in higher Arctic), but contribute <1% to seasonal climatological LWD and precipitation. In contrast, polar-specific ARDTs detect higher AR occurrences and account for up to 16% of seasonal LWD and 41% precipitation. This suggests that algorithms emphasizing extreme events with large anomalies do not necessarily indicate a large climate radiative and precipitation impact.
The response of the climate system to increased greenhouse gases and other radiative perturbations is governed by a combination of fast and slow feedbacks. Slow feedbacks are typically activated in response to changes in ocean temperatures on decadal timescales and manifest as changes in climatic state with no recent historical analogue. However, fast feedbacks are activated in response to rapid atmospheric physical processes on weekly timescales, and they are already operative in the present-day climate. This distinction implies that the physics of fast radiative feedbacks is present in the historical meteorological reanalyses used to train many recent successful machine-learning-based (ML) emulators of weather and climate. In addition, these feedbacks are functional under the historical boundary conditions pertaining to the top-of-atmosphere radiative balance and sea-surface temperatures. Together, these factors imply that we can use historically trained ML weather emulators to study the response of radiative-convective equilibrium (RCE), and hence the global hydrological cycle, to perturbations in carbon dioxide and other well-mixed greenhouse gases. Without retraining on prospective Earth system conditions, we use ML weather emulators to quantify the fast precipitation response to reduced and elevated carbon dioxed concentrations with no recent historical precedent. We show that the responses from historically trained emulators agree with those produced by full-physics Earth System Models (ESMs). In conclusion, we discuss the prospects for and advantages from using ESMs and ML emulators to study fast processes in global climate.
We examine whether tropical cyclones (TCs) obey ordinary Brownian or anomalous diffusion using a huge ensemble (HENS) of hindcasts for summer 2023. Anomalous diffusion has been inferred for actual TCs from the fluctuations in their tracks from the shortest paths between the initiation and termination of each cyclone. We reproduce the same anomalous diffusion power laws connecting spatial position and time using HENS. In addition, we show that the variance in the position of a single TC across HENS since initiation follows a scaling law with time that, in some cases, corresponds to ballistic motion of the TC through the background atmospheric flow. This determination was enabled by the exceptional statistics determined from thousands of plausible yet counterfactual recreations of 34 individual TCs. HENS consists of 7424 15-day hindcasts initiated from observed atmospheric conditions each day from June 1, 2023 to August 31, 2023 using the ECMWF ERA5 meteorological reanalysis. The hindcasts were generated using NVIDIA's Spherical Fourier Neural Operator (SFNO) machine-learning-based weather and climate emulator. We identify tropical cyclones in HENS using a variant of the Tempest Extremes detection and tracking frameworks for TCs with adjustments to the disposable parameters to minimize the numbers of false positives and negatives relative to the International Best Track Archive for Climate Stewardship (IBTrACS) records for TCs observed in summer 2023. We conclude with the implications of our findings for the predictability of TC tracks and landfall locations on lead times of days to weeks.
The summer of 2023 brought record-breaking heat extremes across the globe. In this work, we investigate how much worse these extremes might have become, given identical large-scale conditions. Numerical weather prediction (NWP) ensembles simulate alternate heatwave storylines, but their limited sample size diminishes their statistical coverage of observed heatwaves and ability to represent worst-case events. Using the Spherical Fourier Neural Operator machine learning (ML) weather model, we generated a huge ensemble of 7,424 storyline forecasts of summer temperature extremes. Despite being constrained by the same initial and boundary conditions and using training data from the same modeling system, the ML ensemble produced extreme heatwave scenarios exceeding temperatures from NWP ensembles. For 30
Abrupt snowmelt, triggered by rain-on-snow events or "snow-eater heat waves," can cause flooding, initiate or accelerate snow drought, and affect water availability. However, the characteristics (e.g., area, duration, and frequency), impacts, and trends of snow-eater heat waves have received little attention. To address this gap, we developed a method to identify snow-eater heat waves and estimate their melt potential using 20th Century Reanalysis version 3 air temperature data, the TempestExtremes algorithm, and an operational snowmelt model (SNOW-17) across 1850-2015. Melt season snow-eater heat waves typically last 3 to 5 days, with three to five events, doubling snowmelt rates. Seven of 11 spring superfloods are shown to coincide with snow-eater heat waves. Since the 1850s, snow-eater heat waves have increased in area and frequency, decreased in duration, and shifted earlier in the melt season. Incorporating snow-eater heat-wave impacts into SNOW-17 enhances extreme melt estimates, improving water management support tools.
In a changing climate, artificial intelligence (AI) weather models have the potential to provide cheaper, faster, and more accurate forecasts of high-impact weather events. To realize this potential and gauge trustworthiness, there is a need for more research on how models learn extreme events and how that learning might be improved. Here, we investigate how a Spherical Fourier Neural Operator (SFNO) learns tropical cyclones (TCs) by saving every checkpoint from training and analyzing storm specific metrics. We find evidence that for some storms the SFNO learns information about TC intensity that it loses later in training. This unlearning pattern is associated with anomalously moist environments and may be due to the model unlearning the relationship between moisture and TC intensity. This work provides a first example of leveraging task-specific training dynamics to further our understanding of how AI weather models learn extreme events.
In recent years, machine-learning (ML) models trained on reanalysis data have rivaled physics-based forecast models in terms of performance skill for global weather forecasting. With increased rollout stability, the question of how these models perform for subseasonal to seasonal (S2S, week 3-8) forecasting has emerged. In this study we run a large set of subseasonal hindcasts over 2004-2023 to evaluate two ML weather forecast models at the S2S time scale, SFNO-HENS (Nvidia, fully ML) and NeuralGCM (Google Research, hybrid). Corresponding hindcasts from the European Centre for Medium-Range Weather Forecasts (ECMWF) are used as a baseline for comparison to a physics-based model. Because our focus is on predicting moisture transport over the Western United States between October and March, we evaluate the models' prediction skill for the Madden-Julian Oscillation (MJO) and its associated teleconnections in the North Pacific. We find that both ML models are competitive with the ECWMF model, with comparable skill in predicting the North Pacific large-scale circulation and the MJO at week 3 and beyond. Even though overall the mid-latitude subseasonal prediction skill remains low, the ML models exhibit interesting behavior such as a realistic propagation of the MJO across the Maritime Continent and realistic teleconnections. A SFNO-HENS sensitivity experiment with altered initial conditions in the tropics demonstrates the stability of the model, and it illustrates the capability of ML models to represent important physical processes of the atmosphere at the S2S time scale.
In a warming climate with more frequent severe weather, artificial intelligence (AI) weather models have the potential to provide cheaper, faster, and more accurate forecasts of high- impact weather events. To realize this potential, there is a need for more research on how models learn extreme events and how that learning might be improved. We investigate how a spherical Fourier neural operator model (SFNO) learns extreme weather by saving every checkpoint throughout training and analyzing a collection of 9 extreme weather events including heatwaves, atmospheric rivers, and tropical cyclones. The SFNO learns heatwaves similarly to other weather days, but we find evidence that the model learns information about atmospheric river and tropical cyclone forecasts that it loses later in training. We propose a possible training strategy to improve the forecasting of extreme events by retaining information from earlier training checkpoints, and provide initial evidence of its utility.
Abstract Since the weather is chaotic, it is necessary to forecast an ensemble of future states. Recently, multiple AI weather models have emerged claiming breakthroughs in deterministic skill. Unfortunately, it is hard to fairly compare ensembles of AI forecasts because variations in ensembling methodology become confounding and the baseline data volume is immense. We address this by scoring lagged initial condition ensembles—whereby an ensemble can be constructed from a library of deterministic hindcasts. This allows the first parameter‐free intercomparison of leading AI weather models' probabilistic skill against an operational baseline. Lagged ensembles of the two leading AI weather models, GraphCast and Pangu, perform similarly even though the former outperforms the latter in deterministic scoring. These results are elaborated upon by sensitivity tests showing that commonly used multiple time‐step loss functions damage ensemble calibration.
Studying low-likelihood high-impact extreme weather and climate events in a warming world requires massiveensembles to capture long tails of multi-variate distributions. In combination, it is simply impossible to generatemassive ensembles, of say 10,000 members, using traditional numerical simulations of climate models at highresolution. We describe how to bring the power of machine learning (ML) to replace traditional numericalsimulations for short week-long hindcasts of massive ensembles, where ML has proven to be successful in terms ofaccuracy and fidelity, at five orders-of-magnitude lower computational cost than numerical methods. Becausethe ensembles are reproducible to machine precision, ML also provides a data compression mechanism toavoid storing the data produced from massive ensembles. The machine learning algorithm FourCastNet (FCN) isbased on Fourier Neural Operators and Transformers, proven to be efficient and powerful in modeling a widerange of chaotic dynamical systems, including turbulent flows and atmospheric dynamics. FCN has already beenproven to be highly scalable on GPU-based HPC systems. We discuss our progress using statistics metrics for extremes adopted from operational NWP centers to showthat FCN is sufficiently accurate as an emulator of these phenomena. We also show how to construct hugeensembles through a combination of perturbed-parameter techniques and a variant of bred vectors to generate alarge suite of initial conditions that maximize growth rates of ensemble spread. We demonstrate that theseensembles exhibit a ratio of ensemble spread relative to RMSE that is nearly identical to one, a key metric ofsuccessful near-term NWP systems. We conclude by applying FCN to severe heat waves in the recent climaterecord.
In Part 1, we created an ensemble based on spherical Fourier neural operators. As initial condition perturbations, we used bred vectors, and as model perturbations, we used multiple checkpoints trained independently from scratch. Based on diagnostics that assess the ensemble's physical fidelity, our ensemble has comparable performance to operational weather forecasting systems. However, it requires orders-of-magnitude fewer computational resources. Here in Part 2, we generate a huge ensemble (HENS), with 7424 members initialized each day of summer 2023. We enumerate the technical requirements for running huge ensembles at this scale. HENS precisely samples the tails of the forecast distribution and presents a detailed sampling of internal variability. HENS has two primary applications: (1) as a large dataset with which to study the statistics and drivers of extreme weather and (2) as a weather forecasting system. For extreme climate statistics, HENS samples events 4 sigma away from the ensemble mean. At each grid cell, HENS increases the skill of the most accurate ensemble member and enhances coverage of possible future trajectories. As a weather forecasting model, HENS issues extreme weather forecasts with better uncertainty quantification. It also reduces the probability of outlier events, in which the verification value lies outside the ensemble forecast distribution.
Abrupt snowmelt, triggered by rain-on-snow events or ``snow-eater heatwaves'', can cause flooding, accelerate snow drought, and impact water availability. Yet the characteristics (e.g., area, duration, and frequency), impacts, and trends of snow-eater heatwaves have received relatively little attention. To address this gap, we developed a method to identify snow-eater heatwaves and estimate their melt potential using Twentieth Century Reanalysis Version 3 air temperature data, the TempestExtremes algorithm, and an operational snowmelt model (SNOW-17) across 1850–2015. Melt season snow-eater heatwaves typically last 3–5 days, with 3–5 events, doubling snowmelt rates. Seven of 11 spring superfloods can be linked with snow-eater heatwaves. Since the 1850s, snow-eater heatwaves have increased in area and frequency, decreased in duration, and shifted earlier in the melt season. Incorporating snow-eater heatwave impacts into SNOW-17 improves extreme snowmelt estimates, providing additional tools to support water management.
Atmospheric rivers (ARs) are extreme weather events that play a crucial role in the global hydrological cycle. As a key mechanism of latent heat transport (LHT), they help maintain energy balance in the climate system. While an AR is characterized by a long, narrow corridor of water vapor associated with a low-level jet stream, there is no unambiguous definition of an AR grounded in geophysical fluid dynamics. Therefore, AR identification is currently performed by a large array of expert-defined, threshold-based algorithms. The variety of algorithms introduces uncertainty in the estimated contribution of ARs to LHT. We calculate the instantaneous eddy LHT from moist, poleward anomalies. Based on the dynamics of the large-scale atmospheric circulation, this quantity is a physics-based upper bound that constrains AR projections from the variety of detection algorithms. We quantify the contribution of ARs to transient eddies, stationary eddies, and transient-stationary eddy interactions, and we show the relative contribution of ARs and other processes, such as dry, equatorward transport. We use this upper bound to quantify ARs' frequency, intensity, and temporal variability. In the historical climate, we find that ARs transport 2.21 PW at the latitude of peak meridional transport in Northern Hemisphere winter, with approximately 0.47 PW of temporal variability. By the end of the century in a future climate projection, at this latitude, we find that AR-induced LHT will increase by 0.5 PW and the corresponding temporal variability will increase by 0.14 PW.
Studying low-likelihood high-impact extreme weather events in a warming world is a significant and challenging task for current ensemble forecasting systems. While these systems presently use up to 100 members, larger ensembles could enrich the sampling of internal variability. They may capture the long tails associated with climate hazards better than traditional ensemble sizes. Due to computational constraints, it is infeasible to generate huge ensembles (comprised of 1,000-10,000 members) with traditional, physics-based numerical models. In this two-part paper, we replace traditional numerical simulations with machine learning (ML) to generate hindcasts of huge ensembles. In Part I, we construct an ensemble weather forecasting system based on Spherical Fourier Neural Operators (SFNO), and we discuss important design decisions for constructing such an ensemble. The ensemble represents model uncertainty through perturbed-parameter techniques, and it represents initial condition uncertainty through bred vectors, which sample the fastest growing modes of the forecast. Using the European Centre for Medium-Range Weather Forecasts Integrated Forecasting System (IFS) as a baseline, we develop an evaluation pipeline composed of mean, spectral, and extreme diagnostics. Using large-scale, distributed SFNOs with 1.1 billion learned parameters, we achieve calibrated probabilistic forecasts. As the trajectories of the individual members diverge, the ML ensemble mean spectra degrade with lead time, consistent with physical expectations. However, the individual ensemble members' spectra stay constant with lead time. Therefore, these members simulate realistic weather states, and the ML ensemble thus passes a crucial spectral test in the literature. The IFS and ML ensembles have similar Extreme Forecast Indices, and we show that the ML extreme weather forecasts are reliable and discriminating.
FourCastNet 3 advances global weather modeling by implementing a scalable, geometric machine learning (ML) approach to probabilistic ensemble forecasting. The approach is designed to respect spherical geometry and to accurately model the spatially correlated probabilistic nature of the problem, resulting in stable spectra and realistic dynamics across multiple scales. FourCastNet 3 delivers forecasting accuracy that surpasses leading conventional ensemble models and rivals the best diffusion-based methods, while producing forecasts 8 to 60 times faster than these approaches. In contrast to other ML approaches, FourCastNet 3 demonstrates excellent probabilistic calibration and retains realistic spectra, even at extended lead times of up to 60 days. All of these advances are realized using a purely convolutional neural network architecture tailored for spherical geometry. Scalable and efficient large-scale training on 1024 GPUs and more is enabled by a novel training paradigm for combined model- and data-parallelism, inspired by domain decomposition methods in classical numerical models. Additionally, FourCastNet 3 enables rapid inference on a single GPU, producing a 90-day global forecast at 0.25°, 6-hourly resolution in under 20 seconds. Its computational efficiency, medium-range probabilistic skill, spectral fidelity, and rollout stability at subseasonal timescales make it a strong candidate for improving meteorological forecasting and early warning systems through large ensemble predictions.
The rapid rise of deep learning (DL) in numerical weather prediction (NWP) has led to a proliferation of models which forecast atmospheric variables with comparable or superior skill than traditional physics-based NWP. However, among these leading DL models, there is a wide variance in both the training settings and architecture used. Further, the lack of thorough ablation studies makes it hard to discern which components are most critical to success. In this work, we show that it is possible to attain high forecast skill even with relatively off-the-shelf architectures, simple training procedures, and moderate compute budgets. Specifically, we train a minimally modified SwinV2 transformer on ERA5 data, and find that it attains superior forecast skill when compared against IFS. We present some ablations on key aspects of the training pipeline, exploring different loss functions, model sizes and depths, and multi-step fine-tuning to investigate their effect. We also examine the model performance with metrics beyond the typical ACC and RMSE, and investigate how the performance scales with model size.
Since the weather is chaotic, forecasts aim to predict the distribution of future states rather than make a single prediction. Recently, multiple data driven weather models have emerged claiming breakthroughs in skill. However, these have mostly been benchmarked using deterministic skill scores, and little is known about their probabilistic skill. Unfortunately, it is hard to fairly compare AI weather models in a probabilistic sense, since variations in choice of ensemble initialization, definition of state, and noise injection methodology become confounding. Moreover, even obtaining ensemble forecast baselines is a substantial engineering challenge given the data volumes involved. We sidestep both problems by applying a decades-old idea -- lagged ensembles -- whereby an ensemble can be constructed from a moderately-sized library of deterministic forecasts. This allows the first parameter-free intercomparison of leading AI weather models' probabilistic skill against an operational baseline. The results reveal that two leading AI weather models, i.e. GraphCast and Pangu, are tied on the probabilistic CRPS metric even though the former outperforms the latter in deterministic scoring. We also reveal how multiple time-step loss functions, which many data-driven weather models have employed, are counter-productive: they improve deterministic metrics at the cost of increased dissipation, deteriorating probabilistic skill. This is confirmed through ablations applied to a spherical Fourier Neural Operator (SFNO) approach to AI weather forecasting. Separate SFNO ablations modulating effective resolution reveal it has a useful effect on ensemble dispersion relevant to achieving good ensemble calibration. We hope these and forthcoming insights from lagged ensembles can help guide the development of AI weather forecasts and have thus shared the diagnostic code.
Atmospheric rivers (ARs) are extreme weather events that can alleviate drought or cause billions of US dollars in flood damage. By transporting significant amounts of latent energy towards the poles, they are crucial to maintaining the climate system's energy balance. Since there is no first-principle definition of an AR grounded in geophysical fluid mechanics, AR identification is currently performed by a multitude of expert-defined, threshold-based algorithms. The variety of AR detection algorithms has introduced uncertainty into the study of ARs, and the thresholds of the algorithms may not generalize to new climate datasets and resolutions. We train convolutional neural networks (CNNs) to detect ARs while representing this uncertainty; we name these models ARCNNs. To detect ARs without requiring new labeled data and labor-intensive AR detection campaigns, we present a semi-supervised learning framework based on image style transfer. This framework generalizes ARCNNs across climate datasets and input fields. Using idealized and realistic numerical models, together with observations, we assess the performance of the ARCNNs. We test the ARCNNs in an idealized simulation of a shallow-water fluid in which nearly all the tracer transport can be attributed to AR-like filamentary structures. In reanalysis and a high-resolution climate model, we use ARCNNs to calculate the contribution of ARs to meridional latent heat transport, and we demonstrate that this quantity varies considerably due to AR detection uncertainty.