We couple Forward Flux Sampling (FFS), a non-equilibrium rare-event technique from statistical mechanics, to a neural weather emulator (SDL-WXFormer, 1° grid spacing) to estimate conditional tropical cyclogenesis rates, or how often a tropical cyclone achieves a hurricane-level central pressure, without modifying model dynamics. Tropical cyclogenesis rates vary by orders of magnitude across regimes, yet direct ensemble sampling cannot resolve this variability at operationally feasible ensemble sizes. FFS decomposes the rare disturbance to mature cyclone intensification path into a flux through an initial interface pressure and a product of conditional crossing probabilities across four intermediate interface pressures. We use the 1° emulator because FFS requires O(10^4) model trajectories per initial condition, and because the model's calibrated stochastic layers provide the necessary exploratory spread. Applied to 98 Atlantic basin initial conditions spanning 21 August - 8 October 2022, FFS resolves genesis rates spanning nearly three orders of magnitude, capturing a seasonal cycle qualitatively consistent with observations. A self-consistency check comparing FFS rates to independent direct-sampling rates yields a mean ratio of 1.03 +/- 0.15 across all initial conditions. Computational enhancement factors range from 3X (most active environment) to 140X (most suppressed), with a geometric mean of 14X. Three case studies illustrate the physical diagnostics the method provides: the rate-limiting step is initial tropical organization for the Earl environment, uniformly high crossing probabilities for the Fiona precursor environment, and a compound barrier at the final intensification stages for the Ian environment. More efficient emulators would enable application of FFS to finer resolutions.
Artificial Intelligence (AI) weather prediction (AIWP) models are powerful tools for medium-range forecasts but often lack physical consistency, leading to outputs that violate conservation laws. This study introduces a set of novel physics-based schemes designed to enforce the conservation of global dry air mass, moisture budget, and total atmospheric energy in AIWP models during both training and inference. The schemes are highly modular, allowing for seamless integration into a wide range of AI model architectures. Forecast experiments are conducted to demonstrate the benefit of conservation schemes using FuXi, an example AIWP model, modified and adapted for 1.0 grid spacing. Verification results show that the conservation schemes can guide the model in producing forecasts that obey conservation laws. The forecast skills of upper-air and surface variables are also improved, with longer forecast lead times receiving larger benefits. Notably, large performance gains are found in the total precipitation forecasts, owing to the reduction of drizzle bias. The proposed conservation schemes establish a foundation for implementing other physics-based schemes in the future. They also provide a new way to integrate atmospheric domain knowledge into the design and refinement of AIWP models.
Climate simulations, at all grid resolutions, rely on approximations that encapsulate the forcing due to unresolved processes on resolved variables, known as parameterizations. Parameterizations often lead to inaccuracies in climate models, with significant biases in the physics of key climate phenomena. Advances in artificial intelligence (AI) are now directly enabling the learning of unresolved processes from data to improve the physics of climate simulations. Here, we introduce a flexible framework for developing and implementing physics- and scale-aware machine learning parameterizations within climate models. We focus on the ocean and sea-ice components of a state-of-the-art climate model by implementing a spectrum of data-driven parameterizations, ranging from complex deep learning models to more interpretable equation-based models. Our results showcase the viability of AI-driven parameterizations in operational models, advancing the capabilities of a new generation of hybrid simulations, and include prototypes of fully coupled atmosphere-ocean-sea-ice hybrid simulations. The tools developed are open source, accessible, and available to all.
Recent advancements in artificial intelligence (AI) numerical weather prediction (NWP) have transformed atmospheric modeling. AI NWP models outperform state-of-the-art conventional NWP models like the European Center for Medium Range Weather Forecasting’s (ECMWF) Integrated Forecasting System (IFS) on several global metrics while requiring orders of magnitude fewer computational resources. However, existing AI NWP models still face limitations due to training datasets and dynamic timestep choices, often leading to artifacts that affect performance. To begin to address these challenges, we introduce the Community Research Earth Digital Intelligence Twin (CREDIT) framework, developed at the NSF National Center for Atmospheric Research (NCAR). CREDIT is a flexible, scalable, foundational research platform for training and deploying AI NWP models, providing an end-to-end pipeline for data preprocessing, model training, and evaluation. The CREDIT framework supports both existing architectures and the development of new models. We showcase this flexibility with WXFormer, a novel multiscale vision transformer designed to predict atmospheric states while mitigating common AI NWP pitfalls through techniques like spectral normalization, intelligent padding, and multi-step training. Additionally, we train the FuXi architecture within the CREDIT framework for comparison. Our results demonstrate that both FuXi and WXFormer, trained on hybrid sigma-pressure level ERA5 sampled at 6-h intervals, generally achieve better performance than the IFS High-Resolution (IFS HRES) on 10-day forecasts, offering potential improvements in efficiency and accuracy. The modular nature of CREDIT fosters collaboration, enabling researchers to experiment with models, datasets, and training options.
Here we explore the relative contribution of the Madden-Julian Oscillation (MJO) and El Niño Southern Oscillation (ENSO) to midlatitude subseasonal predictive skill of upper atmospheric circulation over the North Pacific, using an inherently interpretable neural network applied to pre-industrial control runs of the Community Earth System Model version 2. We find that this interpretable network generally favors the state of ENSO, rather than the MJO, to make correct predictions on a range of subseasonal lead times and predictand averaging windows. Moreover, the predictability of positive circulation anomalies over the North Pacific is comparatively lower than that of their negative counterparts, especially evident when the ENSO state is important. However, when ENSO is in a neutral state, our findings indicate that the MJO provides some predictive information, particularly for positive anomalies. We identify three distinct evolutions of these MJO states, offering fresh insights into opportune forecasting windows for MJO teleconnections.
AbstractArtificial intelligence (AI) and machine learning (ML) pose a challenge for achieving science that is both reproducible and replicable. The challenge is compounded in supervised models that depend on manually labeled training data, as they introduce additional decision‐making and processes that require thorough documentation and reporting. We address these limitations by providing an approach to hand labeling training data for supervised ML that integrates quantitative content analysis (QCA)—a method from social science research. The QCA approach provides a rigorous and well‐documented hand labeling procedure to improve the replicability and reproducibility of supervised ML applications in Earth systems science (ESS), as well as the ability to evaluate them. Specifically, the approach requires (a) the articulation and documentation of the exact decision‐making process used for assigning hand labels in a “codebook” and (b) an empirical evaluation of the reliability” of the hand labelers. In this paper, we outline the contributions of QCA to the field, along with an overview of the general approach. We then provide a case study to further demonstrate how this framework has and can be applied when developing supervised ML models for applications in ESS. With this approach, we provide an actionable path forward for addressing ethical considerations and goals outlined by recent AGU work on ML ethics in ESS.
The Madden–Julian Oscillation (MJO) is a dominant source of subseasonal atmospheric variability in the tropics and significantly impacts global weather and climate predictability. Changes in its activity and predictability due to human-induced global climate change have profound implications for future global weather prediction. Here we investigate changes in MJO predictability in reanalysis and climate model data and find that MJO predictability has increased over the past century. This increase can be attributed to anthropogenic warming and continues during the twenty-first century in projections. The increased predictability is accompanied by stronger MJO amplitude, more regular oscillation patterns and organized eastward propagation under global warming. Our results suggest that greenhouse warming will increase the predictability of the MJO, with far-reaching consequences for global weather prediction.
Spatial navigation involves the formation of coherent representations of a map-like space, while simultaneously tracking current location in a primarily unsupervised manner. Despite a plethora of neurophysiological experiments revealing spatially-tuned neurons across the mammalian neocortex and subcortical structures, it remains unclear how such representations are acquired in the absence of explicit allocentric targets. Drawing upon the concept of predictive learning, we utilize a biologically plausible learning rule which utilizes sensory-driven observations with internally-driven expectations and learns through a contrastive manner to better predict sensory information. The local and online nature of this approach is ideal for deployment to neuromorphic hardware for edge-applications. We implement this learning rule in a network with the feedforward and feedback pathways known to be necessary for spatial navigation. After training, we find that the receptive fields of the modeled units resemble experimental findings, with allocentric and egocentric representations in the expected order along processing streams. These findings illustrate how a local and self-supervised learning method for predicting sensory information can extract latent structure from the environment.
Robust quantification of predictive uncertainty is critical for understanding factors that drive weather and climate outcomes. Ensembles provide predictive uncertainty estimates and can be decomposed physically, but both physics and machine learning ensembles are computationally expensive. Parametric deep learning can estimate uncertainty with one model by predicting the parameters of a probability distribution but do not account for epistemic uncertainty.. Evidential deep learning, a technique that extends parametric deep learning to higher-order distributions, can account for both aleatoric and epistemic uncertainty with one model. This study compares the uncertainty derived from evidential neural networks to those obtained from ensembles. Through applications of classification of winter precipitation type and regression of surface layer fluxes, we show evidential deep learning models attaining predictive accuracy rivaling standard methods, while robustly quantifying both sources of uncertainty. We evaluate the uncertainty in terms of how well the predictions are calibrated and how well the uncertainty correlates with prediction error. Analyses of uncertainty in the context of the inputs reveal sensitivities to underlying meteorological processes, facilitating interpretation of the models. The conceptual simplicity, interpretability, and computational efficiency of evidential neural networks make them highly extensible, offering a promising approach for reliable and practical uncertainty quantification in Earth system science modeling. In order to encourage broader adoption of evidential deep learning in Earth System Science, we have developed a new Python package, MILES-GUESS (https://github.com/ai2es/miles-guess), that enables users to train and evaluate both evidential and ensemble deep learning.
We develop and compare model-error representation schemes derived from data assimilation increments and nudging tendencies in multidecadal simulations of the Community Atmosphere Model, version 6. Each scheme applies a bias correction during simulation runtime to the zonal and meridional winds. We quantify the extent to which such online adjustment schemes improve the model climatology and variability on daily to seasonal timescales. Generally, we observe about a 30% improvement to annual upper-level zonal winds, with largest improvements in boreal spring (around 35%) and winter (around 47%). Despite only adjusting the wind fields, we additionally observe around 20% improvement to annual precipitation over land, with the largest improvements in boreal fall (around 36%) and winter (around 25%), and around 50% improvement to annual sea-level pressure, globally. With mean-state adjustments alone, the dominant pattern of boreal low-frequency variability over the Atlantic (the North Atlantic Oscillation) is significantly improved. Additional stochasticity increases the modal explained variances further, which brings the variability closer to the observed value. A streamfunction tendency decomposition reveals that the improvement is due to an adjustment to the high- and low-frequency eddy-eddy interaction terms. In the Pacific, the mean-state adjustment alone led to an erroneous deepening of the Aleutian low, but this was remedied with the addition of stochastically selected tendencies. Finally, from a practical standpoint, we discuss the performance of using data assimilation increments versus nudging tendencies for an online model-error representation. We develop and compare online model-error representation schemes derived from data assimilation increments and nudging tendencies. Generally, we observe significant improvements to annual upper-level zonal wind. Despite only adjusting the wind fields, we additionally observe an improvement to annual precipitation over land and annual sea-level pressure. We additionally quantify improvements to model variability over the North Atlantic and North Pacific. Finally, we discuss the performance of using data assimilation increments versus nudging tendencies for an online model-error representation. image
AbstractHere we explore the relative contribution of the Madden‐Julian Oscillation (MJO) and El Niño Southern Oscillation (ENSO) to midlatitude subseasonal predictive skill of upper atmospheric circulation over the North Pacific, using an inherently interpretable neural network applied to pre‐industrial control runs of the Community Earth System Model version 2. We find that this interpretable network generally favors the state of ENSO, rather than the MJO, to make correct predictions on a range of subseasonal lead times and predictand averaging windows. Moreover, the predictability of positive circulation anomalies over the North Pacific is comparatively lower than that of their negative counterparts, especially evident when the ENSO state is important. However, when ENSO is in a neutral state, our findings indicate that the MJO provides some predictive information, particularly for positive anomalies. We identify three distinct evolutions of these MJO states, offering fresh insights into opportune forecasting windows for MJO teleconnections.
Abstract. Here we present a framework for the first atmosphere-connected supermodel using state-of-the-art atmospheric models. The Community Atmosphere Model (CAM) versions 5 and 6 exchange information interactively while running, a process known as supermodeling. The primary goal of this approach is to synchronize the models, allowing them to compensate for each other's systematic errors in real-time, in part by increasing the dimentionality of the system. In this study, we examine a single untrained supermodel where each model version is equally weighted in creating pseudo-observations. We demonstrate that the models synchronize well without decreased variability, particularly in storm track regions, across multiple timescales and for variables where no information has been exchanged. Synchronization is less pronounced in the tropics, and in regions of lesser synchronization we observe a decrease in high-frequency variability. Additionally, the low-frequency modes of variability (North Atlantic Oscillation and Pacific North American Pattern) are not degraded compared to the base models. For some variables, the mean bias is reduced compared to control simulations of each model version as well as the non-interactive ensemble mean.
Recent advancements in artificial intelligence (AI) for numerical weather prediction (NWP) have significantly transformed atmospheric modeling. AI NWP models outperform traditional physics-based systems, such as the Integrated Forecast System (IFS), across several global metrics while requiring fewer computational resources. However, existing AI NWP models face limitations related to training datasets and timestep choices, often resulting in artifacts that reduce model performance. To address these challenges, we introduce the Community Research Earth Digital Intelligence Twin (CREDIT) framework, developed at NSF NCAR. CREDIT provides a flexible, scalable, and user-friendly platform for training and deploying AI-based atmospheric models on high-performance computing systems. It offers an end-to-end pipeline for data preprocessing, model training, and evaluation, democratizing access to advanced AI NWP capabilities. We demonstrate CREDIT's potential through WXFormer, a novel deterministic vision transformer designed to predict atmospheric states autoregressively, addressing common AI NWP issues like compounding error growth with techniques such as spectral normalization, padding, and multi-step training. Additionally, to illustrate CREDIT's flexibility and state-of-the-art model comparisons, we train the FUXI architecture within this framework. Our findings show that both FUXI and WXFormer, trained on six-hourly ERA5 hybrid sigma-pressure levels, generally outperform IFS HRES in 10-day forecasts, offering potential improvements in efficiency and forecast accuracy. CREDIT's modular design enables researchers to explore various models, datasets, and training configurations, fostering innovation within the scientific community.
Humans and other animals are able to quickly generalize latent dynamics of spatiotemporal sequences, often from a minimal number of previous experiences. Additionally, internal representations of external stimuli must remain stable, even in the presence of sensory noise, in order to be useful for informing behavior. In contrast, typical machine learning approaches require many thousands of samples, and generalize poorly to unexperienced examples, or fail completely to predict at long timescales. Here, we propose a novel neural network module which incorporates hierarchy and recurrent feedback terms, constituting a simplified model of neocortical microcircuits. This microcircuit predicts spatiotemporal trajectories at the input layer using a temporal error minimization algorithm. We show that this module is able to predict with higher accuracy into the future compared to traditional models. Investigating this model we find that successive predictive models learn representations which are increasingly removed from the raw sensory space, namely as successive temporal derivatives of the positional information. Next, we introduce a spiking neural network model which implements the rate-model through the use of a recently proposed biological learning rule utilizing dual-compartment neurons. We show that this network performs well on the same tasks as the mean-field models, by developing intrinsic dynamics that follow the dynamics of the external stimulus, while coordinating transmission of higher-order dynamics. Taken as a whole, these findings suggest that hierarchical temporal abstraction of sequences, rather than feed-forward reconstruction, may be responsible for the ability of neural systems to quickly adapt to novel situations.
Reliably quantifying uncertainty in precipitation forecasts remains a critical challenge. This work examines the application of a deep learning (DL) architecture, Unet, for postprocessing deterministic numerical weather predictions of precipitation to improve their skills and for deriving forecast uncertainty. Daily accumulated 0-4-day precipitation fore-casts are generated from a 34-yr reforecast based on the West Weather Research and Forecasting (West-WRF) mesoscale model, developed by the Center for Western Weather and Water Extremes. The Unet learns the distributional parameters associated with a censored, shifted gamma distribution. In addition, the DL framework is tested against state-of-the-art benchmark methods, including an analog ensemble, nonhomogeneous regression, and mixed-type meta-Gaussian distribu-tion. These methods are evaluated over four years of data and the western United States. The Unet outperforms the benchmark methods at all lead times as measured by continuous ranked probability and Brier skill scores. The Unet also produces a reliable estimation of forecast uncertainty, as measured by binned spread-skill relationship diagrams. Addition-ally, the Unet has the best performance for extreme events (i.e., the 95th and 99th percentiles of the distribution) and for these cases, its performance improves as more training data are available.SIGNIFICANCE STATEMENT: Accurate precipitation forecasts are critical for social and economic sectors. They also play an important role in our daily activity planning. The objective of this research is to investigate how to use a deep learning architecture to postprocess high-resolution (4 km) precipitation forecasts and generate accurate and reliable forecasts with quantified uncertainty. The proposed approach performs well with extreme cases and its per-formance improves as more data are available in training.
We develop and compare model-error representation schemes derived from data assimilation increments and nudging tendencies in multi-decadal simulations of the community atmosphere model, version 6. Each scheme applies a bias correction during simulation run-time to the zonal and meridional winds. We quantify to which extent such online adjustment schemes improve the model climatology and variability on daily to seasonal timescales. Generally, we observe a ca. 30% improvement to annual upper-level zonal winds, with largest improvements in boreal spring (ca. 35%) and winter (ca. 47%). Despite only adjusting the wind fields, we additionally observe a ca. 20% improvement to annual precipitation over land, with the largest improvements in boreal fall (ca. 36%) and winter (ca. 25%), and a ca. 50% improvement to annual sea level pressure, globally. With mean state adjustments alone, the dominant pattern of boreal low-frequency variability over the Atlantic (the North Atlantic Oscillation) is significantly improved. Additional stochasticity further increases the modal explained variances, which brings it closer to the observed value. A streamfunction tendency decomposition reveals that the improvement is due to an adjustment to the high- and low-frequency eddy-eddy interaction terms. In the Pacific, the mean state adjustment alone led to an erroneous deepening of the Aleutian low, but this was remedied with the addition of stochastically selected tendencies. Finally, from a practical standpoint, we discuss the performance of using data assimilation increments versus nudging tendencies for an online model-error representation.
The modeling of weather and climate has been a success story. The skill of forecasts continues to improve and model biases continue to decrease. Combining the output of multiple models has further improved forecast skill and reduced biases. But are we exploiting the full capacity of state-of-the-art models in making forecasts and projections? Supermodeling is a recent step forward in the multimodel ensemble approach. Instead of combining model output after the simulations are completed, in a supermodel individual models exchange state information as they run, influencing each other’s behavior. By learning the optimal parameters that determine how models influence each other based on past observations, model errors are reduced at an early stage before they propagate into larger scales and affect other regions and variables. The models synchronize on a common solution that through learning remains closer to the observed evolution. Effectively a new dynamical system has been created, a supermodel, that optimally combines the strengths of the constituent models. The supermodel approach has the potential to rapidly improve current state-of-the-art weather forecasts and climate predictions. In this paper we introduce supermodeling, demonstrate its potential in examples of various complexity, and discuss learning strategies. We conclude with a discussion of remaining challenges for a successful application of supermodeling in the context of state-of-the-art models. The supermodeling approach is not limited to the modeling of weather and climate, but can be applied to improve the prediction capabilities of any complex system, for which a set of different models exists.
Rats readily switch between foraging and more complex navigational behaviors such as pursuit of other rats or prey. These tasks require vastly different tracking of multiple behaviorally significant variables including self-motion state. To explore whether navigational context modulates self-motion tracking, we examined self-motion tuning in posterior parietal cortex neurons during foraging versus visual target pursuit. Animals performing the pursuit task demonstrate predictive processing of target trajectories by anticipating and intercepting them. Relative to foraging, pursuit yields multiplicative gain modulation of self-motion tuning and enhances self-motion state decoding. Self-motion sensitivity in parietal cortex neurons is, on average, history dependent regardless of behavioral context, but the temporal window of self-motion integration extends during target pursuit. Finally, many self-motion-sensitive neurons conjunctively track the visual target position relative to the animal. Thus, posterior parietal cortex functions to integrate the location of navigationally relevant target stimuli into an ongoing representation of past, present, and future locomotor trajectories.