Two or more spatio-temporally co-located meteorological/climatological extremes (co-occurring extremes) place far greater stress on human and ecological systems than any single extreme could. This was observed during the California drought of 2011–2015 where multiple years of negative precipitation anomalies occurred simultaneously with positive temperature anomalies resulting in California’s worst drought on observational record. The large-scale drivers which modulate the occurrence of extremes in two or more variables remains largely unexplored. Using California wintertime (November–April) temperature and precipitation as a case study, we apply a novel, nonparametric conditional probability distribution method that allows for evaluation of complex, multivariate, and nonlinear relationships that exist among temperature, precipitation, and various indicators of large-scale climate variability and change. We find that multivariate variability and statistics of temperature and precipitation exhibit strong spatial variation across scales that are often treated as being homogeneous. Further, we demonstrate that the multivariate statistics of temperature and precipitation are highly non-stationary and therefore require more robust and sophisticated statistical techniques for accurate characterization. Of all the indicators of the large-scale climate conditions we studied, the dipole index explains the greatest fraction of multivariate variability in the co-occurrence of California wintertime extremes in temperature and precipitation.
In a warming climate, where climate adaptation and mitigation strategies are increasingly critical, understanding changes in major global atmospheric moisture transport mechanisms, such as atmospheric rivers (ARs), is essential for assessing shifts in the hydrological cycle. This study examines the frequency and impacts of ARs under the shared socioeconomic pathway (SSP2-4.5) warming scenario and the stratospheric aerosol injection (SAI) scenario (ARISE-SAI-1.5). Our findings indicate that under SAI-1.5, ARs retreat from inland regions, and the occurrence of high-impact ARs (Category 3 and above) decreases, although some uncertainty remains regarding the response time of ARs to SAI. However, under future climate without SAI (SSP2-4.5), ARs penetrate further inland and there are higher numbers of high-impact ARs (> Category 3 ARs). In the northern hemisphere oceans, SAI-1.5 leads to a gradual increase in AR frequency compared to SSP2-4.5, whereas the southern hemisphere oceans exhibit the opposite trend. Also, extreme AR-associated precipitation is reduced under SAI-1.5 relative to SSP2-4.5, whereas beneficial precipitation is projected to increase under SAI-1.5. The contrasting responses associated with AR location and intensity highlights the need for further research to better understand the underlying drivers of AR changes before SAI can be considered in policy decisions affecting global moisture transport mechanisms.
Many atmospheric river detection tools (ARDTs) have been developed over the past few decades to identify atmospheric rivers (ARs). Different ARDTs have been observed to capture a variety of frequencies, shapes, and sizes of ARs. Due to this, questions have arisen about the underlying phenomena associated with the detected ARs: do all ARDTs detect the same meteorological phenomena? In this paper, we assess eight ARDTs and investigate the underlying synoptic scale phenomena during landfalling ARs along the west coast of North America. We find that during landfalling AR events, prevalent low‐pressure and high‐pressure systems converge and enhance moisture influx toward the landfalling site. We identify that all eight ARDTs identify AR conditions associated with baroclinic waves, with the region of intense integrated vapor transport (IVT) located downstream of the upper level (500 hPa) trough. The magnitude of IVT is enhanced by the strength of the pressure gradients in the confluence region. Although the ARDTs assessed agree on the general phenomena, there are however subtle differences in each ARDT per the clustering analysis we performed. We conclude that the eight ARDTs identify similar underpinning synoptic scale meteorological phenomena.
During the last four decades, global warming has statistically significant intensified extreme precipitation events in the Midwestern United States (defined here as the region covering Illinois, Indiana, Ohio, and Kentucky), leading to increased risks to human life, property, and infrastructure. To enable climate change adaptation and resilience across various economic and social sectors in this region, updated information about future climate changes, specifically at finer spatial scales, is essential. Leveraging a new 150-year dynamical downscaling data set at convection-permitting resolution, this study introduces a framework to construct the projected future intensity-duration-frequency (IDF) curves of heavy precipitation, which are prominent tools for infrastructure design and water resources management. This framework generates IDF curves at both sub-daily and multi-day duration utilizing hourly in situ observations as well as quantile-based statistical techniques in bias-correction and return levels selection. The assumption of non-stationarity in the distribution parameter fitting process is also implemented in this workflow. Compared to historical IDF curves for 1980-2022, future projected IDF curves for 2058-2100 under Representative Concentration Pathway (RCP) 4.5 and RCP 8.5 scenarios indicate an average intensity increase of approximately 15% and 25%, respectively, across 74 stations, considering both annual and seasonal timescales. Future projections suggest that extreme precipitation events may become more severe across six investigated return periods, with longer return periods showing a greater increase. The frequency of future extreme precipitation events in the Midwest region is also projected to double. Furthermore, current results reveal spatial heterogeneity of future trends across stations owing to the high-resolution input data set.
AI models are criticized as being black boxes, potentially subjecting climate science to greater uncertainty. Explainable artificial intelligence (XAI) has been proposed to probe AI models and increase trust. In this review and perspective paper, we suggest that, in addition to using XAI methods, AI researchers in climate science can learn from past successes in the development of physics-based dynamical climate models. Dynamical models are complex but have gained trust because their successes and failures can sometimes be attributed to specific components or sub-models, such as when model bias is explained by pointing to a particular parameterization. We propose three types of understanding as a basis to evaluate trust in dynamical and AI models alike: (1) instrumental understanding, which is obtained when a model has passed a functional test; (2) statistical understanding, obtained when researchers can make sense of the modeling results using statistical techniques to identify input–output relationships; and (3) component-level understanding, which refers to modelers' ability to point to specific model components or parts in the model architecture as the culprit for erratic model behaviors or as the crucial reason why the model functions well. We demonstrate how component-level understanding has been sought and achieved via climate model intercomparison projects over the past several decades. Such component-level understanding routinely leads to model improvements and may also serve as a template for thinking about AI-driven climate science. Currently, XAI methods can help explain the behaviors of AI models by focusing on the mapping between input and output, thereby increasing the statistical understanding of AI models. Yet, to further increase our understanding of AI models, we will have to build AI models that have interpretable components amenable to component-level understanding. We give recent examples from the AI climate science literature to highlight some recent, albeit limited, successes in achieving component-level understanding and thereby explaining model behavior. The merit of such interpretable AI models is that they serve as a stronger basis for trust in climate modeling and, by extension, downstream uses of climate model data.
We present a new atmospheric river (AR) analysis and benchmarking tool, namely Atmospheric River Metrics Package (ARMP). It includes a suite of new AR metrics that are designed for quick analysis of AR characteristics via statistics in gridded climate datasets such as model output and reanalysis. This package can be used for climate model evaluation in comparison with reanalysis and observational products. Integrated metrics such as mean bias and spatial pattern correlation are efficient for diagnosing systematic AR biases in climate models. For example, the package identifies the fact that, in CMIP5 and CMIP6 (Coupled Model Intercomparison Project Phases 5 and 6) models, AR tracks in the South Atlantic are positioned farther poleward compared to ERA5 reanalysis, while in the South Pacific, tracks are generally biased towards the Equator. For the landfalling AR peak season, we find that most climate models simulate a completely opposite seasonal cycle over western Africa. This tool can also be used for identifying and characterizing structural differences among different AR detectors (ARDTs). For example, ARs detected with the Mundhenk algorithm exhibit systematically larger size, width, and length compared to the TempestExtremes (TE) method. The AR metrics developed from this work can be routinely applied for model benchmarking and during the development cycle to trace performance evolution across model versions or generations and set objective targets for the improvement of models. They can also be used by operational centers to perform near-real-time climate and extreme event impact assessments as part of their forecast cycle.
Atmospheric rivers (ARs) are filamentary structures within the atmosphere that account for a substantial portion of poleward moisture transport and play an important role in Earth's hydroclimate. However, there is no one quantitative definition for what constitutes an atmospheric river, leading to uncertainty in quantifying how these systems respond to global change. This study seeks to better understand how different AR detection tools (ARDTs) respond to changes in climate states utilizing single-forcing climate model experiments under the aegis of the Atmospheric River Tracking Method Intercomparison Project (ARTMIP). We compare a simulation with an early Holocene orbital configuration and another with CO2 levels of the Last Glacial Maximum to a preindustrial control simulation to test how the ARDTs respond to changes in seasonality and mean climate state, respectively. We find good agreement among the algorithms in the AR response to the changing orbital configuration, with a poleward shift in AR frequency that tracks seasonal poleward shifts in atmospheric water vapor and zonal winds. In the low CO2 simulation, the algorithms generally agree on the sign of AR changes, but there is substantial spread in their magnitude, indicating that mean-state changes lead to larger uncertainty. This disagreement likely arises primarily from differences between algorithms in their thresholds for water vapor and its transport used for identifying ARs. These findings warrant caution in ARDT selection for paleoclimate and climate change studies in which there is a change to the mean climate state, as ARDT selection contributes substantial uncertainty in such cases.
Object‐based identification algorithms for atmospheric features are commonly utilized to attribute global precipitation. This study employs a systematic approach to examine feature co‐occurrences and their relationships to mean and extreme precipitation. Four features are identified using existing data sets for atmospheric rivers (ARs), mesoscale convective systems (MCSs), low‐pressure systems (LPSs), and fronts (FTs). Often, a single atmospheric phenomenon satisfies the criteria set by multiple feature identification algorithms, yielding an association between precipitation and multiple features. Over the extra‐tropics, the number of features attributed to a single event typically increases with precipitation intensity. Over two‐thirds of the precipitation is from co‐occurring features, with a considerable fraction related to AR‐FT co‐occurrences. Over the tropics, about one‐quarter of precipitation is associated with co‐occurring features, with LPS‐MCS co‐occurrences contributing substantially in monsoon regions. MCSs are the leading single‐feature contributors over tropical land and oceans. In the extra‐tropics, FTs, ARs, and their co‐occurrences account for over half of the total precipitation over oceans. AR‐FT‐MCS and FT‐MCS co‐occurrences contribute to extremes (precipitation exceeding the 95th percentile) over both oceans (over 30%) and land (over 20%). Any combination of features involving MCSs shows a larger contribution to high percentiles of precipitation intensity. A case analysis indicates that AR‐FT‐MCS co‐occurrences exhibit convective instability and deep vertical motion, suggesting that the feature trackers and reanalysis are capturing physics relevant to both convective and frontal systems. The results here emphasize the need for simultaneous identifications of multiple features when attributing precipitation to atmospheric phenomena.
The rapid rise of deep learning (DL) in numerical weather prediction (NWP) has led to a proliferation of models which forecast atmospheric variables with comparable or superior skill than traditional physics-based NWP. However, among these leading DL models, there is a wide variance in both the training settings and architecture used. Further, the lack of thorough ablation studies makes it hard to discern which components are most critical to success. In this work, we show that it is possible to attain high forecast skill even with relatively off-the-shelf architectures, simple training procedures, and moderate compute budgets. Specifically, we train a minimally modified SwinV2 transformer on ERA5 data, and find that it attains superior forecast skill when compared against IFS. We present some ablations on key aspects of the training pipeline, exploring different loss functions, model sizes and depths, and multi-step fine-tuning to investigate their effect. We also examine the model performance with metrics beyond the typical ACC and RMSE, and investigate how the performance scales with model size.
Abstract Atmospheric rivers (ARs) significantly impact the hydrological cycle and associated extremes in western continental regions. Recent studies suggest ARs also influence water resources and extremes in continental interiors. AR detection tools indicate that AR conditions are relatively frequent in areas east of the Rocky Mountains. The origin of these ARs, whether from synoptic‐scale waves or mesoscale processes, is unclear. This study uses meteorological composite maps and transects of AR conditions during the four seasons. The analysis reveals that ARs east of the Rockies are associated with long‐wave, baroclinic Rossby waves. This result demonstrates that eastern North American ARs are dynamically similar to their western coastal counterparts, though mechanisms for vertical moisture flux differ between the two. These findings provide a foundation for understanding future climate change and ARs in this region and offer new methods for evaluating climate model simulations.
Understanding the regional and temporal variability of atmospheric river (AR) seasonality is crucial for preparedness and mitigation of extreme events. Previously thought to peak mainly in winter, recent research reveals that ARs exhibit region-specific seasonality. However, AR analysis is heavily influenced by the chosen detection algorithm. Our study examines how AR seasonality varies based on both location, year and algorithm selection. We investigate the link between year-to-year consistency of peak AR activity and the presence of a dominant seasonal pattern. We categorize regions based on their year-to-year seasonality characteristics, including consistent patterns (e.g., India, Central Asia), patterns with occasional outliers (e.g., British Columbia coast, Gulf of Alaska), and regions lacking a clear dominant season of peak AR frequency (e.g., South Atlantic, parts of Australia). Hence, not all regions exhibit a consistent seasonal cycle of AR activity. Additionally, different algorithms may detect a consistent seasonal pattern for the same region but disagree on the exact dominant season. This is exemplified by the conflicting results obtained for China. Integrated Vapor Transport (IVT) often corroborates consistent or inconsistent patterns across regions. In conclusion, this study suggests that variations in the consistency of seasonal patterns are related not only to the detection technique but also to atmospheric circulation, synoptic and low-frequency anomalies. Understanding the variations in the consistency of seasonal pattern in areas like Britain remains challenging due to algorithmic and physical differences. These findings emphasize the need for a multi-faceted approach to AR research, considering not just detection methodologies but also regional characteristics and atmospheric processes. Understanding the specific reasons for inconsistent seasonal patterns is an important next step for future research to improve forecasts and preparedness.
Atmospheric rivers (ARs) are extreme weather events that can alleviate drought or cause billions of US dollars in flood damage. By transporting significant amounts of latent energy towards the poles, they are crucial to maintaining the climate system's energy balance. Since there is no first-principle definition of an AR grounded in geophysical fluid mechanics, AR identification is currently performed by a multitude of expert-defined, threshold-based algorithms. The variety of AR detection algorithms has introduced uncertainty into the study of ARs, and the thresholds of the algorithms may not generalize to new climate datasets and resolutions. We train convolutional neural networks (CNNs) to detect ARs while representing this uncertainty; we name these models ARCNNs. To detect ARs without requiring new labeled data and labor-intensive AR detection campaigns, we present a semi-supervised learning framework based on image style transfer. This framework generalizes ARCNNs across climate datasets and input fields. Using idealized and realistic numerical models, together with observations, we assess the performance of the ARCNNs. We test the ARCNNs in an idealized simulation of a shallow-water fluid in which nearly all the tracer transport can be attributed to AR-like filamentary structures. In reanalysis and a high-resolution climate model, we use ARCNNs to calculate the contribution of ARs to meridional latent heat transport, and we demonstrate that this quantity varies considerably due to AR detection uncertainty.
Many atmospheric river detectors (ARDTs) have been developed over the past few decades to capture atmospheric rivers (ARs). However, different ARDTs have been observed to capture different frequencies, shapes and sizes of ARs. Due to this, many questions including investigating the underlying phenomena for ARs in the ARDTs have been posed. In this paper, we assess four different ARDTs and investigate the underlying meteorological phenomena during landfalling ARs. We find that during landfalling ARs events, there exists a prevalent low-pressure and high-pressure confluence that enhances moisture influx toward the landfalling site. The strength of the pressure gradient in the confluence region enhances the influx of the integrated vapor transport. The four ARDTs predominantly capture similar atmospheric processes, nonetheless, they have statistically different magnitudes.
A comprehensive understanding of human-induced changes to rainfall is essential for water resource management and infrastructure design. However, at regional scales, existing detection and attribution studies are rarely able to conclusively identify human influence on precipitation. Here we show that anthropogenic aerosol and greenhouse gas (GHG) emissions are the primary drivers of precipitation change over the United States. GHG emissions increase mean and extreme precipitation from rain gauge measurements across all seasons, while the decadal-scale effect of global aerosol emissions decreases precipitation. Local aerosol emissions further offset GHG increases in the winter and spring but enhance rainfall during the summer and fall. Our results show that the conflicting literature on historical precipitation trends can be explained by offsetting aerosol and greenhouse gas signals. At the scale of the United States, individual climate models reproduce observed changes but cannot confidently determine whether a given anthropogenic agent has increased or decreased rainfall.
Abstract. AI models are criticized as being black boxes, potentially subjecting climate science to greater uncertainty. Explainable artificial intelligence (XAI) has been proposed to probe AI models and increase trust. In this Perspective, we suggest that, in addition to using XAI methods, AI researchers in climate science can learn from past successes in the development of physics-based dynamical climate models. Dynamical models are complex but have gained trust because their successes and failures can be attributed to specific components or sub-models, such as when model bias is explained by pointing to a particular parameterization. We propose three types of understanding as a basis to evaluate trust in dynamical and AI models alike: (1) instrumental understanding, which is obtained when a model has passed a functional test; (2) statistical understanding, which is obtained when researchers can make sense of the modelling results using statistical techniques to identify input-output relationships; and (3) Component-level understanding, which refers to modelers’ ability to point to specific model components or parts in the model architecture as the culprit for erratic model behaviors or as the crucial reason why the model functions well. We demonstrate how component-level understanding has been sought and achieved via climate model intercomparison projects over the past several decades. Such component-level of understanding routinely leads to model improvements and may also serve as a template for thinking about AI-driven climate science. Currently, XAI methods can help explain the behaviors of AI models by focusing on the mapping between input and output, thereby increasing the statistical understanding of AI models. Yet, to further increase our understanding of AI models, we will have to build AI models that have interpretable components amenable to component-level understanding. We give recent examples from the AI climate science literature to highlight some recent, albeit limited, successes in achieving component-level understanding and thereby explaining model behaviour. The merit of such interpretable AI models is that they serve as a stronger basis for trust in climate modeling and, by extension, downstream uses of climate model data.
Abstract. We present a suite of new atmospheric river (AR) metrics that are designed for quick analysis of AR characteristics and statistics in gridded climate datasets such as model output and reanalysis. This package is expected to be particularly useful for climate model evaluation. The metrics include mean bias and spatial pattern correlation, which are efficient for diagnosing systematic AR biases in climate models. For example, the package identifies that in CMIP5 and CMIP6 models, AR tracks in the south Atlantic are positioned farther poleward compared to the ERA5 reanalysis, while in the south Pacific, tracks are generally biased towards the equator. For the landfalling AR peak season, we find that most climate models simulate a completely opposite seasonal cycle over western Africa. This tool is also useful for identifying and characterizing structural differences among different AR detectors (ARDTs). For example, ARs detected with the Mundhenk algorithm exhibit systematically larger size, width and length compared to the TempestExtremes (TE) method. The AR metrics developed from this work can be routinely applied for model benchmarking and during the development cycle to trace performance evolution across model versions or generations and set objective targets for the improvement of models. They can also be used by operational centers to perform near real-time climate and extreme events impact assessment as part of their forecast cycle.
Machine learning (ML) is a revolutionary technology with demonstrable applications across multiple disciplines. Within the Earth science community, ML has been most visible for weather forecasting, producing forecasts that rival modern physics-based models. Given the importance of deepening our understanding and improving predictions of the Earth system on all time scales, efforts are now underway to develop forecasting models into Earth-system models (ESMs), capable of representing all components of the coupled Earth system (or their aggregated behavior) and their response to external changes. Modeling the Earth system is a much more difficult problem than weather forecasting, not least because the model must represent the alternate (e.g., future) coupled states of the system for which there are no historical observations. Given that the physical principles that enable predictions about the response of the Earth system are often not explicitly coded in these ML-based models, demonstrating the credibility of ML-based ESMs thus requires us to build evidence of their consistency with the physical system. To this end, this paper puts forward five recommendations to enhance comprehensive, standardized, and independent evaluation of ML-based ESMs to strengthen their credibility and promote their wider use.
In late June, 2021, a devastating heatwave affected the US Pacific Northwest and western Canada, breaking numerous all-time temperature records by large margins and directly causing hundreds of fatalities. The observed 2021 daily maximum temperature across much of the U.S. Pacific Northwest exceeded upper bound estimates obtained from single-station temperature records even after accounting for anthropogenic climate change, meaning that the event could not have been predicted under standard univariate extreme value analysis assumptions. In this work, we utilize a flexible spatial extremes model that considers all stations across the Pacific Northwest domain and accounts for the fact that many stations simultaneously experience extreme temperatures. Our analysis incorporates the effects of anthropogenic forcing and natural climate variability in order to better characterize time-varying changes in the distribution of daily temperature extremes. We show that greenhouse gas forcing, drought conditions and large-scale atmospheric modes of variability all have significant impact on summertime maximum temperatures in this region. Our model represents a significant improvement over corresponding single-station analysis, and our posterior medians of the upper bounds are able to anticipate more than 96% of the observed 2021 high station temperatures after properly accounting for extremal dependence. Supplementary materials accompanying this paper appear online.
Abstract Although the 2021 Western North America (WNA) heat wave was predicted by weather forecast models, questions remain about whether such strong events can be simulated by global climate models (GCMs) at different model resolutions. Here, we analyze sets of GCM simulations including historical and future periods to check for the occurrence of similar events. High‐ and low‐resolution simulations both encounter challenges in reproducing events as extreme as the observed one, particularly under the present climate. Relatively stronger amplitudes are observed during the future periods. Furthermore, high‐ and low‐resolution short initialized GCM simulations are both able to reasonably predict such strong events and their associated high‐pressure ridge over the WNA with a 1 week forecast lead time. Moisture sensitivity experiments further indicate a drier atmospheric moisture condition results in substantially higher near‐surface temperatures in the simulated heat events.