Abstract We assess the skill of daily machine learning (ML)-based probabilistic convective hazard predictions generated from two experimental medium-range (i.e., extending to day 8) convection-allowing ensemble (CAE) systems to determine if such predictions are skillful over the CONUS. CAE forecasts were generated in real time for events in spring 2023 and 2024, and convective hazard predictions were produced using neural networks, trained with CAE output and observed storm reports (i.e., reports of tornadoes, hail, and convectively induced wind gusts). The skill of the two ML-based hazard forecasts was evaluated relative to climatology, CAE-based updraft helicity (UH) forecasts, and ML-based predictions generated from the NOAA operational Global Ensemble Forecast System (GEFS). The CAE ML-based hazard forecasts outperformed both the UH-based forecasts and climatology through day 8. When evaluated against the GEFS ML, CAE-based hazard forecasts were skillful for days 1–4, beyond which differences were small and not statistically significant. All hazard forecasts were more skillful in spring 2024 than in spring 2023, likely due to stronger large-scale forcing that enhanced predictability. The stronger forcing led to reduced benefit of the CAE-based forecasts relative to the GEFS, with benefits extending to days 3–4 in 2024 but extending through day 5 in 2023. Finally, ML interpretability experiments revealed that CAE storm-scale predictors were weighted more than environmental predictors for days 1–2, while the opposite was true for days 3–5. Together, these results document the skill and value of CAEs for ML-based hazard prediction, as well as interannual skill variations due to different weather regimes that can influence skill differences and NWP system intercomparisons. Significance Statement Forecasters often rely on numerical models to provide guidance when making predictions of thunderstorms and their hazards, such as tornadoes. Typically, the most advanced of these models only provide forecasts a day or two into the future. Here, we examine numerical model forecasts that extend further, up to 8 days into the future, and assess whether these new types of thunderstorm hazard forecasts can provide beneficial guidance at these longer lead times.
Improving the skill of medium-range (3-8-day) severe weather prediction is crucial for mitigating societal impacts. This study introduces a novel approach leveraging decoder-only transformer networks to postprocess artificial intelligence (AI)-based weather forecasts, specifically from the Pangu-Weather model, for improved severe weather guidance. Unlike traditional postprocessing methods that use a dense neural network to predict the probability of severe weather using discrete forecast samples, our method treats forecast lead times as sequential "tokens," enabling the transformer to learn complex temporal relationships within the evolving atmospheric state. We compare this approach against postprocessing of the Global Forecast System (GFS) using both a traditional dense neural network and our transformer, as well as configurations that exclude convective parameters to fairly evaluate the impact of using the Pangu-Weather AI model. Results demonstrate that the transformer-based postprocessing significantly enhances forecast skill compared to dense neural networks. Furthermore, AI-driven forecasts, particularly Pangu-Weather initialized from high-resolution analysis, exhibit superior performance to GFS in the medium range, even without explicit convective parameters. Our approach offers improved accuracy, and reliability, which also provides interpretability through feature attribution analysis, advancing medium-range severe weather prediction capabilities.
An ensemble post-processing method is developed to improve the probabilistic forecasts of extreme precipitation events across the conterminous United States (CONUS). The method combines a 3-D Vision Transformer (ViT) for bias correction with a Latent Diffusion Model (LDM), a generative Artificial Intelligence (AI) method, to post-process 6-hourly precipitation ensemble forecasts and produce an enlarged generative ensemble that contains spatiotemporally consistent precipitation trajectories. These trajectories are expected to improve the characterization of extreme precipitation events and offer skillful multi-day accumulated and 6-hourly precipitation guidance. The method is tested using the Global Ensemble Forecast System (GEFS) precipitation forecasts out to day 6 and is verified against the Climate-Calibrated Precipitation Analysis (CCPA) data. Verification results indicate that the method generated skillful ensemble members with improved Continuous Ranked Probabilistic Skill Scores (CRPSSs) and Brier Skill Scores (BSSs) over the raw operational GEFS and a multivariate statistical post-processing baseline. It showed skillful and reliable probabilities for events at extreme precipitation thresholds. Explainability studies were further conducted, which revealed the decision-making process of the method and confirmed its effectiveness on ensemble member generation. This work introduces a novel, generative-AI-based approach to address the limitation of small numerical ensembles and the need for larger ensembles to identify extreme precipitation events.
An ensemble post-processing method is developed for the probabilistic prediction of severe weather (tornadoes, hail, and wind gusts) over the conterminous United States (CONUS). The method combines conditional generative adversarial networks (CGANs), a type of deep generative model, with a convolutional neural network (CNN) to post-process convection-allowing model (CAM) forecasts. The CGANs are designed to create synthetic ensemble members from deterministic CAM forecasts, and their outputs are processed by the CNN to estimate the probability of severe weather. The method is tested using High-Resolution Rapid Refresh (HRRR) 1--24 hr forecasts as inputs and Storm Prediction Center (SPC) severe weather reports as targets. The method produced skillful predictions with up to 20% Brier Skill Score (BSS) increases compared to other neural-network-based reference methods using a testing dataset of HRRR forecasts in 2021. For the evaluation of uncertainty quantification, the method is overconfident but produces meaningful ensemble spreads that can distinguish good and bad forecasts. The quality of CGAN outputs is also evaluated. Results show that the CGAN outputs behave similarly to a numerical ensemble; they preserved the inter-variable correlations and the contribution of influential predictors as in the original HRRR forecasts. This work provides a novel approach to post-process CAM output using neural networks that can be applied to severe weather prediction.
As artificial intelligence (AI) methods are increasingly used to develop new guidance intended for operational use by forecasters, it is critical to evaluate whether forecasters deem the guidance trustworthy. Past trust-related AI research suggests that certain attributes (e.g., understanding how the AI was trained, interactivity, performance) contribute to users perceiving the AI as trustworthy. However, little research has been done to examine the role of these and other attributes for weather forecasters. In this study, we conducted 16 online interviews with National Weather Service (NWS) forecasters to examine (a) how they make guidance use decisions, and (b) how the AI model technique used, training, input variables, performance, and developers as well as interacting with the model output influenced their assessments of trustworthiness of new guidance. The interviews pertained to either a random forest model predicting probability of severe hail or a 2D-convolutional neural net model predicting probability of storm mode. When taken as a whole, our findings illustrate how forecasters’ assessment of AI guidance trustworthiness is a process that occurs over time rather than automatically or at first introduction. We recommend developers center end users when creating new AI guidance tools, making end users integral to their thinking and efforts. This approach is essential for the development of useful and used tools. The details of these findings can help AI developers understand how forecasters perceive AI guidance and inform AI development and refinement efforts.
While convective storm mode is explicitly depicted in convection-allowing model (CAM) output, subjectively diagnosing mode in large volumes of CAM forecasts can be burdensome. In this work, four machine learning (ML) models were trained to probabilistically classify CAM storms into one of three modes: supercells, quasi-linear convective systems, and disorganized convection. The four ML models included a dense neural network (DNN), logistic regression CNN, and LR were trained with a set of hand-labeled CAM storms, while the semisupervised GMM used updraft helicity and storm size to generate clusters, which were then hand labeled. When evaluated using storms withheld from training, the four classifiers had similar ability to discriminate between modes, but the GMM had worse calibration. The DNN and LR had similar objective performance to the CNN, suggesting that CNN-based methods may not be needed for mode classification tasks. The mode classifications from all four classifiers successfully approximated the known climatology of modes in the United States, including a maximum in supercell occurrence in the U.S. Central Plains. Further, the modes also occurred in environments recognized to support the three different storm morphologies. Finally, storm mode provided useful information about hazard type, e.g., storm reports were most likely with supercells, further supporting the efficacy of the classifiers. Future applications, including the use of objective CAM mode classifications as a novel predictor in ML systems, could potentially lead to improved forecasts of convective hazards.
Herein, 14 severe quasi-linear convective systems (QLCS) covering a wide range of geographical locations and environmental conditions are simulated for both 1- and 3-km horizontal grid resolutions, to further clarify their comparative capabilities in representing convective system features associated with severe weather production. Emphasis is placed on validating the simulated reflectivity structures, cold pool strength, mesoscale vortex characteristics, and surface wind strength. As to the overall reflectivity characteristics, the basic leading-line trailing stratiform structure was often better defined at 1 versus 3 km, but both resolutions were capable of producing bow echo and line echo wave pattern type features. Cold pool characteristics for both the 1- and 3-km simulations were also well replicated for the differing environments, with the 1-km cold pools slightly colder and often a bit larger. Both resolutions captured the larger mesoscale vortices, such as line-end or bookend vortices, but smaller, leading-line mesoscale updraft vortices, that often promote QLCS tornadogenesis, were largely absent in the 3-km simulations. Finally, while maximum surface winds were only marginally well predicted for both resolutions, the simulations were able to reasonably differentiate the relative contributions of the cold pool versus mesoscale vortices. The present results suggest that while many QLCS characteristics can be reasonably represented at a grid scale of 3 km, some of the more detailed structures, such as overall reflectivity characteristics and the smaller leading-line mesoscale vortices would likely benefit from the finer 1-km grid spacing. Significance Statement High-resolution model forecasts using 3-km grid spacing have proven to offer significant forecast guidance enhancements for severe convective weather. However, it is unclear whether additional enhancements can be obtained by decreasing grid spacings further to 1 km. Herein, we compare forecasts of severe quasi-linear convective systems (QLCS) simulated using 1- versus 3-km grids to document the potential value added of such increases in grid resolutions. It is shown that some significant improvements can be obtained in the representation of many QLCS features, especially as regards reflectivity structure and in the development of small, leading-line mesoscale vortices that can contribute to both severe surface wind and tornado production.
Thunderstorm mode strongly impacts the likelihood and predictability of tornadoes and other hazards, and thus is of great interest to severe weather forecasters and researchers. It is often impossible for a forecaster to manually classify all the storms within convection-allowing model (CAM) output during a severe weather outbreak, or for a scientist to manually classify all storms in a large CAM or radar dataset in a timely manner. Automated storm classification techniques facilitate these tasks and provide objective inputs to operational tools, including machine learning models for predicting thunderstorm hazards. Accurate storm classification, however, requires accurate storm segmentation. Many storm segmentation techniques fail to distinguish between clustered storms, thereby missing intense cells, or to identify cells embedded within quasi-linear convective systems that can produce tornadoes and damaging winds. Therefore, we have developed an iterative technique that identifies these constituent storms in addition to traditionally identified storms. Identified storms are classified according to a seven-mode scheme designed for severe weather operations and research. The classification model is a hand-developed decision tree that operates on storm properties computed from composite reflectivity and midlevel rotation fields. These properties include geometrical attributes, whether the storm contains smaller storms or resides within a larger-scale complex, and whether strong rotation exists near the storm centroid. We evaluate the classification algorithm using expert labels of 400 storms simulated by the NSSL Warn-on-Forecast System or analyzed by the NSSL Multi-Radar/Multi-Sensor product suite. The classification algorithm emulates expert opinion reasonably well (e.g., 76% accuracy for supercells), and therefore could facilitate a wide range of operational and research applications.
Uncrewed aircraft system (UAS) observations collected during the 2018 Lower Atmospheric Process Studies at Elevation-a Remotely Piloted Aircraft Team Experiment (LAPSE-RATE) field campaign were assimilated into a high-resolution configuration of the Weather Research and Forecasting Model using an ensemble Kalman filter. The benefit of UAS observations was assessed for a terrain-driven (drainage and upvalley) flow event that occurred within Colorado's San Luis Valley (SLV) using independent observations. The analysis and prediction of the strength, depth, and horizontal extent of drainage flow from the Saguache Canyon and the subsequent transition to upvalley and up-canyon flow were improved relative to that obtained both without data assimilation (benchmark) and when only surface observations were assimilated. Assimilation of UAS observations greatly improved the analyses of vertical variations in temperature, relative humidity, and winds at multiple locations in the northern portion of the SLV, with reductions in both bias and the root-mean-square error of roughly 40% for each variable relative to the benchmark run. Despite these noted improvements, some biases remain that were tied to measurement error and/or the impact of the boundary layer parameterization on vertically spreading the observations, both of which require further exploration. The results presented here highlight how observations obtained with a fleet of profiling UAS improve limited-area, high-resolution analyses and short-term forecasts in complex terrain.
A fifty-member convection allowing ensemble was used to examine environmental factors influencing afternoon convection initiation (CI) and subsequent severe weather on 5 April 2017 during Intensive Observing Period (IOP) 3b of the Verification of Rotation in Tornadoes Experiment in the Southeast (VORTEX-SE). This case produced several weak tornadoes (rated EF1 or less), and numerous reports of significant hail (diameter ≥ 2 inches), ahead of an eastward-moving surface cold front over eastern Alabama and southern Tennessee. Both observed and simulated CI was facilitated by mesoscale lower-tropospheric ascent maximized several tens of km ahead of the cold-frontal position, and the simulated mesoscale ascent was linked to surface frontogenesis in the ensemble mean. Simulated maximum 2-5-km AGL updraft helicity (UHmax) was used as a proxy for severe-weather producing mesocyclones, and considerable variability in UHmax occurred among the ensemble members. Ensemble members with UHmax > 100 m2 s-2 had stronger mesoscale ascent than in members with UHmax < 75 m2 s-2, which facilitated more timely CI by producing greater adiabatic cooling and moisture increases above the PBL. After CI, storms in the larger UHmax members moved northeastward toward a mesoscale region with larger convective available potential energy (CAPE) than in smaller UHmax members. The CAPE differences among members was influenced by differences in location of an antecedent mesoscale convective system, which had a thermodynamically stabilizing influence on the environment toward which storms were moving. Despite providing good overall guidance, the model ensemble overpredicted severe weather likelihoods in northeastern Alabama, where comparisons with VORTEX-SE soundings revealed a positive CAPE bias.
Abrupt changes in wind direction and speed can dramatically impact wildfire development and spread, endangering firefighters. A frequent cause of such wind shifts is outflow from thunderstorms and organised convective systems; thus, their identification and prediction present critical challenges for fire weather forecasters. Here, we develop a methodology and implement it in a software tool that can identify and depict convective outflow boundaries in high-resolution numerical weather prediction (NWP) models to provide guidance for fire weather forecasting. The tool can process model output, objectively identify gust fronts, and graphically display detected gust fronts and similar boundaries in NWP model forecasts. The tool is demonstrated with output from the Weather Research and Forecasting (WRF) model from the operational High-Resolution Rapid Refresh (HRRR) forecasting system and from a WRF ensemble run at the National Center for Atmospheric Research that can provide probabilistic information about model-predicted gust fronts. The tool can identify outflow boundaries in model forecasts of convective events occurring in both simple and complex terrain, both with and without concurrent wildfire activity. With accurate underlying model forecast output, the tool can reliably reveal areas of potential gust front activity and thus provide valuable guidance to incident meteorologists and command personnel.
Five sets of 48-h, 10-member, convection-allowing ensemble (CAE) forecasts with 3-km horizontal grid spacing were systematically evaluated over the conterminous United States with a focus on precipitation across 31 cases. The various CAEs solely differed by their initial condition perturbations (ICPs) and central initial states. CAEs initially centered about deterministic Global Forecast System (GFS) analyses were unequivocally better than those initially centered about ensemble mean analyses produced by a limited-area single-physics, single-dynamics 15-km continuously cycling ensemble Kalman filter (EnKF), strongly suggesting relative superiority of the GFS analyses. Additionally, CAEs with flow-dependent ICPs derived from either the EnKF or multimodel 3-h forecasts from the Short-Range Ensemble Forecast (SREF) system had higher fractions skill scores than CAEs with randomly generated mesoscale ICPs. Conversely, due to insufficient spread, CAEs with EnKF ICPs had worse reliability, discrimination, and dispersion than those with random and SREF ICPs. However, members in theCAEwith SREF ICPs undesirably clustered by dynamic core represented in the ICPs, and CAEs with random ICPs had poor spinup characteristics. Collectively, these results indicate that continuously cycled EnKF mean analyses were suboptimal for CAE initialization purposes and suggest that further work to improve limited-area continuously cycling EnKFs over large regional domains is warranted. Additionally, the deleterious aspects of using both multimodel and random ICPs suggest efforts toward improving spread in CAEs with single-physics, single-dynamics, flow-dependent ICPs should continue.
Valliappa Lakshmanan合作论文数University of Oklahoma2