As of January 2025, the United States has gone over 11 years without the occurrence of a tornado rated enhanced Fujita 5 scale (EF5) on the EF scale, constituting the longest "drought" in F5/EF5-rated tornadoes since the beginning of official records (1950). This article places the drought of 5-rated tornadoes in the context of a long-term tornado climatology. A key breakpoint exists between how the legacy F scale and the EF scale handle the complete destruction and sweeping away of single-family homes, with standard "well-constructed" homes being swept away constituting F5 damage on the F scale but only EF4 damage on the EF scale. To illustrate this point, adjusting the lower bound of EF5 on the EF scale from 201 to 190 mph or increasing all from 190-200-mph EF4s to >200-mph EF5s to account for this breakpoint in the handling of single-family homes would lead to consistent 5-level rating assignments from 1880 to the present day. Furthermore, contextual evidence that was used to aid in identifying 5-level damage in the F-scale and early EF-scale eras could assist in identifying top-tier intensity tornadoes. These find-ings ultimately lead to questions regarding what the highest possible rating of a tornado should represent from both physical and societal perspectives. SIGNIFICANCE STATEMENT: This article explores the lack of EF5 tornadoes in the past decade and methods that could be useful in discriminating top-tier tornado intensities. We show evidence that the lack of EF5-rated tornadoes in the past decade is less due to a weakening of tornadoes and likely attributable to stricter application of the enhanced Fujita scale. An 11-mph downward adjustment in the threshold for EF5 estimated wind speed would make the rate of EF5 ratings since 2013 consistent with a long-time climatology back to 1880. We contemplate the implica-tions of these findings on the tornado climatology and how the meteorological and engineering communities really desire to identify top-tier-intensity tornadoes.
The World Meteorological Organization (WMO) has called for more meaningful warnings to help reduce the impacts of weather-related events. Impact-based forecasts and warnings (IBFW) are being developed by forecasting agencies globally to meet this call. However, there are many challenges facing those implementing such systems. The WMO World Weather Research Programme High Impact Weather project sought to understand the future direction of research on IBFW systems. This research involved a virtual workshop series in late 2022 with over 350 international registrants to identify and analyse challenges that people are facing in developing IBFW systems, and potential solutions.We found that challenges relate to ten themes, in addition to defining the measures of success of an IBFW system Examples of key research gaps are to develop evaluation methods to explore the value of multi-hazard IBFW, in terms of collating data at appropriate scales, and including avoided losses, behavioural responses, and unconventional observations. We need to explore the value of using quantitative approaches in comparison to more efficient qualitative approaches, as well as of dynamic exposure and vulnerability data sets, and tailored warnings. We must investigate how to effectively communicate uncertainty and explore the governance of underpinning data.Further research on these topics will assist with the successful implementation of more meaningful warnings globally, whilst considering the feasibility and effectiveness of the efforts involved. This is our contribution to reducing the impacts of future hazards, at a time where climate-related events are expected to increase in severity.
Risks associated with rare events that occur infrequently but have serious impacts, such as tornadoes, are difficult to communicate. The U.S. National Weather Service Storm Prediction Center (SPC) communicates tornado risks by forecasting the absolute likelihood of tornadoes within 25 mi of a point in their regularly issued convective outlooks. The most common forecast likelihood of tornadoes in these outlooks is subjectively low, at 2% and 5%. Studies of probabilistic risk communication for natural disasters have suggested that normalizing the absolute likelihood of rare events by their baseline rate of occurrence can help users better understand small absolute changes in their risk of negative impacts. This study seeks to develop and investigate the distribution of relative risk, defined as the absolute likelihood divided by the climatological likelihood for tornadoes within 25 mi of a point, across the contiguous United States using the 1950-2021 SPC tornado report database. The analysis reveals that relative risk values vary greatly across time and space, primarily due to the annual and regional climatology of tornado events. Overall, relative risk may be able to provide useful context for tornado risk communication in areas that have higher rates of tornado occurrence, but it can be greatly inflated in regions with low climatological risks. Future work should seek to understand how broadcast meteorologists, emergency managers, and members of the public interact with relative risk information and in doing so identify the types of tornado events where relative risk improves or complicates risk messaging. SIGNIFICANCE STATEMENT: Forecasts for rare events like tornadoes are difficult to communicate to weather messaging recipients because of their very low forecast likelihoods. However, risk communication literature suggests that relative risk, which compares the forecast probability of a hazard to how likely a hazard is at a given time, can add important context to risk messages. By calculating relative risk for all reported tornadoes from 1950 to 2021, we observe that relative risk values are highest in areas that infrequently see tornadoes, and extremely large values of relative risk can occur when a single tornado occurs at a time of year and place where no others are recorded. Future work should identify how individuals react to relative risk values in tornado forecasts.
Correlation was examined between detrended monthly surface temperature and monthly [E]F-1+ tornadoes and tornado days for several contiguous US regions during the period 1954-2022. This relatively simple, yet robust, analysis indicated that regional temperature fluctuations are moderately-to-strongly correlated with tornado days during some months and in certain regions. In general, surface temperatures during boreal cool (warm) season had a positive (negative) correlation with tornado days. Implications for using a continuous, simple scalar variable such as surface temperature for tornado prediction are discussed, as well as the potential utility for understanding changes in tornado frequency due to climate variability and change.
In this work, we use 8 years (2014-21) of Operational Programme for the Exchange of Weather Radar Information (OPERA) radar data, ESWD severe weather reports, and arrival time difference (ATD) lightning detection network (ATDnet) data to create a climatology of quasi-linear convective systems (QLCSs) across Europe. In the first step, 15-min radar scans were used to identify 1475 QLCS polygons. Severe weather reports, lightning data, and morphological properties were used to classify QLCSs according to their intensity into 1151 marginal (78.0%), 272 moderate (18.5%), and 52 derecho (3.5%) events. The manual evaluation led to the recognition of QLCS morphological and precipitation archetypes, areal extent, duration, speed, forward motion, width, length, accompanying hazards, injuries, and fatalities. Results indicate that QLCSs are the most frequent during summer in central Europe, while in southern Europe, their occurrence is extended to late autumn. A bow echo feature occurred in around 29% of QLCS cases, while a mesoscale convective vortex occurred in almost 9%. Among precipitation modes, trailing and embedded stratiform types accounted for around 50% of QLCSs. The most frequent hazard accompanying QLCSs was lightning (taking up on average 94.4% of the area impacted by QLCS), followed by severe wind gusts (7.9%), excessive precipitation (6.1%), large hail (2.9%), and tornadoes (0.5%). Derechos had the largest coverage of severe wind reports (49.8%), while back-building QLCSs were the most prone to excessive precipitation events (13.5%). QLCSs caused 104 fatalities and 886 injuries. Severe wind gusts were responsible for 87.6% of fatalities and 73.6% of injuries. Nearly half of all fatalities and injuries were associated with only the 10 most impactful QLCS events, mostly warm-season derechos.
The National Weather Service plays a critical role in alerting the public when dangerous weather occurs. Tornado warnings are one of the most publicly visible products the NWS issues given the large societal impacts tor-nadoes can have. Understanding the performance of these warnings is crucial for providing adequate warning during tornadic events and improving overall warning performance. This study aims to understand warning performance during the lifetimes of individual storms (specifically in terms of probability of detection and lead time). For example, does probability of detection vary based on if the tornado was the first produced by the storm, or the last? We use tornado outbreak data from 2008 to 2014, archived NEXRAD radar data, and the NWS verification database to associate each tornado report with a storm object. This approach allows for an analysis of warning performance based on the chronological order of tornado occurrence within each storm. Results show that the probability of detection and lead time increase with later tornadoes in the storm; the first tornadoes of each storm are less likely to be warned and on average have less lead time. Probability of detection also decreases overnight, especially for first tornadoes and storms that only produce one tornado. These results are important for understanding how tornado warning performance varies during individual storm life cycles and how upstream forecast products (e.g., Storm Prediction Center tornado watches, mesoscale discussions, etc.) may increase warning confidence for the first tornado produced by each storm.
<p>In 1963, Fred Sanders wrote &#8220;it is urged the probability be acknowledged as the proper internal language of forecasters.&#8221; As such, we can consider all forecasts to be based on probabilities. Nevertheless, as Sanders discussed, issuance of probabilistic forecasts to users is not completely clear. In many cases, users may prefer a categorical forecast and, as such, the internal language of probabilities must be translated into an external language of categorical forecasts. Although there are simple ways to threshold probabilities of a dichotomous (yes/no) event into a yes/no forecast, doing so in a way that is consistent between the two expressions is not always easy. This becomes even more complex when not all users have the same decision threshold and, as a result, may have different preferences for where the optimal threshold is set.</p> <p>A typical example of this problem is the area of tornado forecasting, especially for short range (<1 hour) forecasts of events, such as is the case for tornado warnings in the United States. Here, I develop and explore some simple (&#8220;toy&#8221;) models of probability distributions for the forecast and occurrence of tornadoes, as well as the user decision problems. Since each of the three models have one (occurrence given a forecast) or two (forecast probability and distribution of user decisions) parameters, it is relatively easy to construct probabilistic models that can be thresholded to mimic tornado warning performance over time in the United States. In addition, the apparent global benefits (or losses) can be estimated for different probabilistic forecast distributions related to different user decisions compared to dichotomous forecasts.</p> <p>Although this methodology could not directly produce tornado warnings in an operational setting, it provides some limits on how a future system based on probabilities could be constrained. The process allows for the development of thresholded forecasts that are consistent with the underlying probabilistic information.</p>
Many tornadoes are unreported because of lack of observers or are underrated in intensity, width, or track length because of lack of damage indicators. These reporting biases substantially degrade estimates of tornado frequency and thereby undermine important endeavors such as studies of climate impacts on tornadoes and cost-benefit analyses of tornado damage mitigation. Building on previous studies, we use a Bayesian hierarchical modeling framework to estimate and correct for tornado reporting biases over the central United States during 1975-2018. The reporting biases are treated as a univariate function of population density. We assess how these biases vary with tornado intensity, width, and track length and over the analysis period. We find that the frequencies of tornadoes of all kinds, but especially stronger or wider tornadoes, have been substantially underestimated. Most strikingly, the Bayesian model estimates that there have been approximately 3 times as many tornadoes capable of (E)F2+ damage as have been recorded as (E)F2+ [(E)F indicates a rating on the (enhanced) Fujita scale]. The model estimates that total tornado frequency changed little over the analysis period. Statistically significant trends in frequency are found for tornadoes within certain ranges of intensity, pathlength, and width, but it is unclear what proportion of these trends arise from changes in damage survey practices. Simple analyses of the tornado database corroborate many of the inferences from the Bayesian model. Significance StatementPrior studies have shown that the probabilities of a tornado being reported and of its intensity, track length, and width being accurately estimated are strongly correlated with the local population density. We have developed a sophisticated statistical model that accounts for these population-dependent tornado reporting biases to improve estimates of tornado frequency in the central United States. The bias-corrected tornado frequency estimates differ markedly from the official tornado climatology and have important implications for tornado risk assessment, damage mitigation, and studies of climate change impacts on tornado activity.
Approximately 1–3 violent tornadoes hit Europe each decade. We use the ERA5 reanalysis and WRF model to reconstruct environments for 12 cases between 1957 and 2021. Violent tornadoes in Europe occur in a variety of synoptic and mesoscale patterns, but they share environmental similarities with significant tornadoes from the United States. Downscaling simulations to 3 km grid‐spacing showed the added value of improved resolution in representing local convective environments. In 8 out of 12 simulations, the model indicated updraft helicity (UH) tracks in a favorable convective environment in spatial (+/−50 km) and temporal (+/−3 hr) proximity to the tornado report. Tornadoes were accompanied by a mean 0–6 km wind shear of 24.3 m s −1 and mean convective available potential energy of 1678 J kg −1 . The combination of UH tracks with convective environments offers promising results for operational forecasting in Europe, and should be explored in future studies.
Downbursts are strong downdrafts of negatively buoyant air associated with convective storms and are capable of producing severe near-surface winds. Microbursts and macrobursts are subcategories of downbursts with the horizontal extent of damaging winds smaller or larger than 4 km, respectively. From January 2000 to June 2020, the Severe Weather Event Reports provided by the National Centers for Environmental Information (hereafter: Storm Events Database) contained 927 downburst, 914 microburst, and only 27 macroburst entries. We found a spatial variability of reported downbursts that is unlikely to be a result of natural processes, but rather artificially caused by the population density. An example of this bias is the abrupt decline in the number of reported events between southern and northern Arizona. Combining the Storm Events Database, ERA5 reanalysis and lightning data from the National Lightning Detection Network, we showed that cold pool strength, low-level lapse rates, WINDEX, lifted condensation level, DCAPE, WMAXSHEAR, derecho composite parameter, 2-m temperature, delta theta-e and mean low-level relative humidity demonstrate some value in downburst prediction. By combining the best predictor (cold pool strength) with the least correlated WMAXSHEAR, we created a downburst environment index (DEI) and used it to model climatological frequency of favorable downburst environments. Our analysis has shown that favorable downburst environments conditioned on lightning are the most frequent during summer over Southwest and Southeast with the most extreme environments across Great Plains. The vertical profiles of theta-e for the downburst events from reanalysis are further compared against nonsevere thunderstorms and rawinsonde data from four downburst field measurement campaigns. The results show that changes in theta-e over the lowest 200 hPa are the most important for downburst formation.
A series of webinars and panel discussions were conducted on the topic of the evolving role of humans in weather prediction and communication, in recognition of the 100th anniversary of the founding of the AMS. One main theme that arose was the inevitability that new tools using artificial intelligence will improve data analysis, forecasting, and communication. We discussed what tools are being created, how they are being created, and how the tools will potentially affect various duties for operational meteorologists in multiple sectors of the profession. Even as artificial intelligence increases automation, humans will remain a vital part of the forecast process as that process changes over time. Additionally, both university training and professional development must be revised to accommodate the evolving forecasting process, including addressing the need for computing and data skills (including artificial intelligence and visualization), probabilistic and ensemble forecasting, decision support, and communication skills. These changing skill sets necessitate that both the U.S. Government’s Meteorologist General Schedule 1340 requirements and the AMS standards for a bachelor’s degree need to be revised. Seven recommendations are presented for student and forecaster preparation and career planning, highlighting the need for students and operational meteorologists to be flexible lifelong learners, acquire new skills, and be engaged in the changes to forecast technology in order to best serve the user community throughout their careers. The article closes with our vision for the ways that humans can maintain an essential role in weather prediction and communication, highlighting the interdependent relationship between computers and humans.
United States tornado records form the basis for a variety of meteorological, climatological and disaster-risk analyses, but how reliable are they in light of changing standards for rating, as with the 2007 transition of Fujita (F) to Enhanced Fujita (EF) damage scales? To what extent are recorded tornado metrics subject to such influences that may be nonmeteorological in nature? While addressing these questions with utmost thoroughness is too large of a task for any one study, and may not be possible given the many variables and uncertainties involved, some variables that are recorded in large samples are ripe for new examination. We assess basic tornado-path characteristics—damage rating, length, width, and occurrence time, as well as some combined and derived measures—for a 24-yr period of constant path-width recording standard that also coincides with National Weather Service modernization and the WSR-88D deployment era. The middle of that period (in both time and approximate tornado counts) crosses the official switch from F to EF. At least minor shifts in all assessed path variables are associated directly with that change, contrary to the intent of EF implementation. Major and essentially stepwise expansion of tornadic path widths occurred immediately upon EF usage, and widths have expanded still further within the EF era. We also document lesser increases in path lengths, and in tornadoes rated at least EF1 compared to EF0. These apparently secular changes in the tornado data can impact research dependent on bulk tornado-path characteristics and damage-assessment results.
In any discussion of forecast evaluation, it is tempting to fall back on statements reflecting unverified assumptions: "this tornado warning had lower skill because the underlying meteorology reflected a complicated or atypical scenario," or "that forecast performed worse than we would have expected given the straightforward setup." These statements of what is and is not a reasonable expectation for warning skill are particularly relevant as the meteorological community's focus has begun to emphasize non-classic storm environments (e.g., tornadoes spawned by quasi-linear convective systems). In this paper, we build a proof-of-concept methodology to quantify the effect of the near-storm environment on tornado warning skill, and we then test these methods on a 15-yr dataset composed of tens of thousands of tornado events and warnings over the contiguous United States. Our findings include that significant tornadoes rated (E)F2+ have a higher probability of detection (POD) than expected based on their near-storm environments, that nocturnal tornadoes have both worse POD and false alarm ratio (FAR) than even their marginal near-storm environments would suggest, and that tornadoes occurring during the summer months also show worse POD and FAR than their environment-based expectation. Quantifying these shifts in performance in an environmental skill score framework allows us to target the situations in which the greatest improvements may be possible, in terms of forecaster training and/or conceptual models. This work also highlights the essential question that should always be asked in the context of forecast verification: what, exactly, is the baseline standard to which we are comparing forecast performance?
While many studies have looked at the quality of forecast products, few have attempted to understand the relationship between them. We begin to consider whether or not such an influence exists by analyzing storm-based tornado warning product metrics with respect to whether they occurred within a severe weather watch and, if so, what type of watch they occurred within.The probability of detection, false alarm ratio, and lead time all show a general improvement with increasing watch severity. In fact, the probability of detection increased more as a function of watch-type severity than the change in probability of detection during the time period of analysis. False alarm ratio decreased as watch type increased in severity, but with a much smaller magnitude than the difference in probability of detection. Lead time also improved with an increase in watch-type severity. Warnings outside of any watch had a mean lead time of 5.5 minutes, while those inside of a particularly dangerous situation tornado watch had a mean lead time of 15.1 minutes. These results indicate that the existence and type of severe weather watch may have an influence on the quality of tornado warnings. However, it is impossible to separate the influence of weather watches from possible differences in warning strategy or differences in environmental characteristics that make it more or less challenging to warn for tornadoes. Future studies should attempt to disentangle these numerous influences to assess how much influence intermediate products have on downstream products.
The term “tornado outbreak” appeared in the meteorological literature in the 1950s and was used to highlight severe weather events with multiple tornadoes. The exact meaning of “tornado outbreak,” however, evolved over the years. Depending on the availability of scientific data, technological advancements, and the intended purpose of these definitions, authors offered a diverse set of approaches to shape the perception and applications of the term “tornado outbreak.” This paper reviews over 200 peer-reviewed publications—by decade—to outline the evolving nature of the “tornado outbreak” definition and to examine the changes in the “tornado outbreak” definition or its perception. A final discussion highlights the importance, limitations, and potential future evolution of what defines a “tornado outbreak.”
Long-term trends in the historical frequency of environments supportive of atmospheric convection are unclear, and only partially follow the expectations of a warming climate. This uncertainty is driven by the lack of unequivocal changes in the ingredients for severe thunderstorms (i.e., conditional instability, sufficient low-level moisture, initiation mechanism, and vertical wind shear). ERA5 hybrid-sigma data allow for superior characterization of thermodynamic parameters including convective inhibition, which is very sensitive to the number of levels in the lower troposphere. Using hourly data we demonstrate that long-term decreases in instability and stronger convective inhibition cause a decline in the frequency of thunderstorm environments over the southern United States, particularly during summer. Conversely, increasingly favorable conditions for tornadoes are observed during winter across the Southeast. Over Europe, a pronounced multidecadal increase in low-level moisture has provided positive trends in thunderstorm environments over the south, central, and north, with decreases over the east due to strengthening convective inhibition. Modest increases in vertical wind shear and storm-relative helicity have been observed over northwestern Europe and the Great Plains. Both continents exhibit negative trends in the fraction of environments with likely convective initiation. This suggests that despite increasing instability, thunderstorms in a warming climate may be less likely to develop due to stronger convective inhibition and lower relative humidity. Decreases in convective initiation and resulting precipitation may have long-term implications for agriculture, water availability, and the frequency of severe weather such as large hail and tornadoes. Our results also indicate that trends observed over the United States cannot be assumed to be representative of other continents.
Increasing tornado warning skill in terms of the probability of detection and false alarm ratio remains an important operational goal. Although many studies have examined tornado warning performance in a broad sense, less focus has been placed on warning performance within sub-daily convective events. In this study, we use the NWS tornado verification database to examine tornado warning performance by order-of-tornado within each convective day. We combine this database with tornado reports to relate warning performance to environmental characteristics. On convective days with multiple tornadoes, the first tornado is warned significantly less often than the middle and last tornadoes. More favorable kinematic environmental characteristics, like increasing 0–1-km shear and storm-relative helicity, are associated with better warning performance related to the first tornado of the convective day. Thermodynamic and composite parameters are less correlated to warning performance. During tornadic events, over half of false alarms occur after the last tornado of the day decays, and false alarms are twice as likely to be issued during this time than before the first tornado forms. These results indicate that forecasters may be better “primed” (or more prepared) to issue warnings on middle and last tornadoes of the day, and must overcome a higher threshold to warn on the first tornado of the day. To overcome this challenge, using kinematic environmental characteristics and intermediate products on the watch-to-warning scale may help.
In this study we compared 3.7 mln rawinsonde observations from 232 stations over Europe and North America with proximal vertical profiles from ERA5 and MERRA2 to examine how well reanalysis depicts observed convective parameters. Larger differences between soundings and reanalysis are found for thermodynamic theoretical parcel parameters, low-level lapse rates and low-level wind shear. In contrast, reanalysis best represents temperature and moisture variables, mid-tropospheric lapse rates, and mean wind. Both reanalyses underestimate CAPE, low-level moisture and wind shear, particularly when considering extreme values. Overestimation is observed for low-level lapse rates, mid-tropospheric moisture and the level of free convection. Mixed-layer parcels have overall better accuracy when compared to most-unstable, especially considering convective inhibition and lifted condensation level. Mean absolute error for both reanalyses has been steadily decreasing over the last 39 years for almost every analyzed variable. Compared to MERRA2, ERA5 has higher correlations and lower mean absolute errors. MERRA2 is typically drier and less unstable over central Europe and the Balkans, with the opposite pattern over western Russia. Both reanalyses underestimate CAPE and CIN over the Great Plains. Reanalyses are more reliable for lower elevations stations and struggle along boundaries such as coastal zones and mountains. Based on the results from this and prior studies we suggest that ERA5 is likely one of the most reliable available reanalysis for exploration of convective environments, mainly due to its improved resolution. For future studies we also recommend that computation of convective variables should use model levels that provide more accurate sampling of the boundary-layer conditions compared to less numerous pressure levels.
Globally, thunderstorms are responsible for a significant fraction of rainfall, and in the mid-latitudes often produce extreme weather, including large hail, tornadoes and damaging winds. Despite this importance, how the global frequency of thunderstorms and their accompanying hazards has changed over the past 4 decades remains unclear. Large-scale diagnostics applied to global climate models have suggested that the frequency of thunderstorms and their intensity is likely to increase in the future. Here, we show that according to ERA5 convective available potential energy (CAPE) and convective precipitation (CP) have decreased over the tropics and subtropics with simultaneous increases in 0–6 km wind shear (BS06). Conversely, rawinsonde observations paint a different picture across the mid-latitudes with increasing CAPE and significant decreases to BS06. Differing trends and disagreement between ERA5 and rawinsondes observed over some regions suggest that results should be interpreted with caution, especially for CAPE and CP across tropics where uncertainty is the highest and reliable long-term rawinsonde observations are missing.