The quality of weather forecasts has improved considerably in recent decades as models can better represent the complexity of the Earth’s climate system, benefitting from assimilation of comprehensive Earth observation data and increased computational resources. Analysis of errors is an integral part of numerical weather prediction to produce better quality forecasts. The Earth’s climate, being a highly complex interacting system, often gives rise to significant statistical relationships between the states of the climate at distant geographical locations. Likewise, correlated errors in forecasting the state of the system can arise from predictable relationships between forecast errors at various regions resulting from an underlying systematic or random process. Estimation of error correlations is very important for producing quality forecasts and is a key issue for data assimilation. However, the size of the corresponding correlation matrix is larger than what is possible to represent on geographical maps in order to diagnose its full spatial variation. In this work, we propose an approach based on complex network theory to quantitatively study the spatiotemporal coherent structures of medium-range forecast errors of different climate variables. We demonstrate that the spatial variation of the network measures computed from the error correlation matrix can provide insights into the origin of forecast errors in a climate variable by identifying spatially coherent patterns of regions having common sources of error. Notably, the network topology of forecast errors of a climate variable is significantly different from those of random networks corresponding to a deterministic phenomenon which the model fails to simulate adequately. This is especially important to reveal the spatial heterogeneity of the errors – for example, the forecast errors of outgoing long-wave radiation in tropical regions can be correlated across very long distances, indicating an underlying climate mechanism as the source of the error. Additionally, we highlight that these structures of forecast errors may not always be directly derivable from the spatiotemporal co-variability pattern of the corresponding climate variable, contrary to the expectations that the patterns should resemble each other. We further employ other common statistical tools such as, empirical orthogonal functions, to support these findings. Our results underline the potential of complex networks as a very promising diagnostic tool to gain better understanding of the spatial variation, origin, and propagation of forecast errors.
We evaluate the skill and jumpiness of the ECMWF medium-range ensemble (ENS) in predicting tropical cyclone genesis in the Atlantic basin. Focusing on the probabilistic performance of the ENS, we assess how far in advance the ENS can predict genesis, quantify the consistency (jumpiness) from run to run, and investigate what factors influence the skill and consistency. We find that first indications of genesis are picked up at least 7 days ahead in 50% of the observed cases, although strong signals often only appear less than 3 days before genesis. There are significant regional differences, with observed genesis events predicted 2-3 days earlier in the eastern Atlantic than in other areas. The genesis probabilities can be jumpy from run to run, and the jumpiest cases are in the more skillful regions (central and eastern Atlantic) and for situations where the initial signal for genesis appears at longer lead time. In the eastern Atlantic, there is a tendency for the ENS tracks to reach tropical storm strength earlier and further east than observed; this model bias can affect both skill and jumpiness of the genesis forecasts. Our results provide guidance to forecasters on how to use and interpret the ENS predictions. Areas for future work include the link between early intensification in the eastern Atlantic and African easterly wave activity, the relationship between skill and the TC development pathways, and the impact of systematic analysis differences between 0000 UTC and 1200 UTC on forecast intensity. SIGNIFICANCE STATEMENT: Forecasting where and when tropical cyclones will appear increases the lead time at which decision-makers can begin to take preparatory mitigating action. Numerical weather prediction models can provide important guidance but sometimes are not consistent from one run to the next. We evaluate the skill and consistency of a state-of-the-art global model in predicting the formation of tropical cyclones up to 10 days ahead and provide guidance to forecasters on how to use and interpret the model predictions. We show that the formation of tropical cyclones can be predicted 2-3 days earlier in the eastern Atlantic than in the western Atlantic and identify some of the factors influencing both skill and consistency.
Abstract. Severe heat waves lasting for weeks and expanding over hundreds of kilometres in horizontal scale have many harmful impacts on health, ecosystems, societies, and economy. Under the ongoing climate change heat waves are becoming even longer and hotter, and as proactive adaptation, the development of early warning services is essential. Weather forecasts in extended range (2 weeks to 1 month) tend to indicate a higher skill in predicting warm extremes than average temperature events in Europe. We verified hindcasts of the European Centre for Medium-Range Weather Forecasts (ECMWF) in forecasting heat wave days, i.e., periods with the 5-day mean temperature being above its 90th percentile. The verification was done in 5° × 2° resolution over Europe, based on the forecast week (1 to 4 weeks). In the first forecast week, it is evident that across Europe, the accuracy of ECMWF heat wave forecasts surpasses that of a mere climatological forecast. Even into the second week, in many places in Europe, the ECMWF forecasts prove to be more reliable than their statistical counterparts. However, if we extend the forecast lead time to 3–4 weeks, predictability begins to lower to such a level that it can no longer be said, with the exception of Southeastern Europe, that the forecasts in general were statistically significantly better than the statistical forecast. Nonetheless, intense and prolonged heat waves during the third forecast weeks appear to have a higher-than-average level of predictability.
We investigate the run-to-run consistency (jumpiness) of ensemble forecasts of tropical cyclone tracks from three global centers: ECMWF, the Met Office, and NCEP. We use a divergence function to quantify the change in cross -track position between consecutive ensemble forecasts initialized at 12-h intervals. Results for the 2019-21 North Atlantic hurricane season show that the jumpiness varied substantially between cases and centers, with no common cause across the different ensemble systems. Recent upgrades to the Met Office and NCEP ensembles reduced their overall jumpiness to match that of the ECMWF ensemble. The average divergence over the set of cases provides an objective measure of the expected change in cross-track position from one forecast to the next. For example, a user should expect on average that the ensemble mean position will change by around 80-90 km in the cross-track direction between a forecast for 120 h ahead and the updated forecast made 12 h later for the same valid time. This quantitative information can support users' decision-making, for example, in deciding whether to act now or wait for the next forecast. We did not find any link between jumpiness and skill, indicating that users should not rely on the consistency between successive forecasts as a measure of confidence. Instead, we suggest that users should use ensemble spread and probabilistic information to assess forecast uncertainty, and consider multimodel combinations to reduce the effects of jumpiness. SIGNIFICANCE STATEMENT: Forecasting the tracks of tropical cyclones is essential to mitigate their impacts on society. Numerical weather prediction models provide valuable guidance, but occasionally there is a large jump in the predicted track from one run to the next. This jumpiness complicates the creation and communication of consistent forecast advisories and early warnings. In this work we aim to better understand forecast jumpiness and we provide practical information to forecasters to help them better use the model guidance. We show that the jumpiest cases are different for different modeling centers, that recent model upgrades have reduced forecast jumpiness, and that there is not a strong link between jumpiness and forecast skill.
The term jet stream generally refers to a narrow region of intense winds near the top of the midlatitude or subtropical troposphere. It is in the midlatitude jet stream where instabilities and waves may develop into synoptic-scale systems, which in turn makes accurately resolving the structure of the jet stream and associated features critical for atmospheric development, predictability, and impacts, such as extreme precipitation and winds. Using dropwindsonde observations collected during the Atmospheric River Reconnaissance (AR Recon) campaign from 2020 to 2022, this study assesses the North Pacific jet stream structure in the European Centre for Medium-Range Weather Forecasts (ECMWF) Integrated Forecasting System (IFS). Results show that the IFS has a slow-wind bias on the lead times assessed, with the strongest winds (& GE;50 m & BULL;s(-1)) having a bias of up to -1.88 m & BULL;s(-1) on forecast day 4. Also, the IFS cannot resolve the sharp potential vorticity (PV) gradient across the jet stream and tropopause, and this PV gradient weakens with forecast lead time. Cases with larger wind biases are characterized by higher PV biases and PV biases tend to be larger for cases with a higher horizontal PV gradient. These results suggest that further model-based experiments are needed to identify and address these biases, which could ultimately yield increased forecast accuracy.
Understanding error properties is an essential part in numerical weather prediction. Predictable relationship between errors of different regions due to some underlying systematic or random process can give rise to correlated errors. Estimation of error correlation is crucial for improvement of forecasts. However, the size of the corresponding correlation matrix is larger than what is possible to represent on geographical maps in order to diagnose its full spatial variation. Here, we propose a complex network‐based approach to analyse forecast error correlations that enables us to estimate the spatially varying component of the error. A quantitative study of the spatio‐temporal coherent structures of medium‐range forecast errors of different climate variables using network measures can reveal common sources of errors. Such information is crucial, especially in cases such as the outgoing long‐wave radiation, in which errors are correlated across very long distances, indicating an underlying climate mechanism as the source of the error. We show that the spatial patterns of forecast error co‐variability may not be the same as that of the corresponding climate variable itself, thereby implying that the mechanisms behind the correlated errors may be different from the climate processes responsible for the spatio‐temporal interactions of the climate variable. Our results highlight the importance of diagnosing the full spatial variation of error correlations to understand the origin and propagation of forecast errors, and demonstrate complex networks to be a promising diagnostic tool in this regard.
The relative operating characteristic (ROC) curve is a popular diagnostic tool in forecast verification, with the area under the ROC curve (AUC) used as a verification metric measuring the discrimination ability of a forecast. Along with calibration, discrimination is deemed as a fundamental probabilistic forecast attribute. In particular, in ensemble forecast verification, AUC provides a basis for the comparison of potential predictive skill of competing forecasts. While this approach is straightforward when dealing with forecasts of common events (e.g. probability of precipitation), the AUC interpretation can turn out to be oversimplistic or misleading when focusing on rare events (e.g. precipitation exceeding some warning criterion). How should we interpret AUC of ensemble forecasts when focusing on rare events? How can changes in the way probability forecasts are derived from the ensemble forecast affect AUC results? How can we detect a genuine improvement in terms of predictive skill? Based on verification experiments, a critical eye is cast on the AUC interpretation to answer these questions. As well as the traditional trapezoidal approximation and the well-known bi-normal fitting model, we discuss a new approach which embraces the concept of imprecise probabilities and relies on the subdivision of the lowest ensemble probability category.
WeatherVolume 76, Issue 10 p. 355-355 Obituary Obituary: Anders Persson David Richardson, Corresponding Author david.richardson@ecmwf.int david.richardson@ecmwf.intSearch for more papers by this authorKen Mylne, Search for more papers by this author David Richardson, Corresponding Author david.richardson@ecmwf.int david.richardson@ecmwf.intSearch for more papers by this authorKen Mylne, Search for more papers by this author First published: 10 August 2021 https://doi.org/10.1002/wea.4054Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onEmailFacebookTwitterLinked InRedditWechat No abstract is available for this article. Volume76, Issue10Special Issue: COP26October 2021Pages 355-355 RelatedInformation
A key aim of observational campaigns is to sample atmosphere–ocean phenomena to improve understanding of these phenomena, and in turn, numerical weather prediction. In early 2018 and 2019, the Atmospheric River Reconnaissance (AR Recon) campaign released dropsondes and radiosondes into atmospheric rivers (ARs) over the northeast Pacific Ocean to collect unique observations of temperature, winds, and moisture in ARs. These narrow regions of water vapor transport in the atmosphere—like rivers in the sky—can be associated with extreme precipitation and flooding events in the midlatitudes. This study uses the dropsonde observations collected during the AR Recon campaign and the European Centre for Medium-Range Weather Forecasts (ECMWF) Integrated Forecasting System (IFS) to evaluate forecasts of ARs. Results show that ECMWF IFS forecasts 1) were colder than observations by up to 0.6 K throughout the troposphere; 2) have a dry bias in the lower troposphere, which along with weaker winds below 950 hPa, resulted in weaker horizontal water vapor fluxes in the 950–1000-hPa layer; and 3) exhibit an underdispersiveness in the water vapor flux that largely arises from model representativeness errors associated with dropsondes. Four U.S. West Coast radiosonde sites confirm the IFS cold bias throughout winter. These issues are likely to affect the model’s hydrological cycle and hence precipitation forecasts.
Spatial variability of precipitation is analyzed to characterize to what extent precipitation observed at a single location is representative of precipitation over a larger area. Characterization of precipitation representativeness is made in probabilistic terms using a parametric approach, namely, by fitting a censored shifted gamma distribution to observation measurements. Parameters are estimated and analyzed for independent precipitation datasets, among which one is based on high-density gauge measurements. The results of this analysis serve as a basis for accounting for representativeness error in an ensemble verification process. Uncertainty associated with the scale mismatch between forecast and observation is accounted for by applying a perturbed-ensemble approach before the computation of scores. Verification results reveal a large impact of representativeness error on precipitation forecast reliability and skill estimates. The parametric model and estimated coefficients presented in this study could be used directly for forecast postprocessing to partly compensate for the limitation of any modeling system in terms of precipitation subgrid-scale variability.
An expected benefit of ensemble forecasts is that a sequence of consecutive forecasts valid for the same time will be more consistent than an equivalent sequence of individual forecasts. Inconsistent (jumpy) forecasts can cause users to lose confidence in the forecasting system. We present a first systematic, objective evaluation of the consistency of the European Centre for Medium‐Range Weather Forecasts (ECMWF) ensemble using a measure of forecast divergence that takes account of the full ensemble distribution. Focusing on forecasts of the North Atlantic Oscillation and European Blocking regimes up to 2 weeks ahead, we identify occasional large inconsistency between successive runs, with the largest jumps tending to occur at 7–9 days lead. However, care is needed in the interpretation of ensemble jumpiness. An apparent clear flip‐flop in a single index may hide a more complex predictability issue which may be better understood by examining the ensemble evolution in phase space.
The strength of the stratospheric polar vortex influences the surface weather in the Northern Hemisphere in winter; a weaker (stronger) than average stratospheric polar vortex is connected to negative (positive) Arctic Oscillation (AO) and colder (warmer) than average surface temperatures in northern Europe within weeks or months. This holds the potential for forecasting in that timescale. We investigate here if the strength of the stratospheric polar vortex at the start of the forecast could be used to improve the extended-range temperature forecasts of the European Centre for Medium-Range Weather Forecasts (ECMWF) and to find periods with higher prediction skill scores. For this, we developed a stratospheric wind indicator (SWI) based on the strength of the stratospheric polar vortex and the phase of the AO during the following weeks. We demonstrate that there was a statistically significant difference in the observed surface temperature in northern Europe within the 3–6 weeks, depending on the SWI at the start of the forecast. When our new SWI was applied in post-processing the ECMWF's 2-week mean temperature reforecasts for weeks 3–4 and 5–6 in northern Europe during boreal winter, the skill scores of those weeks were slightly improved. This indicates there is some room for improving the extended-range forecasts, if the stratosphere–troposphere links were better captured in the modelling. In addition to this, we found that during the boreal winter, in cases where the polar vortex was weak at the start of the forecast, the mean skill scores of the 3–6 weeks' surface temperature forecasts were higher than average.
IMproving PRedictions and management of hydrological EXtremes (IMPREX) was a European Union Horizon 2020 project that ran from September 2015 to September 2019. IMPREX aimed to improve society’s ability to anticipate and respond to future extreme hydrological events in Europe across a variety of uses in the water-related sectors (flood forecasting, drought risk assessment, agriculture, navigation, hydropower and water supply utilities). Through the engagement with stakeholders and continuous feedback between model outputs and water applications, progress was achieved in better understanding the way hydrological predictions can be useful to (and operationally incorporated into) problem-solving in the water sector. The work and discussions carried out during the project nurtured further reflections toward a common vision for hydrological prediction. In this article, we summarized the main findings of the IMPREX project within a broader overview of hydrological prediction, providing a vision for improving such predictions. In so doing, we first presented a synopsis of hydrological and weather forecasting, with a focus on medium-range to seasonal scales of prediction for increased preparedness. Second, the lessons learned from IMPREX were discussed. The key findings were the gaps highlighted in the global observing system of the hydrological cycle, the degree of accuracy of hydrological models and the techniques of post-processing to correct biases, the origin of seasonal hydrological skill in Europe and user requirements of hydrometeorological forecasts to ensure their appropriate use in decision-making models and practices. Last, a vision for how to improve these forecast systems/products in the future was expounded, including advancing numerical weather and hydrological models, improved earth monitoring and more frequent interaction between forecasters and users to tailor the forecasts to applications. We conclude that if these improvements can be implemented in the coming years, earth system and hydrological modelling will become more skillful, thus leading to socioeconomic benefits for the citizens of Europe and beyond.
We investigate how users combine objective probabilities with their own subjective feelings when deciding how to act on weather forecast information. Results are based on two scenarios investigated at a Live Science event held by the Royal Meteorological Society. When deciding whether to go to the beach with the possibility of warm, dry weather, we find that users attempt to identify their ‘Bayes Action’: the one which minimises their expected negative feeling or utility. Key factors are the ‘thrill’ of a nice day at the beach and the ‘pain’ of coping with, for example, children in wet weather, and the costs of travel. The users' threshold probabilities for deciding to go to the beach thus approximately define their distribution of cost/loss ratios. This is used to calculate a ‘User Brier Score’ (UBS): a measure of the overall utility to society, and which could be used to guide forecast system development. When applied to operational ensemble forecasts issued by the European Centre for Medium‐Range Weather Forecasts (ECMWF) over the period 1995–2018, the UBS tends to be higher (i.e., worse) than the Brier Score, largely because users tended not to exhibit high cost/loss ratios. When deciding whether to leave a campsite in the face of potentially dangerous gales, users try to find a balance between the ‘regret’ of serious injury and the ‘pain’ of spoiling an enjoyable holiday. Some users decide to stay even at high probabilities of serious consequences – partly due to a lack of experience. On the other hand, forecasts suffer from ‘complete misses’ – where probabilities of zero are accompanied by non‐negligible outcome frequencies. These dominate the overall Brier Score. The frequency of complete misses halved over the period 1995–2018: a welcome improvement for users who do wish to avoid danger at low probabilities.
Atmospheric rivers lie behind many extreme precipitation and flood episodes in the mid-latitudes. Better forecasts of atmospheric rivers and their impacts could help with preparedness. Here we argue that a comprehensive and systematic observational campaign could help advance numerical weather prediction, and thereby provide a path towards much improved forecasts of atmospheric rivers. We envision an interdisciplinary European–American observational campaign in the North Atlantic to identify and address numerical weather prediction errors in atmospheric rivers, and the associated extratropical cyclones. Insights gained could be applied in other regions. With improved understanding of the physiography of river basins and insights into their flood response to extreme precipitation, the impacts of atmospheric rivers can also be forecast more reliably.