Ensemble weather forecasts provide a probabilistic description of the future state of the atmosphere and give users flow-dependent estimates of forecast uncertainty. Here, we introduce AIFS-CRPS, an ensemble variant of the machine-learned Artificial Intelligence Forecasting System (AIFS) developed at ECMWF. Its loss function is the almost fair Continuous Ranked Probability Score (afCRPS). It is based on a proper score, the CRPS, but approximately removes the bias in the score due to finite ensemble size yet avoids a degeneracy of the fair CRPS. The trained model is stochastic and can generate as many exchangeable members as desired. For medium-range forecasts AIFS-CRPS outperforms the physics-based Integrated Forecasting System (IFS) ensemble for the majority of variables and lead times. For subseasonal forecasts, AIFS-CRPS outperforms the IFS ensemble before calibration and is competitive with the IFS ensemble when forecasts are evaluated as anomalies to remove the influence of model biases.
We assess the impact of a multi-scale loss formulation for training probabilistic machine-learned weather forecasting models. The multi-scale loss is tested in AIFS-CRPS, a machine-learned weather forecasting model developed at the European Centre for Medium-Range Weather Forecasts (ECMWF). AIFS-CRPS is trained by directly optimising the almost fair continuous ranked probability score (afCRPS). The multi-scale loss better constrains small scale variability without negatively impacting forecast skill. This opens up promising directions for future work in scale-aware model training.
Probabilistic forecast models can be machine-learned from data using loss functions based on scoring rules such as the Continuous Ranked Probability Score (CRPS). This note summarises a preliminary study comparing versions of AIFS-CRPS, a global weather forecast model, trained with different univariate and multivariate scoring rules that aim to explicitly represent scale-awareness in the loss function. In the first part, we compare the (almost) fair CRPS, a fair global energy score, and a graph energy score based on node neighbourhoods. Across standard verification metrics, forecast skill is broadly similar. In the extratropics we find only small differences, while in the tropics the graph energy score setup performs somewhat better and the global energy score shows some degradation. These results suggest that multivariate scores are a viable alternative to CRPS-based training for global machine-learned weather forecasting. In the second part of the study, we analyse how different scoring rules and scale-aware loss constraints shape the spectra of forecast fields. It is apparent that any form of explicit scale-awareness improves realism. Here, the largest differences are likely associated with different effective weights per scale.
We introduce a probabilistic diffusion-based method for global atmospheric downscaling implemented within the Anemoi framework. The approach transforms low-resolution ensemble forecasts into high-resolution ensembles by learning the conditional distribution of finer-scale residuals, defined as the difference between the high-resolution fields and the interpolated low-resolution inputs. The system is trained on reforecast pairs from ECMWF IFS, using coarse fields at 100 km to reconstruct fine-scale variability at 30 km resolution. The bulk of the training focuses on recovering small-scale structures, while fine-tuning in high-noise regimes enables the generation of extremes. Evaluation against the medium-range IFS ensemble target shows that the model increases probabilistic skill (FCRPS) for surface variables, reproduces target power spectra at small scales, captures physically consistent multivariate relationships such as wind-pressure coupling, and generates extreme values consistent with those of the target ensemble in tropical cyclones.
Monte Carlo techniques are the method of choice for making probabilistic predictions of an outcome in several disciplines. Usually, the aim is to generate calibrated predictions which are statistically indistinguishable from the outcome. Developers and users of such Monte Carlo predictions are interested in evaluating the degree of calibration of the forecasts. Here, we consider predictions of $p$-dimensional outcomes sampling a multivariate Gaussian distribution and apply the Box ordinate transform (BOT) to assess calibration. However, this approach is known to fail to reliably indicate calibration when the sample size n is moderate. For some applications, the cost of obtaining Monte-Carlo estimates is significant, which can limit the sample size, for instance, in model development when the model is improved iteratively. Thus, it would be beneficial to be able to reliably assess calibration even if the sample size n is moderate. To address this need, we introduce a fair, sample size- and dimension-dependent version of the Gaussian sample BOT. In a simulation study, the fair Gaussian sample BOT is compared with alternative BOT versions for different miscalibrations and for different sample sizes. Results confirm that the fair Gaussian sample BOT is capable of correctly identifying miscalibration when the sample size is moderate in contrast to the alternative BOT versions. Subsequently, the fair Gaussian sample BOT is applied to two to 12-dimensional predictions of temperature and vector wind using operational ensemble forecasts of the European Centre for Medium-Range Weather Forecasts (ECMWF). Firstly, perfectly reliable situations are considered where the outcome is replaced by a forecast that samples the same distribution as the members in the ensemble. Secondly, the BOT is computed using estimates of the actual temperature and vector wind from ECMWF analyses.
In evaluating multivariate probabilistic forecasts predicting vector quantities such as a weather variable at multiple locations or a wind vector, an important step is the assessment of their calibration and reliability. Here, we focus on the Gaussian Box ordinate transform (BOT), which is appropriate if the forecasts and observations are multivariate normal. The BOT is based on the Mahalanobis distance of the observation vector and the estimated Gaussian mean and asymptotically standard uniform if the forecasts and the observation are drawn from the same multivariate Gaussian law. However, for small ensemble sizes combined with high dimensionality, deviation from uniformity is substantial even for reliable forecasts, resulting in hump-shaped or triangular BOT histograms. To circumvent this problem, we derive an ensemble size and dimension-dependent fair version of the Gaussian BOT, where the uniformity holds for any combination of these parameters. With the help of a simulation study, first, we assess the behaviour of the fair BOT for various dimensions, ensemble sizes, and types of calibration misspecification. Then, using ensemble forecasts of vectors consisting of multiple combinations of upper-air weather variables, we demonstrate the usefulness of the fair BOT when multivariate normality is only an approximation.*Research was supported by the Hungarian National Research, Development and Innovation Office under Grant No. K142849.
Long-range ensemble forecasts are typically verified as anomalies with respect to a lead-time dependent climatological mean to remove the influence of systematic biases. However, common methods for calculating anomalies result in statistical inconsistencies between forecast and verification anomalies, even for a perfectly reliable ensemble. It is important to account for these systematic effects when evaluating ensemble forecast systems, particularly when tuning a model to improve the reliability of forecast anomalies or when comparing spread-error diagnostics between systems with different reforecast periods. Here, we show that unbiased variances and spread-error ratios can be recovered by deriving estimators that are consistent with the values that would be achieved when calculating anomalies relative to the true, but unknown, climatological mean. An elegant alternative is to construct forecast climatologies separately for each member, which ensures that forecast and verification anomalies are defined relative to reference climatological means with the same sampling uncertainty. This alternative approach has no impact on forecast ensemble means but systematically modifies the total variance and ensemble spread of forecast anomalies in such a way that anomaly-based spread-error ratios are unbiased without any explicit correction for climatology sample size. Furthermore, the improved statistical consistency of forecast and verification anomalies means that probabilistic forecast skill is optimised when the underlying forecast is also perfectly reliable. Alternative methods for anomaly calculation can thus impact probabilistic forecast skill, especially when anomalies are defined relative to climatologies with a small sample size. Finally, we demonstrate the equivalence of anomalies calculated using different methods after applying an unbiased statistical calibration.
Multivariate probabilistic verification is concerned with the evaluation of joint probability distributions of vector quantities such as a weather variable at multiple locations or a wind vector for instance. The logarithmic score (also known as ignorance score) is a proper score that is useful in this context. In order to apply the logarithmic score to ensemble forecasts, a choice for the density is required. Here, we are interested in the specific case when the density is multivariate normal with mean and covariance given by the ensemble mean and ensemble covariance respectively. Under the assumptions of multivariate normality and exchangeability of the ensemble members, a relationship is derived that describes how the logarithmic score depends on ensemble size. It permits one to estimate the score in the limit of infinite ensemble size from a small ensemble and thus produces a fair logarithmic score for multivariate ensemble forecasts under the assumption of normality. This generalises a study from 2018 that derived the ensemble size adjustment of the logarithmic score in the univariate case. An application to medium-range forecasts examines the usefulness of the ensemble size adjustments when multivariate normality is only an approximation. Predictions of vectors consisting of several different combinations of upper air variables are considered. Logarithmic scores are calculated for these vectors using ECMWF's daily extended-range forecasts, which consist of a 100-member ensemble. The probabilistic forecasts of these vectors are verified against operational European Centre for Medium-Range Weather Forecasts (ECMWF) analyses in the northern midlatitudes in autumn 2023. Scores are computed for ensemble sizes from 8 to 100. The fair logarithmic scores of ensembles with different cardinalities are very close, in contrast to the unadjusted scores, which decrease considerably with ensemble size. This provides evidence for the practical usefulness of the relationships derived.
Over the last three decades, ensemble forecasts have become an integral part of forecasting the weather. They provide users with more complete information than single forecasts as they permit to estimate the probability of weather events by representing the sources of uncertainties and accounting for the day-to-day variability of error growth in the atmosphere. This paper presents a novel approach to obtain a weather forecast model for ensemble forecasting with machine-learning. AIFS-CRPS is a variant of the Artificial Intelligence Forecasting System (AIFS) developed at ECMWF. Its loss function is based on a proper score, the Continuous Ranked Probability Score (CRPS). For the loss, the almost fair CRPS is introduced because it approximately removes the bias in the score due to finite ensemble size yet avoids a degeneracy of the fair CRPS. The trained model is stochastic and can generate as many exchangeable members as desired and computationally feasible in inference. For medium-range forecasts AIFS-CRPS outperforms the physics-based Integrated Forecasting System (IFS) ensemble for the majority of variables and lead times. For subseasonal forecasts, AIFS-CRPS outperforms the IFS ensemble before calibration and is competitive with the IFS ensemble when forecasts are evaluated as anomalies to remove the influence of model biases.
The large data volumes available in weather forecasting make post-processing an attractive field for applying machine learning. In turn, novel statistical machine learning methods that can be used to generate uncertainty information from a deterministic forecast are of great interest to forecast users, especially given the computational cost of running high resolution ensembles. In this work, we show how one such method, a Bayesian Neural Network (BNN), can be used to post-process a single global high resolution forecast for 2m temperature. This methodology improves both the accuracy of the forecast and adds uncertainty information, by predicting the distribution of the forecast error relative to its own analysis.Here we assess both model and data uncertainty using two different BNN approaches. In the first approach, the BNN’s parameters are defined to be distributions rather than deterministic parameters, thereby generating an ensemble of models that can be used to quantify model uncertainty. In the second approach, the BNN remains deterministic but predicts a distribution rather than a deterministic output thereby quantifying data uncertainty. Our BNN results are benchmarked against simpler statistical methods, as well as statistics from the ECMWF operational ensemble.Finally, in order to add trustworthiness to the BNN predictions, we apply an explainable AI technique (Layerwise Relevance Propagation) so as to understand whether the variables on which the BNN bases its prediction are physically reasonable or whether it is instead learning spurious correlations.
The stochastically perturbed parametrisation tendency (SPPT) scheme is a well‐established technique in ensemble forecasting to address model uncertainty by introducing perturbations into the tendencies provided by the physics parametrisations. The magnitude of the perturbations scales with the local net parametrisation tendency, resulting in large perturbations where diabatic processes are active. Rapidly ascending air streams, such as warm conveyor belts (WCBs) and organized tropical convection, are often driven by cloud diabatic processes and are therefore prone to such perturbations. This study investigates the effects of SPPT and initial condition perturbations on rapidly ascending air streams by computing trajectories in sensitivity experiments with the European Centre for Medium‐Range Weather Forecasts (ECMWF) ensemble prediction system, which are set up to disentangle the effects of initial conditions and physics perturbations. The results demonstrate that SPPT systematically increases the frequency of rapidly ascending air streams. The effect is observed globally, but is enhanced in regions where the latent heating along the trajectories is larger. Despite the frequency changes, there are only minor modifications to the physical properties of the trajectories due to SPPT. In contrast to SPPT, initial condition perturbations do not affect WCBs and tropical convection systematically. An Eulerian perspective on vertical velocities reveals that SPPT increases the frequency of strong upward motions compared with experiments with unperturbed model physics. Consistent with the altered vertical motions, precipitation rates are also affected by the model physics perturbations. The unperturbed control member shows the same characteristics as the experiments without SPPT regarding rapidly ascending air streams. We make use of this to corroborate the findings from the sensitivity experiments by analyzing the differences between perturbed and unperturbed members in operational ensemble forecasts of ECMWF. Finally, we give an explanation of how symmetric, zero‐mean perturbations can lead to a unidirectional response when applied in a nonlinear system.
The Stochastically Perturbed Parametrisations scheme (SPP) represents model uncertainty in numerical weather prediction by introducing stochastic perturbations into the physical parametrisation schemes. The perturbations are constructed in such a way that the internal consistency of the physical parametrisation schemes is preserved. We developed a revised version of SPP for the Integrated Forecasting System of the European Centre for Medium‐Range Weather Forecasts (ECMWF). The revised version introduces perturbations to additional quantities and modifies the probability distributions sampled by the scheme. Medium‐range ensemble forecasts with the revised SPP are considerably more skilful than ensemble forecasts with the original implementation of SPP. The revised version of SPP is similar, in terms of forecast skill, to the Stochastically Perturbed Parametrisation Tendency scheme (SPPT), which is currently used to represent model uncertainty in the operational ECMWF ensemble forecasts.
Reducing the numerical precision of the forecast model of the Integrated Forecasting System (IFS) of the European Centre for Medium‐Range Weather Forecasts (ECMWF) from double to single precision results in significant computational savings without negatively affecting forecast accuracy. The computational savings allow to increase the vertical resolution of the operational ensemble forecasts from 91 to 137 levels earlier than anticipated and before the next upgrade of ECMWF's high‐performance computing facility. This upgrade to 137 levels harmonises the vertical resolution of the medium‐range deterministic forecasts and the medium‐range and extended‐range ensemble forecasts. Increasing the vertical resolution of the ensemble forecasts substantially improves forecast skill for all lead times as well as the mean of the model climate. ECMWF's ensemble and deterministic forecasts will run operationally at single precision from IFS model cycle 47R2 onwards.
The philosophy of forecast verification is rather different between deterministic and probabilistic verification metrics: generally speaking, deterministic metrics measure differences, whereas probabilistic metrics assess reliability and sharpness of predictive distributions. This article considers the root‐mean‐square error (RMSE), which can be seen as a deterministic metric, and the probabilistic metric Continuous Ranked Probability Score (CRPS), and demonstrates that under certain conditions, the CRPS can be mathematically expressed in terms of the RMSE when these metrics are aggregated. One of the required conditions is the normality of distributions. The other condition is that, while the forecast ensemble need not be calibrated, any bias or over/underdispersion cannot depend on the forecast distribution itself. Under these conditions, the CRPS is a fraction of the RMSE, and this fraction depends only on the heteroscedasticity of the ensemble spread and the measures of calibration. The derived CRPS–RMSE relationship for the case of perfect ensemble reliability is tested on simulations of idealised two‐dimensional barotropic turbulence. Results suggest that the relationship holds approximately despite the normality condition not being met.
Most of the precipitation formation in extratropical cyclones occurs in the warm sector along an elongated air stream ahead of the cold front - the so-called warm conveyor belt (WCB). The WCB ascends slantwise from the planetary boundary layer into the upper troposphere, where its outflow interacts with the upper-level jet and modifies the Rossby wave structure. The ascent of WCBs is strongly driven by cloud-condensational processes, which are parametrized in numerical weather prediction models, and is therefore associated with forecast uncertainty. In the European Centre for Medium-Range Weather Forecasts (ECMWF) ensemble prediction system (EPS), model uncertainty related to parametrizations is represented by the so-called stochastically perturbed parametrization tendencies (SPPT)-scheme, which introduces multiplicative noise to the physics tendencies. In this study, we investigate the systematic effect of the SPPT-scheme on rapidly ascending air streams in the extratropics (i.e. WCBs) and on tropical convection by conducting sensitivity experiments with the ECMWF EPS based on the Integrated Forecasting System (IFS) model. The comparison of an experiment with an operational setup (initial condition and model physics perturbations) to one where model physics perturbations are switched off demonstrates that the SPPT-scheme systematically influences the activity of WCBs and tropical convection. Globally, rapidly ascending air streams, which are detected by applying trajectory analysis in each ensemble member, are enhanced by about 37% when SPPT is activated. Also the dynamical and physical characteristics of the trajectories are systematically modified: the latent heat release and the ascent speed are increased, while the outflow latitude is decreased. This systematic modulation is stronger in the tropics and weaker in the extratropics. A detailed investigation of vertical velocities indicates that SPPT increases the frequency of relatively strong upward motion related to WCBs and tropical convection, while slower upward motion is suppressed compared to the unperturbed experiment. Despite the symmetric, zero-mean nature of the perturbations, the response of rapidly ascending air streams to the SPPT-scheme is systematically unidirectional, pointing towards non-linearities in the underlying processes. This study shows that process-oriented diagnostics of weather systems help to advance the understanding of upscale impacts of the ensemble configuration on the representation of the large-scale circulation in numerical models.
Improving ensemble forecasts is a complex process which involves proper scores such as the continuous ranked probability score (CRPS). A homogeneous Gaussian (hoG) model is introduced in order to better understand the characteristics of the CRPS. An analytical formula is derived for the expected CRPS of an ensemble in the hoG model. The score is a function of the variance of the error of the ensemble mean, the mean error of the ensemble mean and the ensemble variance. The hoG model also provides a score decomposition into reliability and resolution components. We examine whether the hoG model provides a useful approximation of the CRPS when applied to operational ECMWF medium‐range ensemble forecasts. The hoG approximation describes the spatial variations of the CRPS well while moderately overestimating the mean score. Seasonal averages over large domains are within 10% of the actual CRPS. Furthermore, the ability to approximate score changes is evaluated by (a) comparing raw ensemble forecasts with postprocessed ensemble forecasts, and (b) by examining score changes associated with a recent upgrade of the IFS. Overall, the hoG approximation predicts the actual CRPS changes well. One of the main anticipated applications of the hoG approximation are new diagnostics in verification software used by NWP developers routinely. The purpose of the diagnostics is to help developers explain impacts of forecast system changes on the CRPS in terms of the changes in mean error, changes in error variance and changes in ensemble variance. The diagnostics require little additional computational resources compared to the alternative of verifying postprocessed versions of the ensemble forecasts. Therefore, it will be feasible to apply the diagnostics easily to all variables that are examined as part of the model development process.
This talk focusses on progress in ensemble forecasting methodology (Part I) and ensemble verification methodology (Part II). Operational ECMWF ensemble forecasts are global predictions from days to months ahead. At all forecast ranges, model uncertainties are represented stochastically with the Stochastically Perturbed Parametrization Tendency scheme (SPPT). Recently, considerable progress has been made in developing the Stochastically Perturbed Parametrization scheme (SPP). The SPP scheme offers improved physical consistency by naturally preserving the local conservation properties for energy and moisture of the unperturbed version of the corresponding parametrization. In contrast, the SPPT scheme lacks such local conservation properties, mainly because the scheme does not perturb fluxes at the surface and at the top of the atmosphere consistently with the tendency perturbations in the column. NWP research and development relies on scoring rules to judge whether or not a change to the forecast systems results in better ensemble forecasts. A new tool will be presented that can improve the understanding of score differences between sets of forecasts for a widely used proper score, the Continuous Ranked Probability Score (CRPS). An analytical expression has been derived for the CRPS when a homogeneous Gaussian (hoG) forecast-observation distribution is considered. This leads to an approximation of the CRPS when actual verification data are considered, which deviate from a homogeneous Gaussian distribution. The hoG approximation of the CRPS permits a useful decomposition of score differences. The methodology will be illustrated with verification data for medium-range weather forecasts.
ECMWF’s medium-range forecasts of near-surface weather parameters, such as 2 m temperature, humidity and 10 m wind speed, have become more skilful over the years, following the trend of improvements in the forecast skill of upper-air fields. However, they are still affected by systematic errors which have proved difficult to eliminate. Systematic forecast errors in temperature and humidity near the surface can be better understood by also examining errors higher up in the atmospheric boundary layer and in the soil. Meteorological observatories, also known as super-sites, provide long-term observational records of such vertical profiles. ECMWF started to use data from super-sites more systematically to evaluate the quality of forecasts in the lowest part of the atmosphere (up to 100m) and in the soil, in an attempt to disentangle sources of forecast error in near-surface weather parameters. Findings for 2-metre temperature errors in ECMWF forecasts at European super-sites suggest that the errors are partly the result of the model exchanging too much energy between the atmosphere and the land. However, the influence of other factors, such as errors resulting from the representation of vegetation in semi-arid areas and from small-scale variations in vegetation and soil type near measurement stations, mean that it is difficult to adjust the energy exchange in a way which leads to an overall error reduction on the European scale.