Accurate estimation of the upper tail of a distribution is crucial in seismology, where estimating the probability of extreme earthquake magnitudes is vital for risk assessment and mitigation. Traditional statistical methods often overlook expert knowledge, particularly regarding physical upper bounds on earthquake magnitudes. This paper introduces a novel methodology for estimating the upper tail distribution, integrating experts' knowledge on the physical processes through a conservative bound on the worst possible earthquakes. The methodology combines rigorous statistical techniques with expert judgement, creating a hybrid model that complements existing data-driven methods and enhances the reliability of tail estimates. We demonstrate the benefits of incorporating experts' knowledge through the application to data on human-induced earthquakes in the Netherlands. Within this paper, we focus on seismological magnitude modelling; however, the proposed methodology has the potential to be implemented as a generic extreme value approach for multiple problem settings.
We assess the value of calibrating forecast models for significant wave height HS, wind speed W and mean spectral wave period T-m for forecast horizons between zero and 168 h from a commercial forecast provider, to improve forecast performance for a location in the central North Sea. We consider two straightforward calibration models, linear regression (LR) and non-homogeneous Gaussian regression (NHGR), incorporating deterministic, control and ensemble mean forecast covariates. We show that relatively simple calibration models (with at most three covariates) provide good calibration and that addition of further covariates cannot be justified. Optimal calibration models (for the forecast mean of a physical quantity) always make use of the deterministic forecast and ensemble mean forecast for the same quantity, together with a covariate associated with a different physical quantity. The selection of optimal covariates is performed independently per forecast horizon, and the set of optimal covariates shows a large degree of consistency across forecast horizons. As a result, it is possible to specify a consistent model to calibrate a given physical quantity, incorporating a common set of three covariates for all horizons. For NHGR models of a given physical quantity, the ensemble forecast standard deviation for that quantity is skilful in predicting forecast error standard deviation, strikingly so for HS. We show that the consistent LR and NHGR calibration models facilitate reduction in forecast bias to near zero for all of HS, W and T-m, and that there is little difference between LR and NHGR calibration for the mean. Both LR and NHGR models facilitate reduction in forecast error standard deviation relative to naive adoption of the (uncalibrated) deterministic forecast, with NHGR providing somewhat better performance. Distributions of standardised residuals from NHGR are generally more similar to a standard Gaussian than those from LR.
The covXtreme software provides functionality for estimation of marginal and conditional extreme value models, non-stationary with respect to covariates, and environmental design contours. Generalised Pareto (GP) marginal models of peaks over threshold are estimated, using a piecewise-constant representation for the variation of GP threshold and scale parameters on the (potentially multidimensional) covariate domain of interest. The conditional variation of one or more associated variates, given a large value of a single conditioning variate, is described using the conditional extremes model of Heffernan and Tawn (2004), the slope term of which is also assumed to vary in a piecewise constant manner with covariates. Optimal smoothness of marginal and conditional extreme value model parameters with respect to covariates is estimated using cross-validated roughness-penalised maximum likelihood estimation. Uncertainties in model parameter estimates due to marginal and conditional extreme value threshold choice, and sample size, are quantified using a bootstrap resampling scheme. Estimates of environmental contours using various schemes, including the direct sampling approach of Huseby et al. 2013, are calculated by simulation or numerical integration under fitted models. The software was developed in MATLAB for metocean applications, but is applicable generally to multivariate samples of peaks over threshold data. The software and case study data can be downloaded from GitHub, with an accompanying user guide.
Soils are the largest terrestrial store of carbon, storing more carbon than the atmosphere and the biosphere combined. Soil carbon plays a key role in the delivery of a wide range of ecosystem services including climate regulation, food production, water quality and regulation and as such is often used as a proxy for ‘soil health’. International initiatives such as ‘Carbon 4 per mille’ highlight the potential for carbon sequestration in soils as a mechanism for climate mitigation, and the UK’s NetZero target depends on significant land-based carbon sequestration. Therefore, a need exists to quantify present-day soil carbon stocks at both regional and national scales to guide policy decisions and provide a baseline to enable estimates of carbon sequestration potential. To meet this need Digital Soil Maps (DSMs) have gained significant provenance, providing high-resolution maps through spatial extrapolation of observed data to regional, national and global scales. These maps are created by applying data-science methods to observational point data and associated covariates to create a predictive model. The model is used to extrapolate the prediction over the area for which covariate information is available. The predictive models often indicate impressively high levels of accuracy based on test/validation data. However, due to differences in both the range of data, methods and covariates used to drive predictive models, multiple DSMs created for the same areas are unlikely to be identical, which is indicative of the uncertainty associated with these mapped products. Much like with process-based models, there is a need to understand which data-science methodology is most suitable for a given research question and provide clarity on the magnitude of uncertainty associated with predictions. In this study, we quantify uncertainty in DSMs as a result of methodological choice; we apply several approaches (Random forest, Gaussian Process, Generalised Additive Model, Neural Network and Linear Regression) to create multiple predictive models of SOC concentration across the UK. By allowing the models to select from identical input data we provide a fair comparison of each approach through isolating uncertainty in DSMs as a result of methodological choice. In addition to accuracy assessment of each of the generated DSMs, we evaluate the suitability of each of these methods for DSM application. Most crucially, we highlight the need for caution in relation to the assumed levels of accuracy of generated DSMs when considering only standard validation statistics, and the limitations of these approaches when data has bi-modal distribution, a common feature of data that encompasses both mineral and organic soils. Whilst standard statistics evaluating the overall accuracy of the DSMs are highly significant, levels of accuracy across land use classifications vary considerably. Our study highlights the need for increased transparency in communication of uncertainty and limitations of derived map products.
The design and reanalysis of offshore and coastal structures usually requires the estimation of return values for dominant metocean variables (such as significant wave height) and associated values for other variables (such as peak spectral period or wind speed) from a finite sample of data; these are typically estimated using extreme value analysis. Yet the parameters of extreme value models can only be estimated with error from finite data. Different choices available to summarise uncertain information about the characteristics of the tail of a multivariate distribution in a small number of summary statistics (such as return values and associated values) complicates their estimation, especially for small sample sizes: choices regarding the ordering of mathematical operations lead to estimators of return values and associated values with different finite sample bias and variance characteristics. The current work extends a previous study (Jonathan et al.,2021) into the performance of estimators for marginal return values in the presence of sampling uncertainty, to estimators of associated values based on the bivariate conditional extremes model (Heffernan and Tawn, 2004) and competitors. Using a large designed simulation experiment, we explore the performance of combinations of 12 different estimators and three bivariate model candidates. The rich set of results from the simulation experiment are reported and explained in detail. Briefly: (a) calculation of associated values is only always feasible from small samples using two of the 12 estimators, which should be preferred; (b) estimators exploiting the median rather than the mean to summarise a distribution are more robust, and should also be preferred, especially for small sample sizes; (c) extreme value models incorporating appropriate descriptions of marginal and dependence provide better estimation of associated values for larger sample size; and (d) summarising the joint tail of metocean variables (in terms of return values and associated values) should be avoided where possible, in favour of probabilistic risk analysis of structural failure incorporating full uncertainty propagation.
Methods of computational statistics allow efficient estimation of extreme ocean environments, and facilitate optimal operational decision making. We describe estimation of extreme quantiles of total water level and related quantities from a non-stationary hierarchical model for ocean storms. The model incorporates a directional–seasonal extreme value model for occurrences of storm peak significant wave height, a conditional directional model for within-storm evolution of sea states relative to storm peak, a conditional model for the maximum crest within a sea state, and models for total water level. Importance sampling is used for efficient computation of marginal total water level characteristics. We use the model to estimate an optimal un-manning procedure for a notional North Sea offshore structure in severe conditions.
Digital technology is having a major impact on many areas of society, and there is equal opportunity for impact on science. This is particularly true in the environmental sciences as we seek to understand the complexities of the natural environment under climate change. This perspective presents the outcomes of a summit in this area, a unique cross-disciplinary gathering bringing together environmental scientists, data scientists, computer scientists, social scientists, and representatives of the creative arts. The key output of this workshop is an agreed vision in the form of a framework and associated roadmap, captured in the Windermere Accord. This accord envisions a new kind of environmental science underpinned by unprecedented amounts of data, with technological advances leading to breakthroughs in taming uncertainty and complexity, and also supporting openness, transparency, and reproducibility in science. The perspective also includes a call to build an international community working in this important area.
Reliable estimates of the occurrence rates of extreme events are highly important for insurance companies, government agencies and the general public. The rarity of an extreme event is typically expressed through its return period, i.e. the expected waiting time between events of the observed size if the extreme events of the processes are independent and identically distributed. A major limitation with this measure is when an unexpectedly high number of events occur within the next few months immediately after a T year event, with T large. Such instances undermine the trust in the quality of risk estimates. The clustering of apparently independent extreme events can occur as a result of local non-stationarity of the process, which can be explained by covariates or random effects. We show how accounting for these covariates and random effects provides more accurate estimates of return levels and aids short-term risk assessment through the use of a complementary new risk measure. Supplementary materials accompanying this paper appear online.
Decision‐making in flood risk management is increasingly dependent on access to data, with the availability of data increasing dramatically in recent years. We are therefore moving towards an era of big data, with the added challenges that, in this area, data sources are highly heterogeneous, at a variety of scales, and include a mix of structured and unstructured data. The key requirement is therefore one of integration and subsequent analyses of this complex web of data. This paper examines the potential of a data‐driven approach to support decision‐making in flood risk management, with the goal of investigating a suitable software architecture and associated set of techniques to support a more data‐centric approach. The key contribution of the paper is a cloud‐based data hypercube that achieves the desired level of integration of highly complex data. This hypercube builds on innovations in cloud services for data storage, semantic enrichment and querying, and also features the use of notebook technologies to support open and collaborative scenario analyses in support of decision making. The paper also highlights the success of our agile methodology in weaving together cross‐disciplinary perspectives and in engaging a wide range of stakeholders in exploring possible technological futures for flood risk management.
The concept 'models of everywhere' was first introduced in the mid 2000s as a means of reasoning about the environmental science of a place, changing the nature of the underlying modelling process, from one in which general model structures are used to one in which modelling becomes a learning process about specific places, in particular capturing the idiosyncrasies of that place. At one level, this is a straightforward concept, but at another it is a rich multi-dimensional conceptual framework involving the following key dimensions: models of everywhere, models of everything and models at all times, being constantly re-evaluated against the most current evidence. This is a compelling approach with the potential to deal with epistemic uncertainties and non-linearities. However, the approach has, as yet, not been fully utilised or explored. This paper examines the concept of models of everywhere in the light of recent advances in technology. The paper argues that, when first proposed, technology was a limiting factor but now, with advances in areas such as Internet of Things, cloud computing and data analytics, many of the barriers have been alleviated. Consequently, it is timely to look again at the concept of models of everywhere in practical conditions as part of a trans-disciplinary effort to tackle the remaining research questions. The paper concludes by identifying the key elements of a research agenda that should underpin such experimentation and deployment.
Multivariate extreme value models are used to estimate joint risk in a number of applications, with a particular focus on environmental fields ranging from climatology and hydrology to oceanography and seismic hazards. The semi-parametric conditional extreme value model of Heffernan and Tawn (2004) involving a multivariate regression provides the most suitable of current statistical models in terms of its flexibility to handle a range of extremal dependence classes. However, the standard inference for the joint distribution of the residuals of this model suffers from the curse of dimensionality since in a $d$-dimensional application it involves a $d-1$-dimensional non-parametric density estimator, which requires, for accuracy, a number points and commensurate effort that is exponential in $d$. Furthermore, it does not allow for any partially missing observations to be included and a previous proposal to address this is extremely computationally intensive, making its use prohibitive if the proportion of missing data is non-trivial. We propose to replace the $d-1$-dimensional non-parametric density estimator with a model-based copula with univariate marginal densities estimated using kernel methods. This approach provides statistically and computationally efficient estimates whatever the dimension, $d$ or the degree of missing data. Evidence is presented to show that the benefits of this approach substantially outweigh potential mis-specification errors. The methods are illustrated through the analysis of UK river flow data at a network of 46 sites and assessing the rarity of the 2015 floods in north west England.
Floods in England and Wales have the potential to cause billions of pounds of damage. You might think such extreme events are rare, but they are likely to occur more frequently than expected. By Ross Towe, Jonathan Tawn and Rob Lamb
Spatial extreme value analysis has been an area of rapid growth in the last decade. The focus has been on modelling the spatial componentwise maxima by max-stable processes. Here, we will explain the limitations of these modelling approaches and show how spatial models can be developed that overcome these deficiencies by exploiting the flexible conditional multivariate extremes models of Heffernan and Tawn (2004). We illustrate the benefits of these new spatial models through applications to North Sea wave analysis and to widespread UK river flood risk analysis.
For safe offshore operations, accurate knowledge of the extreme oceanographic conditions is required. We develop a multi-step statistical downscaling algorithm using data from a low resolution global climate model (GCM) and local-scale hindcast data to make predictions of the extreme wave climate in the next 50-year period at locations in the North Sea. The GCM is unable to produce wave data accurately so instead we use its 3-hourly wind speed and direction data. By exploiting the relationships between wind characteristics and wave heights, a downscaling approach is developed to relate the large and local-scale data sets, and hence future changes in wind characteristics can be translated into changes in extreme wave distributions. We assess the performance of the methods using within sample testing and apply the method to derive future design levels over the northern North Sea.
Ensemble is an interdisciplinary research team, working to explore the opportunity of new and emergent digital technologies in understanding, mitigating and adapting to environmental change. Using methods drawn from computer science, environmental science, social science, statistics, art, design and writing the team aim to transform the work of environmental scientists and decision makers, and the experience of communities by addressing themes related to complexity, uncertainty and abstraction of data. This paper discusses activities undertaken within a research ‘sprint’ directed at addressing flood risk management through data driven decision making, communication and community engagement.
Widespread flooding, such as the events in the winter of 2013/2014 in the UK and early summer 2013 in Cent ral Europe, demonst rate clearly how important it is to understand the characterist ics of floods in which mult iple locat ions experience ext reme river flows. Recent developments in mult ivariate stat ist ical modelling help to place such events in a probabilist ic framework. It is now possible to perform joint probability analysis of events defined in terms of physical variables at hundreds of locat ions simultaneously, over mult iple variables (including river flows, rainfall and sea levels), combined with analysis of temporal dependence to capture the evolut ion of events over a large domain. Crit ical const raints on such data-driven methods are the problems of missing data, especially where records over a network are not all concurrent , the joint analysis of several different physical variables, and the choice of suitable t ime scales when combining informat ion from those variables. This paper presents new developments of a high-dimensional condit ional probability model for ext reme river flow events condit ioned on flow and r ainfall observat ions. These are: a new computat ionally efficient paramet ric approach to account for missing data in the joint analysis of ext remes over a large hydromet ric network; a robust approach for the spat ial interpolation of extreme events throughout a large river network,; generat ion of realist ic est imates of ext remes at ungauged locat ions; and, exploit ing rainfall information rat ionally within the stat ist ical model to help improve efficiency. These methodological advances will be illust rated with data from the UK river network and recent events to show how they cont ribute to a flexible and effective framework for flood risk assessment, with applicat ions in the insurance sector and for nat ional-scale emergency planning.
Characterising the joint distribution of extremes of significant wave height and wind speed is critical for reliable design and assessment of marine structures. The extremal dependence of pairs of oceanographic variables can be characterised using one of a number of summary statistics, which describe the two different types of extremal dependence. Quantifying the type of extremal dependence is an essential pre-requisite to joint or spatial extreme value modelling, and ensures that appropriate model forms are employed.We estimate extremal dependence between storm peak significant wave height and storm peak wind speed (He, WS) for locations in a region of the northern North Sea. However, since the extremal dependence itself may vary with storm direction, we introduce new covariate-dependent forms of the extremal dependence measures that account for the direction of the storm.We discuss the implications of all of the estimates for marine design, including specification of joint design criteria for extended spatial domains, and statistical downscaling to incorporate the effects of climate change on design specification.
Discussion of "Statistical Modeling of Spatial Extremes" by A. C. Davison, S. A. Padoan and M. Ribatet [arXiv:1208.3378].