Complex computer codes or models can often be run in a hierarchy of different levels of complexity ranging from the very basic to the sophisticated. The top levels in this hierarchy are typically expensive to run, which limits the number of possible runs. To make use of runs over all levels, and crucially improve predictions at the top level, we use multi-level Gaussian process emulators (GPs). The accuracy of the GP greatly depends on the design of the training points. In this paper, we present a multi-level adaptive sampling algorithm to sequentially increase the set of design points to optimally improve the fit of the GP. The normalised expected leave-one-out cross-validation error is calculated at all unobserved locations, and a new design point is chosen using expected improvement combined with a repulsion function. This criterion is calculated for each model level weighted by an associated cost for the code at that level. Hence, at each iteration, our algorithm optimises for both the new point location and the model level. The algorithm is extended to batch selection as well as single point selection, where batches can be designed for single levels or optimally across all levels.
CIBSE weather files are currently used by the building industry as the standard input data for building performance assessment for the purpose of regulatory compliance in the UK. In this study, the state-of-the-art CIBSE weather files are created with four major improvements incorporated, namely, (1) the enhanced representation of the UK climate through the creation of discriminative climate zones; (2) the latest climate change signals from the UK Climate Projection 2018 (UKCP18); (3) the satellite based solar radiation data from CAMS (Copernicus Atmosphere Monitoring Service) data repository; (4) the up-to-date observation record from 1994 to 2023. The methodology for creating the latest CIBSE weather files is elaborated in detail to enhance the transparency of the new weather data. Evaluated using a simulation case study, the new weather files demonstrate spatial and temporal coherency. The new future weather files enable robust building performance assessment against future climate conditions under different scenarios and will play an important role in designing climate-resilient buildings and delivering a net zero built environment.Practical applications As per the principle of "garbage in, garbage out", weather data plays an instrumental role in streamlining building design to achieve both energy efficiency and thermal comfort. In this study, we present the methodology for the creation of the state-of-the-art CIBSE weather files. The new CIBSE weather files not only employ the update-to-date observation and projection data, but are also grounded on a total of 28 granular climate zones to account for diverse climate characteristics and eliminate the ambiguity with weather data selection. The new files will lay a solid data foundation for future-proofing building design in the UK.
Decision making often uses complex computer codes run at the exa-scale (10e18 flops). Such computer codes or models are often run in a hierarchy of different levels of fidelity ranging from the basic to the very sophisticated. The top levels in this hierarchy are expensive to run, limiting the number of possible runs. To make use of runs over all levels, and crucially improve emulation at the top level, we use multi-level Gaussian process emulators (GPs). We will present a new method of building GP emulators from hierarchies of models. In order to share information across the different levels, l=1,...,L, we define the form of the prior of the l+1th level to be the posterior of the lth level, hence building a Bayesian hierarchical structure for the top Lth level. This enables us to not only learn about the GP hyperparameters as we move up the multi-level hierarchy, but also allows us to limit the total number of parameters in the full model, whilst maintaining accuracy.
Complex numerical models are increasingly being used in healthcare and epidemiology. To represent the complex features, modellers often make the decision to include stochastic behaviour where repeated runs of the model with identical inputs produce different outputs. When computational constraints limit the number of model runs and replications, heteroscedastic Gaussian processes can be used as a fast surrogate, allowing for efficient emulation of varying noise levels across the input space. The accuracy of any emulator is greatly dependent on the design of the training data, where sequential design algorithms increase the number of design points iteratively based on predefined criteria. For stochastic models, the design problem is more challenging due to the possibility of replicates at design points. This article develops a new sequential design method for stochastic models which scales well in high-dimensional input spaces. We build upon an existing method for deterministic models using an expected squared leave-one-out error criterion that balances exploration and replication. We compare our approach with existing sequential design methods as well as applying it to an agent-based model and a COVID-19 model. Results demonstrate that the proposed method performs well in noisy environments, offering a scalable alternative to existing methods.This article is part of the theme issue 'Uncertainty quantification for healthcare and biological systems (Part 2)'.
The ocean biological carbon pump, a significant set of processes in the global carbon cycle, drives the sinking of particulate organic carbon (POC) towards the deep ocean. Global estimates of POC fluxes and an improved understanding of how environmental factors influence organic ocean carbon transport can help quantify how much carbon is sequestered in the ocean and how this can change in different environmental conditions, in addition to improving global carbon and marine ecosystem models. POC fluxes can be derived from observations taken by a variety of in situ instruments such as sediment traps, 234-Thorium tracers and Underwater Vision Profilers. However, the manual and time-consuming nature of data collection leads to limitations of spatial data sparsity on a global scale, resulting in large estimate uncertainties in under-sampled regions.This research takes an observation-driven approach with machine learning and statistical models trained to estimate POC fluxes on a global scale using the in situ observations and well-sampled environmental driver datasets, such as temperature and nutrient concentrations. This approach holds two main benefits: 1) the ability to fill observational gaps on both a spatial and temporal scale and 2) the opportunity to interpret the importance of each environmental factor for estimating POC fluxes, and therefore exposing their relationship to organic carbon transport processes. The models built include random forests, neural networks and Bayesian hierarchical models, where their global POC flux estimates, feature importance and model performances are studied and compared. Additionally, this research explores the use of data fusion methods to combine all three heterogeneous in situ POC flux data sources to achieve improved accuracy and better-informed inferences about organic carbon transport than what is possible using a single data source. By treating the heterogeneous data sources differently, accounting for their biases, and introducing domain knowledge into the models, our data fusion method can not only harness the information from all three data sources, but also gives insights into their key differences.
Weather files are fundamental for building performance applications such as estimating energy demand or predicting the risk of thermal discomfort. Often such weather files are based on specific locations with limited guidance to which file to use which can lead to non-representative climate conditions for a specific building and potentially under or over engineering of systems. Recently, a new approach to generating representative weather files for building performance using climate zones has been proposed. This work presents a comparative analysis of the two approaches. Results from examining the external climate and an example building simulation reveal substantial differences: heating degree days vary by up to 931 degrees days, cooling degree days by up to 500 degrees days, heating energy demand differs by as much as 60%, and overheating exposure in the bedroom exceeds 26 degrees C for up to 45 h more. Overall, we find significant variations in heating and cooling loads can be expected if unrepresentative climate data is used highlighting the importance of selecting appropriate weather data for building performance simulation. The use of zone-based files provides a more nuanced understanding of climate conditions, improving the reliability of building performance assessments and compliance with overheating standards.Practical application Weather files are fundamental for building performance analysis and supporting design decisions including the prediction of the risk of thermal discomfort. There is a need for these weather files to be kept up to date following the continual warming of the environment to ensure that the performance analysis represents external conditions and ensuring the analysis is fit for purpose. A significant limitation of the current approach is the uncertainty of which weather file should be used for any given location particularly in locations such as the UK. This work provides a framework for the development of climate-zones and representative weather files suitable for supporting the benchmarking of building performance across the climate zone removing ambiguity within the industry for the appropriate selection of weather data.
We explore the application of uncertainty quantification methods to agent-based models (ABMs) using a simple sheep and wolf predator-prey model. This work serves as a tutorial on how techniques like emulation can be powerful tools in this context. We also highlight the importance of advanced statistical methods in effectively utilising computationally expensive ABMs. Specifically, we implement stochastic Gaussian processes, Gaussian process classification, sequential design, and history matching to address uncertainties in model input parameters and outputs. Our results show that these methods significantly enhance the robustness, accuracy, and predictive power of ABMs.
Fires occurring over the peatlands in Indonesian Borneo accompanied by droughts have posed devastating impacts on human health, livelihoods, economy and the natural environment, and their prevention requires comprehensive understanding of climate-associated risk. Although it is widely known that the droughts are associated with El Nino events, the onset process of El Nino and thus the drought precursors and their possible changes under the future climate are not clearly understood. Here, we use a causal network approach to quantify the strength of teleconnections to droughts at a seasonal timescale shown in observations and climate models. We portray two drivers of June-July-August (JJA) droughts identified through literature review and causal analysis, namely Nino 3.4 sea surface temperature (SST) in JJA (El Nino Southern Oscillation [abbreviated as ENSO]) and SST anomaly over the eastern North Pacific to the east of the Hawaiian Islands (abbreviated as Pacific SST) in March-April-May (MAM) period. We argue that the droughts are strongly linked to ENSO variability, with drier years corresponding to El Nino conditions. The droughts can be predicted with a lead time of 3 months based on their associations with Pacific SST, with higher SST preceding drier conditions. We find that under the SSP585 scenario, the Coupled Model Intercomparison Phase 6 (CMIP6) multi-model ensembles show significant increase in both the maximum number of consecutive dry days in the Indonesian Borneo region in JJA (p = 0.006) and its linear association with Pacific SST in MAM (p = 0.001) from year 2061 to 2100 compared with the historical baseline. Some models are showing unrealistic amounts of JJA rainfall and underestimate drought risks in Indonesian Borneo and their teleconnections, owing to the underestimation of ENSO amplitude and overestimation of local convections. Our study strengthens the possibility of early warning triggers of fires and stresses the need for taking enhanced climate risk into consideration when formulating long-term policies to eliminate fires. We have identified some predictors (shown in the schematic diagram and equations) that are responsible for the drought and fire risks during boreal summer in Indonesian Borneo. We have also looked at climate model projections and detected the enhancing risks towards the future (shown in the boxplots). These findings are crucial for seasonal forecasting of fires, which enables early preventive actions, and foster long-term resilience building to eradicate fire occurrences. image
Climate zones play an important role in promoting climate responsive building design and implementing climate-specific prescriptions in national building standards and regulations. The existing studies on climate zoning are subject to several limitations, i.e. the incapability of distinguishing microclimates and the lack of consideration of climate change. In this research, we propose a two-tiered ensemble clustering method for the identification of granular climate zones using the projections of future climate. The first tier identifies primary climate zones using a combination of climatic features and geographical locations, whereas the second tier identifies distinct local variations within each primary climate zone using the temperature related features. The proposed ensemble clustering model is applied to the UK to create a mapping of granular climate zones for future proofing building design. The method identified 14 distinct primary zones and distinguished microclimates at a range of scales from large urban areas, such as the Greater London Area, to national parks, such as Dartmoor and the Pennines. The identified mapping resolves two major obstacles in the creation and usage of weather data for building performance assessment in the UK, i.e. the lack of guidance for selecting weather files, and the absence of scientific rationale for representing the UK climate using the current 14 locations.
Global warming and net zero transition are the two biggest challenges currently faced by the building industry in the UK. While the net zero transition primarily focuses on the problems of energy efficiency and heat decarbonization, the rise of global temperature imposes a significant threat to the health and wellbeing of occupants and the industry is obliged to make buildings climate-resilient by testing their designs using future weather files. To improve the quality of the current weather files, a new project has been commissioned by CIBSE to revisit the data and the methodology employed for creating future weather files and produce new CIBSE weather files using the latest UK Climate Projections released in 2018 (UKCP18). In this study, we evaluate the newly produced weather files for overheating risk using building simulation. Two different batches of weather files were curated. The first batch was produced primarily using the existing methodology for creating the UKCP09 based weather files, with an adjustment to accommodate new features of the UKCP18 and an improved procedure for morphing the solar radiation data. The second batch was created through an improved morphing process to better emulate the characteristics of distributions of climatic variables. The differences between the existing UKCP09 and new UCKP18 based weather files are compared by evaluating overheating metrics. The new weather files enable robust building performance assessment against future climate conditions under different scenarios and will play an important role in designing climate-resilient buildings and delivering a net zero built environment. Practical applications As the extreme weather events resulting from climate change become more frequent and intense, they pose significant challenges to the resilience of the built environment and severe threats to the health and wellbeing of the occupants. Climate data, which serves as the foundation for climate risk assessment, plays a critical role in helping the building sector to achieve climate resilience through the means of performance assessment and the channel of regulatory compliance. In this study, the revised future weather files created using the latest UKCP18 climate projections are presented and evaluated using building simulation, as part of the weather file testing programme for quality assurance. The revision of the CIBSE weather files according to the latest climate science, i.e. UKCP18, will enable the building industry to quantify overheating risks with more accurate climate assumptions and better inform decision making about risk mitigation and climate adaptation.
Gaussian processes (GPs) are generally regarded as the gold standard surrogate model for emulating computationally expensive computer-based simulators. However, the problem of training GPs as accurately as possible with a minimum number of model evaluations remains challenging. We address this problem by suggesting a novel adaptive sampling criterion called VIGF (variance of improvement for global fit). The improvement function at any point is a measure of the deviation of the GP emulator from the nearest observed model output. At each iteration of the proposed algorithm, a new run is performed where VIGF is the largest. Then, the new sample is added to the design and the emulator is updated accordingly. A batch version of VIGF is also proposed which can save the user time when parallel computing is available. Additionally, VIGF is extended to the multi-fidelity case where the expensive high-fidelity model is predicted with the assistance of a lower fidelity simulator. This is performed via hierarchical kriging. The applicability of our method is assessed on a bunch of test functions and its performance is compared with several sequential sampling strategies. The results suggest that our method has a superior performance in predicting the benchmark functions in most cases. An implementation of VIGF is available in the dgpsi R package, which can be found on CRAN.
A Gaussian process (GP)-based methodology is proposed to emulate complex dynamical computer models (or simulators). The method relies on emulating the numerical flow map of the system over an initial (short) time step, where the flow map is a function that describes the evolution of the system from an initial condition to a subsequent value at the next time step. This yields a probabilistic distribution over the entire flow map function, with each draw offering an approximation to the flow map. The model output times series is then predicted (under the Markov assumption) by drawing a sample from the emulated flow map (i.e., its posterior distribution) and using it to iterate from the initial condition ahead in time. Repeating this procedure with multiple such draws creates a distribution over the time series. The mean and variance of this distribution at a specific time point serve as the model output prediction and the associated uncertainty, respectively. However, drawing a GP posterior sample that represents the underlying function across its entire domain is computationally infeasible, given the infinite-dimensional nature of this object. To overcome this limitation, one can generate such a sample in an approximate manner using random Fourier features (RFF). RFF is an efficient technique for approximating the kernel and generating GP samples, offering both computational efficiency and theoretical guarantees. The proposed method is applied to emulate several dynamic nonlinear simulators including the well-known Lorenz and van der Pol models. The results suggest that our approach has a promising predictive performance and the associated uncertainty can capture the dynamics of the system appropriately.
Mathematical models are widely used in simulating and analysing the brain's connectivity and dynamics. Within these models, coupled Hopf oscillators are employed to describe the state of a brain network model changes over time and account for the interactions between brain regions. Gaussian Process Emulators provide a cheap and efficient framework to approximate the simulator and produce the prediction of simulator output with associated uncertainty. This paper extends the traditional static Gaussian Process Emulator to a dynamic version and is used to predict the dynamic evolution of the brain network model and quantify its uncertainties. We present different topology of the brain network to study their dynamic behaviours over time.
Probabilistic graphical models are widely used to model complex systems under uncertainty. Traditionally, Gaussian directed graphical models are applied for analysis of large networks with continuous variables as they can provide conditional and marginal distributions in closed form simplifying the inferential task. The Gaussianity and linearity assumptions are often adequate, yet can lead to poor performance when dealing with some practical applications. In this paper, we model each variable in graph G as a polynomial regression of its parents to capture complex relationships between individual variables and with a utility function of polynomial form. We develop a message-passing algorithm to propagate information throughout the network solely using moments which enables the expected utility scores to be calculated exactly. Our propagation method scales up well and enables to perform inference in terms of a finite number of expectations. We illustrate how the proposed methodology works with examples and in an application to decision problems in energy planning and for real-time clinical decision support.
<p>Fires occurring over the peatlands in Indonesian Borneo accompanied by droughts have posed devastating impacts on human health, livelihoods, economy and the natural environment, and their prevention requires a comprehensive understanding of climate-associated risk. We want to strengthen the possibility of early warning triggers of drought, which is a strong predictor of the prevalence of fires, and evaluate the climate risk relevant to the formulation of long-term policies to eliminate fires. Although it is widely known that the droughts are often associated with El Ni&#241;o events, the onset process of El Ni&#241;o and thus the drought precursors and their possible changes under the future climate are not clearly understood. Here we use a causal network approach to quantify the strength of teleconnections to droughts at a seasonal timescale shown in (1) observational and reanalysis data (2) CMIP6 models and (3) seasonal hindcasts. We consider two drivers of JJA droughts identified through literature review and causal analysis, namely Ni&#241;o 3.4 SST in JJA (abbreviated as ENSO) and SST anomaly over the eastern North Pacific to the east of the Hawaiian Islands (abbreviated as Pacific SST) in MAM. The observational and reanalysis data proves that the droughts are strongly linked to ENSO variability, with drier years corresponding to El Ni&#241;o conditions, and droughts can be predicted with a lead time of three months based on their associations with Pacific SST, with higher SST preceding drier conditions. We find that some CMIP6 models are showing unrealistic amounts of JJA rainfall and underestimate drought risks in Indonesian Borneo and their teleconnections, owing to the underestimation of ENSO amplitude and overestimation of local convections. Under the SSP585 scenario, the CMIP6 multi-model ensembles show significant increase in both the maximum number of consecutive dry days in the Indonesian Borneo region in JJA (p = 0.006) and its linear association with Pacific SST in MAM (p = 0.001) from year 2061 &#8211; 2100 compared with the historical baseline. On the other hand, seasonal hindcast models are (1) overestimating the variability of maximum number of consecutive dry days, (2) showing varied skills in simulating the mean rainfall and drought indicators, and (3) underestimating the teleconnections to Borneo droughts, making it difficult to assess the likelihood of unprecedented drought and fire risk under El Ni&#241;o conditions. Our study agrees with previous studies regarding the limited skill of fire risk prediction by state-of-the-art seasonal forecasting models, with their shortfalls caused by a lack of proper representation of relevant teleconnections.</p>
Sensitivity analysis plays an important role in the development of computer models/simulators through identifying the contribution of each (uncertain) input factor to the model output variability. This report investigates different aspects of the variance-based global sensitivity analysis in the context of complex black-box computer codes. The analysis is mainly conducted using two R packages, namely sensobol (Puy et al., 2021) and sensitivity (Iooss et al., 2021). While the package sensitivity is equipped with a rich set of methods to conduct sensitivity analysis, especially in the case of models with dependent inputs, the package sensobol offers a bunch of user-friendly tools for the visualisation purposes. Several illustrative examples are supplied that allow the user to learn both packages easily and benefit from their features.
The estimation of parameters and model structure for informing infectious disease response has become a focal point of the recent pandemic. However, it has also highlighted a plethora of challenges remaining in the fast and robust extraction of information using data and models to help inform policy. In this paper, we identify and discuss four broad challenges in the estimation paradigm relating to infectious disease modelling, namely the Uncertainty Quantification framework, data challenges in estimation, model-based inference and prediction, and expert judgement. We also postulate priorities in estimation methodology to facilitate preparation for future pandemics.
Gaussian Process (GP) emulators are widely used to approximate complex computer model behaviour across the input space. Motivated by the problem of coupling computer models, recently progress has been made in the theory of the analysis of networks of connected GP emulators. In this paper, we combine these recent methodological advances with classical state-space models to construct a Bayesian decision support system. This approach gives a coherent probability model that produces predictions with the measure of uncertainty in terms of two first moments and enables the propagation of uncertainty from individual decision components. This methodology is used to produce a decision support tool for a UK county council considering low carbon technologies to transform its infrastructure to reach a net-zero carbon target. In particular, we demonstrate how to couple information from an energy model, a heating demand model, and gas and electricity price time-series to quantitatively assess the impact on operational costs of various policy choices and changes in the energy market.
In many real-world applications, we are interested in approximating black-box, costly functions as accurately as possible with the smallest number of function evaluations. A complex computer code is an example of such a function. In this work, a Gaussian process (GP) emulator is used to approximate the output of complex computer code. We consider the problem of extending an initial experiment (set of model runs) sequentially to improve the emulator. A sequential sampling approach based on leave-one-out (LOO) cross-validation is proposed that can be easily extended to a batch mode. This is a desirable property since it saves the user time when parallel computing is available. After fitting a GP to training data points, the expected squared LOO (ES-LOO) error is calculated at each design point. ES-LOO is used as a measure to identify important data points. More precisely, when this quantity is large at a point it means that the quality of prediction depends a great deal on that point and adding more samples nearby could improve the accuracy of the GP. As a result, it is reasonable to select the next sample where ES-LOO is maximized. However, ES-LOO is only known at the experimental design and needs to be estimated at unobserved points. To do this, a second GP is fitted to the ES-LOO errors, and where the maximum of the modified expected improvement (EI) criterion occurs is chosen as the next sample. EI is a popular acquisition function in Bayesian optimization and is used to trade off between local and global search. However, it has a tendency towards exploitation, meaning that its maximum is close to the (current) ``best"" sample. To avoid clustering, a modified version of EI, called pseudoexpected improvement, is employed which is more explorative than EI yet allows us to discover unexplored regions. Our results show that the proposed sampling method is promising.
Binary time series data are very common in many applications, and are typically modelled independently via a Bernoulli process with a single probability of success. However, the probability of a success can be dependent on the outcome successes of past events. Presented here is a novel approach for modelling binary time series data called a binary de Bruijn process which takes into account temporal correlation. The structure is derived from de Bruijn Graphs - a directed graph, where given a set of symbols, V, and a 'word' length, m, the nodes of the graph consist of all possible sequences of V of length m. De Bruijn Graphs are equivalent to mth order Markov chains, where the 'word' length controls the number of states that each individual state is dependent on. This increases correlation over a wider area. To quantify how clustered a sequence generated from a de Bruijn process is, the run lengths of letters are observed along with run length properties. Inference is also presented along with two application examples: precipitation data and the Oxford and Cambridge boat race.