Reliable outputs from physics-based models are essential for a wide range of scientific and operational applications, such as energy systems. However, physics-based numerical models often produce systematically inconsistent outputs due to variations in spatial resolution, parameterizations, boundary conditions, and numerical schemes. While these discrepancies pose challenges, the diversity among models also offers complementary insights into the same physical processes. Integrating outputs across multiple models rather than relying on a single source can therefore yield more robust and informative results. This study presents a general two-step framework for correcting and integrating multi-model outputs with observations. First, a copula-based bias correction method is applied to align both the marginal distributions and dependence structure of each model with observational data, using the Bulk-and-Tails (BATs) distribution to capture complex behaviors such as asymmetry and tail dependence. Second, Bayesian Model Averaging (BMA) is used to probabilistically combine the bias-corrected outputs, assigning weights based on each model’s historical predictive performance. The resulting ensemble preserves key statistical features, quantifies uncertainty, and enhances physical realism. This approach improves the fidelity and interpretability of integrated model outputs and is broadly applicable to domains requiring model-data fusion, including Earth system science, infrastructure risk assessment, and environmental monitoring.
Power grid expansion planning requires making large investment decisions in the present that will impact the future cost and reliability of a system exposed to wide-ranging uncertainties. Extreme temperatures can pose significant challenges to providing power by increasing demand and decreasing supply and have contributed to recent major power outages. We propose to address a modeling challenge of such high-impact, low-frequency events with a bi-objective stochastic integer optimization model that finds solutions with different trade-offs between efficiency in normal conditions and risk to extreme events. We propose a conditional sampling approach paired with a risk measure to address the inherent challenge in approximating the risk of low-frequency events within a sampling based approach. We present a model for spatially correlated, county-specific temperatures and a method to generate both unconditional and conditionally extreme temperature samples from this model efficiently. These models are investigated within an extensive case study with realistic data that demonstrates the effectiveness of the bi-objective approach and the conditional sampling technique. We find that spatial correlations in the temperature samples are essential to finding good solutions and that modeling generator temperature dependence is an important consideration for finding efficient, low-risk solutions.
Inverter-based resources (IBRs) data retrieved from physics-based models, such as wind speed, generally suffer from statistical biases in distribution. These biases result from various factors such as simplifications in the model physics, inaccuracies in input data or uncertainties in model parameterizations. Thus, we propose a copula-based bivariate bias correction method for wind speed from physics-based models to improve their reliability and enhance our understanding of downscaled physics-based model data. The proposed statistical method is applied on hourly wind speed data at 10m and 100m heights in Argonne, IL. We compare the method with the commonly applied benchmark models in the literature. The proposed copula-based method outperforms benchmark models and reduces the bias up to 20%.
Abstract Accurate quantification of uncertainty in neural network predictions remains a central challenge for scientific applications involving high‐dimensional, correlated data. While existing methods capture either aleatoric or epistemic uncertainty, few offer closed‐form, multidimensional distributions that preserve spatial correlation while remaining computationally tractable. In this work, we present a framework for training neural networks with a multidimensional Gaussian loss, generating a closed‐form predictive distribution over outputs informed by non‐identically distributed training data. Our approach captures aleatoric uncertainty by iteratively estimating the means and covariance matrices, and is demonstrated on a super‐resolution example out‐of‐training‐sample. We leverage a Fourier representation of the covariance matrix to stabilize network training and preserve spatial correlation. We introduce a novel regularization strategy—referred to as information sharing—that interpolates between image‐specific and global covariance estimates, enabling convergence of the super‐resolution downscaling network trained on image‐specific distributional loss functions. This framework allows for efficient sampling, explicit correlation modeling, and extensions to more complex distribution families all without disrupting prediction performance. We demonstrate the method on a surface wind speed downscaling task and discuss its broader applicability to uncertainty‐aware prediction in scientific models.
Combining strengths from deep learning and extreme value theory can help describe complex relationships between variables where extreme events have significant impacts (e.g., environmental or financial applications). Neural networks learn complicated nonlinear relationships from large datasets under limited parametric assumptions. By definition, the number of occurrences of extreme events is small, which limits the ability of the data-hungry, nonparametric neural network to describe rare events. Inspired by recent extreme cold winter weather events in North America caused by atmospheric blocking, we examine several probabilistic generative models for the entire multivariate probability distribution of daily boreal winter surface air temperature. We propose metrics to measure spatial asymmetries, such as long-range anticorrelated patterns that commonly appear in temperature fields during blocking events. Compared to vine copulas, the statistical standard for multivariate copula modeling, deep learning methods show improved ability to reproduce complicated asymmetries in the spatial distribution of ERA5 temperature reanalysis, including the spatial extent of in-sample extreme events.
We propose a new modeling framework for highly-multivariate spatial processes that synthesizes ideas from recent multiscale and spectral approaches with graphical models. The basis graphical lasso writes a univariate Gaussian process as a linear combination of basis functions weighted with entries of a Gaussian graphical vector whose graph is estimated from optimizing an $\ell_1$ penalized likelihood. This paper extends the setting to a multivariate Gaussian process where the basis functions are weighted with Gaussian graphical vectors. We motivate a model where the basis functions represent different levels of resolution and the graphical vectors for each level are assumed to be independent. Using an orthogonal basis grants linear complexity and memory usage in the number of spatial locations, the number of basis functions, and the number of realizations. An additional fusion penalty encourages a parsimonious conditional independence structure in the multilevel graphical model. We illustrate our method on a large climate ensemble from the National Center for Atmospheric Research's Community Atmosphere Model that involves 40 spatial processes.
Current models for spatial extremes are concerned with the joint upper (or lower) tail of the distribution at two or more locations. Such models cannot account for teleconnection patterns of 2-m surface air temperature (T2m) in North America, where very low temperatures in the contiguous United States may coincide with very high temperatures in Alaska in the wintertime. This dependence between warm and cold extremes motivates the need for a model with opposite-tail dependence in spatial extremes. This work develops a statistical modeling framework that has flexible behav-ior in all four pairings of high and low extremes at pairs of locations. In particular, we use a mixture of rotations of common Archimedean copulas to capture various combinations of four-corner tail dependence. We study teleconnected T2m ex-tremes using ERA5 of daily average 2-m temperature during the boreal winter. The estimated mixture model quantifies the strength of opposite-tail dependence between warm temperatures in Alaska and cold temperatures in the midlatitudes of North America, as well as the reverse pattern. These dependence patterns are shown to correspond to blocked and zonal patterns of midtropospheric flow. This analysis extends the classical notion of correlation-based teleconnections to consid-ering dependence in higher quantiles.
Current models for spatial extremes are concerned with the joint upper (or lower) tail of the distribution at two or more locations. Such models cannot account for teleconnection patterns of two-meter surface air temperature ($T_{2m}$) in North America, where very low temperatures in the contiguous Unites States (CONUS) may coincide with very high temperatures in Alaska in the wintertime. This dependence between warm and cold extremes motivates the need for a model with opposite-tail dependence in spatial extremes. This work develops a statistical modeling framework which has flexible behavior in all four pairings of high and low extremes at pairs of locations. In particular, we use a mixture of rotations of common Archimedean copulas to capture various combinations of four-corner tail dependence. We study teleconnected $T_{2m}$ extremes using ERA5 reanalysis of daily average two-meter temperature during the boreal winter. The estimated mixture model quantifies the strength of opposite-tail dependence between warm temperatures in Alaska and cold temperatures in the midlatitudes of North America, as well as the reverse pattern. These dependence patterns are shown to correspond to blocked and zonal patterns of mid-tropospheric flow. This analysis extends the classical notion of correlation-based teleconnections to considering dependence in higher quantiles.
In traditional extreme value analysis, the bulk of the data is ignored, and only the tails of the distribution are used for inference. Extreme observations are specified as values that exceed a threshold or as maximum values over distinct blocks of time, and subsequent estimation procedures are motivated by asymptotic theory for extremes of random processes. For environmental data, nonstationary behavior in the bulk of the distribution, such as seasonality or climate change, will also be observed in the tails. To accurately model such nonstationarity, it seems natural to use the entire dataset rather than just the most extreme values. It is also common to observe different types of nonstationarity in each tail of a distribution. Most work on extremes only focuses on one tail of a distribution, but for temperature, both tails are of interest. This paper builds on a recently proposed parametric model for the entire probability distribution that has flexible behavior in both tails. We apply an extension of this model to historical records of daily mean temperature at several locations across the United States with different climates and local conditions. We highlight the ability of the method to quantify changes in the bulk and tails across the year over the past decades and under different geographic and climatic conditions. The proposed model shows good performance when compared to several benchmark models that are typically used in extreme value analysis of temperature.
In addition to providing water for nearly 2 billion people, snow drives resource selection by wildlife and influences the behavior and demography of many species. Because snow cover is highly spatially and temporally variable, mapping its extent using currently available satellite data remains a challenge. At present, there are no sensors acquiring daily data of Earth's entire surface at fine spatial resolutions (< 30 m) in wavelengths required for snow cover retrieval, namely: visible, near-infrared, and shortwave infrared. Fine scale observations at 30 m from Landsat are available at 16-day intervals since 1982 and at 8-day intervals since 1999. However, over this duration, snow can accumulate, ablate, or both, making the Landsat data ineffective for many applications. Conversely, the Moderate Resolution Imaging Spectroradiometer (MODIS) atmospherically corrected daily reflectance data, have a coarse spatial resolution of 463 m and thus, are not ideal for snow cover mapping either. This spatial and temporal resolution tradeoff limits the use of these data for a wide range of snow cover applications and indicates a pressing need for data fusion. To address this need, we use a physically-based, spectral-mixture-analysis approach for mapping fractional snow cover (fSCA) and a two-stage random forest algorithm to produce daily 30 m fSCA. We test our algorithm in the US Sierra Nevada and find MODIS fSCA is the most important predictor. We cross validate using 170 Landsat scenes and while snow cover varies immensely in time we find little variation in errors between seasons, a small bias of 0.01, and an overall accuracy of 0.97 with slightly higher precision than recall. This technique for accurate, daily, high-resolution snow cover retrievals could be applied more broadly for analyses of regional energy budget, validating snow cover in global and regional models, and for quantifying changes in the availability of biotic resources in ecosystems.
Many modern spatial models express the stochastic variation component as a basis expansion with random coefficients. Low rank models, approximate spectral decompositions, multiresolution representations, stochastic partial differential equations, and empirical orthogonal functions all fall within this basic framework. Given a particular basis, stochastic dependence relies on flexible modeling of the coefficients. Under a Gaussianity assumption, we propose a graphical model family for the stochastic coefficients by parameterizing the precision matrix. Sparsity in the precision matrix is encouraged using a penalized likelihood framework—we term this approach the basis graphical lasso. Computations follow from a majorization-minimization (MM) approach, a byproduct of which is a connection to the standard graphical lasso. The result is a flexible nonstationary spatial model that is adaptable to very large datasets with multiple realizations. We apply the model to two large and heterogeneous spatial datasets in statistical climatology and recover physically sensible graphical structures. Moreover, the model performs competitively against the popular LatticeKrig model in predictive cross-validation but improves the Akaike information criterion score and a log score for the quality of the joint predictive distribution.
Many modern spatial models express the stochastic variation component as a basis expansion with random coefficients. Low rank models, approximate spectral decompositions, multiresolution representations, stochastic partial differential equations and empirical orthogonal functions all fall within this basic framework. Given a particular basis, stochastic dependence relies on flexible modeling of the coefficients. Under a Gaussianity assumption, we propose a graphical model family for the stochastic coefficients by parameterizing the precision matrix. Sparsity in the precision matrix is encouraged using a penalized likelihood framework. Computations follow from a majorization-minimization approach, a byproduct of which is a connection to the graphical lasso. The result is a flexible nonstationary spatial model that is adaptable to very large datasets. We apply the model to two large and heterogeneous spatial datasets in statistical climatology and recover physically sensible graphical structures. Moreover, the model performs competitively against the popular LatticeKrig model in predictive cross-validation, but substantially improves the Akaike information criterion score.
Geophysical and other natural processes often exhibit non-stationary covariances and this feature is important to take into account for statistical models that attempt to emulate the physical process. A convolution-based model is used to represent non-stationary Gaussian processes that allows for variation in the correlation range and vari- ance of the process across space. Application of this model has two steps: windowed estimates of the covariance function under the as- sumption of local stationary and encoding the local estimates into a single spatial process model that allows for efficient simulation. Specifically we give evidence to show that non-stationary covariance functions based on the Mat`ern family can be reproduced by the Lat- ticeKrig model, a flexible, multi-resolution representation of Gaussian processes. We propose to fit locally stationary models based on the Mat`ern covariance and then assemble these estimates into a single, global LatticeKrig model. One advantage of the LatticeKrig model is that it is efficient for simulating non-stationary fields even at 105 locations. This work is motivated by the interest in emulating spatial fields derived from numerical model simulations such as Earth system models. We successfully apply these ideas to emulate fields that de- scribe the uncertainty in the pattern scaling of mean summer (JJA) surface temperature from a series of climate model experiments. This example is significant because it emulates tens of thousands of loca- tions, typical in geophysical model fields, and leverages embarrassing parallel computation to speed up the local covariance fitting
We extend the work of Robinson and Turner to use hypothesis testing with persistence homology to test for measurable differences in shape between point clouds from three or more groups. Using samples of point clouds from three distinct groups, we conduct a large-scale simulation study to validate our proposed extension. We consider various combinations of groups, samples sizes and measurement errors in the simulation study, providing for each combination the percentage of $p$-values below an alpha-level of 0.05. Additionally, we apply our method to a Cardiotocography data set and find statistically significant evidence of measurable differences in shape between normal, suspect and pathologic health status groups.
Stephen Becker合作论文数University of Colorado Boulder3