L & eacute;vy processes are widely used in financial modeling due to their ability to capture discontinuities and heavy tails, which are common in high-frequency asset return data. However, parameter estimation remains a challenge when associated likelihoods are unavailable or costly to compute. We propose a fast and accurate method for L & eacute;vy parameter estimation using the neural Bayes estimation (NBE) framework-a simulation-based, likelihood-free approach that leverages permutation-invariant neural networks to approximate Bayes estimators. We contribute new theoretical results, showing that NBE results in consistent estimators whose risk converges to that of the Bayes estimator under mild conditions. Moreover, through extensive simulations across several L & eacute;vy models, we show that NBE outperforms traditional methods in both accuracy and runtime, while also enabling two complementary approaches to uncertainty quantification. We illustrate our approach on a challenging high-frequency cryptocurrency return dataset, where the method captures evolving parameter dynamics and delivers reliable and interpretable inference at a fraction of the computational cost of traditional methods. We investigate nearly a decade of high-frequency Bitcoin returns, requiring less than one minute to estimate parameters under the proposed approach.
Ocean mesoscale eddies can be thought of as the “weather” of the ocean and strongly influence the ocean’s physics, chemistry, and biology; they influence other components of the Earth system via air-sea and sea-ice interactions, and are crucial drivers of marine heat waves. Thus, proper modeling of eddies in both historical and future climates is crucial to accurately capturing the Earth system. Climate projections using global coupled models with eddying ocean components are only recently starting to be more widely used. Despite their critical role in understanding and forecasting climate characteristics, these so-called eddy-permitting models have not been explored to verify that resolved eddies are realistic, and thus any downstream scientific testing of hypotheses in biogeochemistry, ocean physics or other associated Earth systems impacted by eddies hinge on this critical assumption. This paper compares observed eddies with lifetimes longer than 6 weeks present in 1/4∘ satellite altimetry data with observed eddies in 1/4∘ reanalysis data and ocean model output.When compared to eddies observed in satellite altimetry data, eddies in reanalysis data and ocean model output are missing almost 30% of the number of eddy trajectories. In addition to missing eddy trajectories, the characteristics of eddies in reanalysis data and ocean model output differ from eddies observed in satellite altimetry data. At a high level, eddies in reanalysis data and ocean model output tend to live longer, are larger, and are weaker than eddies in observed altimetry data. This paper presents a variety of statistics describing these differences both spatially and in global aggregate.
Many atmospheric measurement techniques involve inversion of photon counts detected by multi-spectral sensors spanning the X-ray to microwave regions of the electromagnetic spectrum. Although photon counts follow Poisson statistics, commonly used inversion techniques often rely on statistical assumptions that disregard the Poisson nature of the sensor data, limiting the scientific utility of datasets. Motivated to overcome this limiting assumption, this study focuses on retrieval techniques that involve the ratio of counts received in different sub-bands and introduces a new computationally efficient and robust approach to this type of inverse problem that respects the underlying count statistics. The method assumes that the received photon counts in each channel are a realization of a binned point process, allowing the ratio of the channel intensities to be modeled within a hierarchical Bayesian framework. This allows us to directly incorporate correlation between the bins via the prior that is modeled using a permanental process. It further enables more accurate uncertainty quantification without costly sampling procedures common in Bayesian inversion methods. The method is verified and validated on thermospheric column-integrated neutral temperature retrievals from simulated top-of-atmosphere far-ultraviolet (FUV) disk emission data corresponding to 2-8 November 2018, which includes a minor geomagnetic storm. The sub-bands associated with the N2 Lyman-Birge-Hopfield (2,0) transition are used for the ratio calculation. The method is also demonstrated on calibrated photon counts data from the NASA Global-scale Observations of the Limb and Disk (GOLD) mission from the same time period and from 11 May 2024 during a severe geomagnetic storm. The study demonstrates the method's ability to accurately recover neutral temperature in a variety of geophysical conditions, attesting to its potential to extend the fidelity of neutral temperature retrievals over a broader range of solar zenith angles (0-100 degrees), compared to the current limit of (0-80 degrees) available with existing techniques.
We propose a new approach for the modeling large datasets of nonstationary spatial processes that combines a latent low rank process and a sparse covariance model. The low rank component coefficients are endowed with a flexible graphical Gaussian Markov random field model. The utilization of a low rank and compactly-supported covariance structure combines the full-scale approximation and the basis graphical lasso; we term this new approach the full-scale basis graphical lasso (FSBGL). Estimation employs a graphical lasso-penalized likelihood, which is optimized using a difference-of-convex scheme. We illustrate the proposed approach on synthetic fields as well as with a challenging high-resolution simulation dataset of the thermosphere. In a comparison against state-of-the-art spatial models, the FSBGL performs better at capturing salient features of the thermospheric temperature fields, even with limited available training data.
Statistical modeling and interpolation of space-time processes has gained increasing relevance over the last few years. However, real world data often exhibit characteristics that challenge conventional methods such as nonstationarity and temporal misalignment. For example, high frequency solar irradiance data are typically observed at fine temporal scales, but at sparse spatial sampling, so space-time interpolation is necessary to support solar energy studies. The nonstationarity and phase misalignment of such data challenges extant approaches. We propose random elastic space-time (REST) prediction, a novel method that addresses temporally-varying phase misalignment by combining elastic alignment and conventional kriging techniques. Moreover, uncertainty in both amplitude and phase alignment can be readily quantified in a conditional simulation framework, whereas conventional space-time methods only address amplitude uncertainty. We illustrate our approach on a challenging solar irradiance dataset, where our method demonstrates superior predictive distributions compared to existing geostatistical and functional data analytic techniques.
We present a novel space-time Bayesian hierarchical model (BHM) to reconstruct annual Sea Surface Temperature (SST) over a large domain based on SST at limited proxy (i.e., sediment core) locations. The model is tested in the equatorial Pacific. The BHM leverages Principal Component Analysis to identify dominant space-time modes of contemporary variability of the SST field at the proxy locations and employs these modes in a Gaussian process framework to estimate SSTs across the entire domain. The BHM allows us to model the mean field and covariance, varying in space and time in the process layers of the hierarchy. Using the Markov Chain Monte Carlo (MCMC) method and suitable priors on the model parameters, posterior distributions of the model parameters and, consequently, posterior distributions of the SST fields and the attendant uncertainties are obtained for any desired year. The BHM is calibrated and validated in the contemporary period (1854-2014) and subsequently applied to reconstruct SST fields during the Holocene (0-10 ka). Results are consistent with prior inferences of La Ni & ntilde;a-like conditions during the Holocene. This modeling framework opens exciting prospects for modeling and reconstruction of other fields, such as precipitation, drought indices, and vegetation.
Snowpack in mountainous areas often provides water storage for summer and fall, especially in the Western United States. In situ observations of snow properties in mountainous terrain are limited by cost and effort, impacting both temporal and spatial sampling, while remote sensing estimates provide more complete spacetime coverage. Spatial estimates of fractional snow covered area (fSCA) at 30m are available every 16 days from the series of multispectral scanning instruments on Landsat platforms. Daily estimates at 463m spatial resolution are also available from the Moderate Resolution Imaging Spectroradiometer (MODIS) instrument on the Terra satellite. Fusing Landsat and MODIS fSCA images creates high resolution daily spatial estimates of fSCA that are needed for various uses: to support scientists and managers interested in energy and water budgets for water resources and to understand the movement of animals in a changing climate. Here, we propose a new machine learning approach conditioned on MODIS fSCA, as well as a set of physiographic features, and fit to Landsat fSCA over a portion of the Sierra Nevada USA. The predictions are daily 30m fSCA. The approach relies on two stages of spatially-varying models. The first classifies fSCA into three categories and the second yields estimates within (0, 100) percent fSCA. Separate models are applied and fitted within sub-regions of the study domain. Compared with a recently-published machine learning model (Rittger, Krock, et al., 2021), this approach uses spatially local (rather than global) random forests, and improves the classification error of fSCA by 16%, and fractionally-covered pixel estimates by 18%.
Given the tradeoffs between spatial and temporal resolution, questions about resolution optimality are fundamental to the study of global snow. Answers to these questions will inform future scientific priorities and mission specifications. Heterogeneity of mountain snowpacks drives a need for daily snow cover mapping at the slope scale (=30 m) that is unmet for a variety of scientific users, ranging from hydrologists to the military to wildlife biologists. But finer spatial resolution usually requires coarser temporal or spectral resolution. Thus, no single sensor can meet all these needs. Recently, constellations of satellites and fusion techniques have made noteworthy progress. The efficacy of two such recent advances is examined: (1) a fused MODIS-Landsat product with daily 30 m spatial resolution and (2) a harmonized Landsat 8 and Sentinel 2A and B (HLS) product with 3-4 d temporal and 30 m spatial resolution. State-of-the-art spectral unmixing techniques are applied to surface reflectance products from 1 and 2 to create snow cover and albedo maps. Then an energy balance model was run to reconstruct snow water equivalent (SWE). For validation, lidar-based Airborne Snow Observatory SWE estimates were used. Results show that reconstructed SWE forced with 30 m resolution snow cover has lower bias, a measure of basin-wide accuracy, than the baseline case using MODIS (463 m cell size) but greater mean absolute error, a measure of per-pixel accuracy. However, the differences in errors may be within uncertainties from scaling artifacts, e.g., basin boundary delineation. Other explanations are (1) the importance of daily acquisitions and (2) the limitations of downscaled forcings for reconstruction. Conclusions are as follows: (1) spectrally unmixed snow cover and snow albedo from MODIS continue to provide accurate forcings for snow models and (2) finer spatial and temporal resolution through sensor design, fusion techniques, and satellite constellations are the future for Earth observations, but existing moderate-resolution sensors still offer value.
We propose a new modeling framework for highly-multivariate spatial processes that synthesizes ideas from recent multiscale and spectral approaches with graphical models. The basis graphical lasso writes a univariate Gaussian process as a linear combination of basis functions weighted with entries of a Gaussian graphical vector whose graph is estimated from optimizing an $\ell_1$ penalized likelihood. This paper extends the setting to a multivariate Gaussian process where the basis functions are weighted with Gaussian graphical vectors. We motivate a model where the basis functions represent different levels of resolution and the graphical vectors for each level are assumed to be independent. Using an orthogonal basis grants linear complexity and memory usage in the number of spatial locations, the number of basis functions, and the number of realizations. An additional fusion penalty encourages a parsimonious conditional independence structure in the multilevel graphical model. We illustrate our method on a large climate ensemble from the National Center for Atmospheric Research's Community Atmosphere Model that involves 40 spatial processes.
Lévy processes are useful tools for analysis and modeling of jump‐diffusion processes. Such processes are commonly used in the financial and physical sciences. One approach to building new Lévy processes is through subordination, or a random time change. In this work, we discuss and examine a type of multiply subordinated Lévy process model that we term a deep variance gamma (DVG) process, including estimation and inspection methods for selecting the appropriate level of subordination given data. We perform an extensive simulation study to identify situations in which different subordination depths are identifiable and provide a rigorous theoretical result detailing the behavior of a DVG process as the levels of subordination tend to infinity. We test the model and estimation approach on a data set of intraday 1‐min cryptocurrency returns and show that our approach outperforms other state‐of‐the‐art subordinated Lévy process models.
Tropical cyclones are important drivers of coastal flooding which have severe negative public safety and economic consequences. Due to the rare occurrence of such events, high spatial and temporal resolution historical storm precipitation data are limited in availability. This article introduces a statistical tropical cyclone space-time precipitation generator given limited information from storm track datasets. Given a handful of predictor variables that are common in either historical or simulated storm track ensembles such as pressure deficit at the storm's center, radius of maximalwinds, storm center and direction, and distance to coast, the proposed stochastic model generates space-time fields of quantitative precipitation over the study domain. Statistically novel aspects include that the model is developed in Lagrangian coordinates with respect to the dynamic storm center that uses ideas from low-rank representations along with circular process models. The model is trained on a set of tropical cyclone data from an advanced weather forecasting model over the Gulf of Mexico and southern United States, and is validated by cross-validation. Results show the model appropriately captures spatial asymmetry of cyclone precipitation patterns, total precipitation as well as the local distribution of precipitation at a set of case study locations along the coast. We additionally compare our model against a widely-used statistical forecast, and illustrate that our approach better captures uncertainty, as well as storm characteristics such as asymmetry.
Traditionally the power grid has been a one‐way street with power flowing from large transmission‐connected generators through the distribution network to consumers. This paradigm is changing with the introduction of distributed renewable energy resources (DERs), and with it, the way the grid is managed. There is currently a dearth of high fidelity solar irradiance datasets available to help grid researchers understand how expansion of DERs could affect future power system operations. Realistic simulations of by‐the‐second solar irradiances are needed to study how DER variability affects the grid. Irradiance data are highly non‐stationary and non‐Gaussian, and even modern time series models are challenged by their distributional properties. We develop a subordinated non‐Gaussian stochastic model whose simulations realistically capture the distribution and dependence structure in measured irradiance. We illustrate our approach on a fine resolution dataset from Hawaii, where our approach outperforms standard nonlinear time series models.
We develop a space‐time Bayesian hierarchical modeling (BHM) framework for two flood risk attributes—seasonal daily maximum flow and the number of events that exceed a threshold during a season (NEETM)—at a suite of gauge locations on a river network. The model uses generalized extreme value (GEV) and Poisson distributions as marginals for these flood attributes with non‐stationary parameters. The rate parameters of the Poisson distribution and location, scale, and shape parameters of the GEV are modeled as linear functions of suitable covariates. Gaussian copulas are applied to capture the spatial dependence. The best covariates are selected using the Watanabe‐Akaike information criterion (WAIC). The modeling framework results in the posterior distribution of the flood attributes at all the gauges and various lead times. We demonstrate the utility of this modeling framework to forecast the flood risk attributes during the summer peak monsoon season (July‐August) at five gauges in the Narmada River basin (NRB) of West‐Central India for several lead times (0–3 months). As potential covariates, we consider climate indices such as El Niño–Southern Oscillation (ENSO), the Indian Ocean Dipole (IOD), and the Pacific Warm Pool Region (PWPR) from antecedent seasons, which have shown strong teleconnections with the Indian monsoon. We also include new indices related to the East Pacific and West Indian Ocean regions depending on the lead times. We show useful long lead skill from this modeling approach which has a strong potential to enable robust risk‐based flood mitigation and adaptation strategies 3 months before flood occurrences.
Timely projections of seasonal streamflow extremes can be useful for the early implementation of annual flood risk adaptation strategies. However, predicting seasonal extremes is challenging, particularly under nonstationary conditions and if extremes are correlated in space. The goal of this study is to implement a space–time model for the projection of seasonal streamflow extremes that considers the nonstationarity (interannual variability) and spatiotemporal dependence of high flows. We develop a space–time model to project seasonal streamflow extremes for several lead times up to 2 months, using a Bayesian hierarchical modeling (BHM) framework. This model is based on the assumption that streamflow extremes (3 d maxima) at a set of gauge locations are realizations of a Gaussian elliptical copula and generalized extreme value (GEV) margins with nonstationary parameters. These parameters are modeled as a linear function of suitable covariates describing the previous season selected using the deviance information criterion (DIC). Finally, the copula is used to generate streamflow ensembles, which capture spatiotemporal variability and uncertainty. We apply this modeling framework to predict 3 d maximum streamflow in spring (May–June) at seven gauges in the Upper Colorado River basin (UCRB) with 0- to 2-month lead time. In this basin, almost all extremes that cause severe flooding occur in spring as a result of snowmelt and precipitation. Therefore, we use regional mean snow water equivalent and temperature from the preceding winter season as well as indices of large-scale climate teleconnections – El Niño–Southern Oscillation, Atlantic Multidecadal Oscillation, and Pacific Decadal Oscillation – as potential covariates for 3 d spring maximum streamflow. Our model evaluation, which is based on the comparison of different model versions and the energy skill score, indicates that the model can capture the space–time variability in extreme streamflow well and that model skill increases with decreasing lead time. We also find that the use of climate variables slightly enhances skill relative to using only snow information. Median projections and their uncertainties are consistent with observations, thanks to the representation of spatial dependencies through covariates in the margins and a Gaussian copula. This spatiotemporal modeling framework helps in the planning of seasonal adaptation and preparedness measures as predictions of extreme spring streamflows become available 2 months before actual flood occurrence.
Historically, power has flowed from large power plants to customers. Increasing penetration of distributed energy resources such as solar power from rooftop photovoltaic has made the distribution network a two‐way‐street with power being generated at the customer level. The incorporation of renewables introduces additional uncertainty and variability into the power grid. Distribution network operation studies are being adapted to include renewables; however, such studies require high quality solar irradiance data that adequately reflect realistic meteorological variability. Data from satellite‐based products are spatially complete, but temporally coarse, whereas solar irradiances exhibit high frequency variation at very fine timescales. We propose a new stochastic method for temporally downscaling global horizontal irradiance (GHI) to 1 min resolution, but we do not consider the spatial aspect due to limited availability of the in situ irradiance measurements. Solar irradiance's first and second‐order structures vary diurnally and seasonally, and our model adapts to such nonstationarity. Empirical irradiance data exhibits highly non‐Gaussian behavior; we develop a nonstationary and non‐Gaussian moving average model that is shown to capture realistic solar variability at multiple timescales. We also propose a new estimation scheme based on Cholesky factors of empirical autocovariance matrices, bypassing difficult and inaccessible likelihood‐based approaches. The model is demonstrated for a case study of three locations that are located in diverse climates through the United States. The model is compared against competitors from the literature and is shown to provide better uncertainty and variability quantification on testing data.
Abstract. Given the tradeoffs between spatial and temporal resolution, questions about resolution optimality are fundamental to the study of global snow. Answers to these questions will inform future scientific priorities and mission specifications. Heterogeneity of mountain snowpacks drives a need for daily snow cover mapping at the slope scale (≤ 30 m) that is unmet for a variety of scientific users, ranging from hydrologists to the military to wildlife biologists. But finer spatial resolution usually requires coarser temporal or spectral resolution. Thus, no single sensor can meet all these needs. Recently, constellations of satellites and fusion techniques have made noteworthy progress. The efficacy of two such recent advances is examined: 1) a fused MODIS - Landsat product with daily 30 m spatial resolution; and 2) a harmonized Landsat 8 - Sentinel 2A/B (HLS) product with 2–3 day temporal and 30 m spatial resolution. State-of-art spectral unmixing techniques are applied to surface reflectance products from 1 & 2 to create snow cover and albedo maps. Then an energy balance model was run to reconstruct snow water equivalent (SWE). For validation, lidar-based Airborne Snow Observatory SWE estimates were used. Results show that reconstructed SWE forced with 30 m resolution snow cover has lower bias, a measure of basin-wide accuracy, than the baseline case using MODIS (463 m cell size), but higher mean absolute error, a measure of per-pixel accuracy. However, the differences in errors may be within uncertainties from scaling artifacts e.g., basin boundary delineation. Other explanations are 1) the importance of daily acquisitions and 2) the limitations of downscaled forcings for reconstruction. Conclusions are: 1) spectrally unmixed snow cover and snow albedo from MODIS continue to provide accurate forcings for snow models; and 2) finer spatial and temporal resolution through sensor design, fusion techniques, and satellite constellations are the future for Earth observations.
Surface registration, the task of aligning several multidimensional point sets, is a necessary task in many scientific fields. In this work, a novel statistical approach is developed to solve the problem of nonrigid registration. While the application of an affine transformation results in rigid registration, using a general nonlinear function to achieve nonrigid registration is necessary when the point sets require deformations that change over space. The use of a local likelihood-based approach using windowed Gaussian processes provides a flexible way to accurately estimate the nonrigid deformation. This strategy also makes registration of massive data sets feasible by splitting the data into many subsets. The estimation results yield spatially-varying local rigid registration parameters. Gaussian process surface models are then fit to the parameter fields, allowing prediction of the transformation parameters at unestimated locations, specifically at observation locations in the unregistered data set. Applying these transformations results in a global, nonrigid registration. A penalty on the transformation parameters is included in the likelihood objective function. Combined with smoothing of the local estimates from the surface models, the nonrigid registration model can prevent the problem of overfitting. The efficacy of the nonrigid registration method is tested in two simulation studies, varying the number of windows and number of points, as well as the type of deformation. The nonrigid method is applied to a pair of massive remote sensing elevation data sets exhibiting complex geological terrain, with improved accuracy and uncertainty quantification in a cross validation study versus two rigid registration methods.
Recent studies predict significant decreases in the future levelized cost of energy (LCOE) of offshore wind energy, much of which is attributed to anticipated cost reductions from technological innovation. This study evaluates the spatial variability of LCOE caused by technology-induced decreases in a range of capital, operational, and financial cost categories. A spatial cost model of fixed-bottom and floating offshore wind plants is used to model the impact across thousands of potential United States sites. A specified change in an individual turbine subsystem cost produces a range of LCOE outcomes due to the varying geospatial characteristics of the considered sites and the nonlinear, interactive dependency on these input parameters; for example, a 10.8% improvement in net capacity factor can reduce LCOE by between 6% and 20% at different sites. This work expands upon the existing offshore wind literature, which typically evaluates cost sensitivities at a single site and does not consider the spatial variance in LCOE. The results suggest that the impact of technological innovations can be considerable and should be considered on a spatial as well as temporal basis when prioritizing technology innovation research or funding decisions to advance offshore wind technologies in the United States.
Stephen Becker合作论文数University of Colorado Boulder3