This study examines the impact of incorporating cryptocurrencies into global asset portfolios using ensemble approaches and a tracing strategy. We considered cryptocurrency ratios of 1%, 3%, and 5% for including cryptocurrencies. Benchmarking was performed using classical portfolio optimization strategies such as minimum variance portfolio (MVP), maximum diversification portfolio (MDP), equal risk contribution portfolio (ERCP), and hierarchical risk parity (HRP). The ensemble methods and tracing strategies we evaluated were the equally weighted portfolio (EWP), the linear combination portfolio (LCP), the return tracing portfolio (RTP), and the return volatility tracing portfolio (RVTP). EWP averages the weights of classical methods, while LCP combines the objective functions of three optimization methods. RTP and RVTP represent tracing strategy portfolios with monthly rebalancing, selecting the best-performing portfolio based on cumulative returns or a combination of cumulative returns and annualized volatility. Our findings reveal that increasing the cryptocurrency allocation improves performance metrics in ensemble portfolios but also leads to higher risk. In addition, including cryptocurrencies reduces transaction fees, especially evident in the LCP with a 5% allocation. In the case of a 3-month RTP, HRP emerged as the preferred strategy, outperforming the use of HRP alone. In the case of a 6-month RVTP, MVP remained the preferred choice, consistently achieving lower volatility.
As the risk posed by climate change becomes increasingly evident, countries across the world are constantly seeking alternative energy sources. Wind energy has substantial potential for future energy portfolios without having negative impacts on the environment. In developing nationwide and worldwide energy plans, understanding the spatio-temporal pattern of wind is crucial. We analyze wind vectors in the region of East Asia from the fifth-generation ECMWF atmospheric reanalysis. To model the wind vectors, we consider Tukey g-and-h transformation-based non-Gaussian processes, along with multivariate covariance functions. The proposed model can address non-Gaussian features and nonstationary dependence structures of wind vectors. In addition, a two-step inference scheme coupled with the composite likelihood method is applied to handle the computational issues posed by a large dataset. In the first step, we fit the temporal dependence structures of data with a location-specific non-Gaussian time series model. This allows us to remove substantial amounts of nonstationary variations in both space and time, and thus, relatively simple covariance models can handle large and complicated data in the second step. We show that the proposed method with a covariance structure reflecting the nonstationarity due to the latitude difference and the land–ocean difference leads to better predictions for wind speed as well as wind potential, which is crucial for planning wind power generation.
Hepatitis A is a water-borne infectious disease that frequently occurs in unsanitary environments. However, paradoxically, those who have spent their infancy in a sanitary environment are more susceptible to hepatitis A because they do not have the opportunity to acquire natural immunity. In Korea, hepatitis A is prevalent because of the distribution of uncooked seafood, especially during hot and humid summers. In general, the transmission of hepatitis A is known to be dynamically affected by socioeconomic, environmental, and weather-related factors and is heterogeneous in time and space. In this study, we aimed to investigate the spatio-temporal variation of hepatitis A and the effects of socioeconomic and weather-related factors in Korea using a flexible spatio-temporal model. We propose a Bayesian Poisson regression model coupled with spatio-temporal variability to estimate the effects of risk factors. We used weekly hepatitis A incidence data across 250 districts in Korea from 2016 to 2019. We found spatial and temporal autocorrelations of hepatitis A indicating that the spatial distribution of hepatitis A varied dynamically over time. From the estimation results, we noticed that the districts with large proportions of males and foreigners correspond to higher incidences. The average temperature was positively correlated with the incidence, which is in agreement with other studies showing that the incidences in Korea are noticeable in spring and summer due to the increased outdoor activity and intake of stale seafood. To the best of our knowledge, this study is the first to suggest a spatio-temporal model for hepatitis A across the entirety of Korean. The proposed model could be useful for predicting, preventing, and controlling the spread of hepatitis A.
This study aims to improve the performance of voice spoofing attack detection through self-supervised pre-training. Supervised learning needs appropriate input variables and corresponding labels for constructing the machine learning models that are to be applied. It is necessary to secure a large number of labeled datasets to improve the performance of supervised learning processes. However, labeling requires substantial inputs of time and effort. One of the methods for managing this requirement is self-supervised learning, which uses pseudo-labeling without the necessity for substantial human input. This study experimented with contrastive learning, a well-performing self-supervised learning approach, to construct a voice spoofing detection model. We applied MoCo’s dynamic dictionary, SimCLR’s symmetric loss, and COLA’s bilinear similarity in our contrastive learning framework. Our model was trained using VoxCeleb data and voice data extracted from YouTube videos. Our self-supervised model improved the performance of the baseline model from 6.93% to 5.26% for a logical access (LA) scenario and improved the performance of the baseline model from 0.60% to 0.40% for a physical access (PA) scenario. In the case of PA, the best performance was achieved when random crop augmentation was applied, and in the case of LA, the best performance was obtained when random crop and random shifting augmentations were considered.
We set up a general framework for modeling non-Gaussian multivariate stochastic processes by transforming underlying multivariate Gaussian processes. This general framework includes multivariate spatial random fields, multivariate time series, and multivariate spatio-temporal processes, whereas the respective univariate processes can also be seen as special cases. We advocate joint modeling of the transformation and the cross-/auto-correlation structure of the latent multivariate Gaussian process, for better estimation and prediction performance. We provide two useful models, the Tukey g-and-h transformed vector autoregressive model and the sinh-arcsinh-transformed multivariate Matérn random field. We evaluate them with a simulation study. Finally, we apply the two models to a wind data set for modeling the two perpendicular components of wind speed vectors. Both the simulation study and data analysis show the advantages of the joint modeling approach.
Quantifying the uncertainty of wind energy potential from climate models is a time-consuming task and requires considerable computational resources. A statistical model trained on a small set of runs can act as a stochastic approximation of the original climate model, and can assess the uncertainty considerably faster than by resorting to the original climate model for additional runs. While Gaussian models have been widely employed as means to approximate climate simulations, the Gaussianity assumption is not suitable for winds at policy-relevant (i.e., sub-annual) time scales. We propose a trans-Gaussian model for monthly wind speed that relies on an autoregressive structure with a Tukey g-and-h transformation, a flexible new class that can separately model skewness and tail behavior. This temporal structure is integrated into a multi-step spectral framework that can account for global nonstationarities across land/ocean boundaries, as well as across mountain ranges. Inferences are achieved by balancing memory storage and distributed computation for a big data set of 220 million points. Once the statistical model was fitted using as few as five runs, it can generate surrogates rapidly and efficiently on a simple laptop. Furthermore, it provides uncertainty assessments very close to those obtained from all available climate simulations (40) on a monthly scale.
Wind has the potential to make a significant contribution to future energy resources. Locating the sources of this renewable energy on a global scale is however extremely challenging, given the difficulty to store very large data sets generated by modern computer models. We propose a statistical model that aims at reproducing the data-generating mechanism of an ensemble of runs via a Stochastic Generator (SG) of global annual wind data. We introduce an evolutionary spectrum approach with spatially varying parameters based on large-scale geographical descriptors such as altitude to better account for different regimes across the Earth's orography. We consider a multistep conditional likelihood approach to estimate the parameters that explicitly accounts for nonstationary features while also balancing memory storage and distributed computation. We apply the proposed model to more than 18 million points of yearly global wind speed. The proposed SG requires orders of magnitude less storage for generating surrogate ensemble members from wind than does creating additional wind fields from the climate model, even if an effective lossy data compression algorithm is applied to the simulation output.
Wind has the potential to make a significant contribution to future energy resources. Locating the sources of this renewable energy on a global scale is however extremely challenging, given the difficulty to store very large data sets generated by modern computer models. We propose a statistical model that aims at reproducing the data-generating mechanism of an ensemble of runs via a Stochastic Generator (SG) of global annual wind data. We introduce an evolutionary spectrum approach with spatially varying parameters based on large-scale geographical descriptors such as altitude to better account for different regimes across the Earth's orography. We consider a multi-step conditional likelihood approach to estimate the parameters that explicitly accounts for nonstationary features while also balancing memory storage and distributed computation. We apply the proposed model to more than 18 million points of yearly global wind speed. The proposed SG requires orders of magnitude less storage for generating surrogate ensemble members from wind than does creating additional wind fields from the climate model, even if an effective lossy data compression algorithm is applied to the simulation output.
Statistical models used in geophysical, environmental, and climate science applications must reflect the curvature of the spatial domain in global data. Over the past few decades, statisticians have developed covariance models that capture the spatial and temporal behavior of these global data sets. Though the geodesic distance is the most natural metric for measuring distance on the surface of a sphere, mathematical limitations have compelled statisticians to use the chordal distance to compute the covariance matrix in many applications instead, which may cause physically unrealistic distortions. Therefore, covariance functions directly defined on a sphere using the geodesic distance are needed. We discuss the issues that arise when dealing with spherical data sets on a global scale and provide references to recent literature. We review the current approaches to building process models on spheres, including the differential operator, the stochastic partial differential equation, the kernel convolution, and the deformation approaches. We illustrate realizations obtained from Gaussian processes with different covariance structures and the use of isotropic and nonstationary covariance models through deformations and geographical indicators for global surface temperature data. To assess the suitability of each method, we compare their log-likelihood values and prediction scores, and we end with a discussion of related research problems.
There is a growing interest in developing covariance functions for processes on the surface of a sphere due to wide availability of data on the globe. Utilizing the one-to-one mapping between the Euclidean distance and the great circle distance, isotropic and positive definite functions in a Euclidean space can be used as covariance functions on the surface of a sphere. This approach, however, may result in physically unrealistic distortion on the sphere especially for large distances. We consider several classes of parametric covariance functions on the surface of a sphere, defined with either the great circle distance or the Euclidean distance, and investigate their impact upon spatial prediction. We fit several isotropic covariance models to simulated data as well as real data from NCEP/NCAR reanalysis on the sphere. We demonstrate that covariance functions originally defined with the Euclidean distance may not be adequate for some global data.
There have been noticeable advancements in developing parametric covariance models for spatial and spatio-temporal data with various applications to environmental problems. However, literature on covariance models for processes defined on the surface of a sphere with great circle distance as a distance metric is still sparse, due to its mathematical difficulties. It is known that the popular Matérn covariance function, with smoothness parameter greater than 0.5, is not valid for processes on the surface of a sphere with great circle distance. We introduce an approach to produce Matérn-like covariance functions for smooth processes on the surface of a sphere that are valid with great circle distance. The resulting model is isotropic and positive definite on the surface of a sphere with great circle distance, with a natural extension for nonstationarity case. We present extensive numerical comparisons of our model, with a Matérn covariance model using great circle distance as well as chordal distance. We apply our new covariance model class to sea level pressure data, known to be smooth compared to other climate variables, from the CMIP5 climate model outputs.
When the underlying asset price process follows a Lévy process, the market becomes incomplete, in which the option pricing can be a complicated problem. This paper proposes a method of asymptotic option pricing when the underlying asset price process follows a pure-jump Lévy process. We express the option price as the expected value of the discounted payoff and expand it at the Black–Scholes price assuming that the price process converges weakly to the Black–Scholes model. The price can be approximated by a formula with 4 parameters, which can easily be estimated using option prices observed in the market. The proposed price explains the market option data better than the Black–Scholes price in real data application with KOSPI 200.