The paper presents for the first time a methodology for solving supervised learning problems, such as classification and regression, based on deep Gaussian mixture models (DGMMs). We use a self-supervised approach to construct a classifier as well as a semi-supervised one for a regressor. More than 20 public UCI datasets with various parameters were used for testing. It has been demonstrated that the greatest increase in classification accuracy of 37.69% is achieved by using the ensemble of DGMM and extreme gradient boosting (XGBoost). The accuracy of this method exceeds that of the combination of GMM and SVM by 14.51% . The DGMM regression (DGMMR) analogue of the Gaussian mixture model regression (GMMR) is introduced as a semi-supervised learning algorithm. On the test data, the best results were shown by the ensemble of DGMMR and XGBoost regression. The accuracy of this method exceeded the combination with support vector machines regression (SVR), as well as variants of GMMR with SVR and linear regression with SVR by 3.58% , 11.63% and 32.78% , respectively.
The paper introduces a probability-informed methodology for the segmentation of synthetic aperture radar (SAR) images in the case of small sample learning. It assumes that the amount of training data is limited to several hundred or thousand elements, which prevents the effective training of state-of-the-art neural network (NN) models. This is a typical problem for real SAR images whose characteristics depend significantly on the sensors used to produce them and cannot always be repeated within open available datasets. To solve this problem, we propose NN models called Probability-Informed Neural Networks (PrINNs). As part of our approach, we introduce the use of probability models as a source of additional features for data. Specifically, the training dataset is enriched by modeling the pixel brightness using a finite normal mixture. We prove that such an extension can reduce errors in the learning process theoretically. The resulting enriched dataset is segmented using attention-based convolutional NNs or visual transformers. Then, post-processing is implemented based on another probability model—quadtree, which is a special case of random Markov fields. As we have theoretically demonstrated, this part of PrINNs is analogous to the graph-convolutional NNs with fixed weights. Using open SAR images obtained by different radars (namely, Sentinel-1, Capella, ESAR and HRSID) with various types of underlying surfaces, the possibility of improving segmentation quality based on PrINNs is demonstrated. We tested various combinations of methods from the PrINNs architecture, and in all cases, the PrINN approach we proposed was superior to any other combination of these methods. From the point of view of the achieved accuracy metrics, the mean F_1 score increased up to 19.24% , and the median F_1 score was improved up to 9.57% . Some further architectural improvements to PrINNs are also discussed in the paper.
The advancement of cloud computing technologies has positioned virtual machine (VM) migration as a critical area of research, essential for optimizing resource management, bolstering fault tolerance, and ensuring uninterrupted service delivery. This paper offers an exhaustive analysis of VM migration processes within cloud infrastructures, examining various migration types, server load assessment methods, VM selection strategies, ideal migration timing, and target server determination criteria. We introduce a queuing theory-based model to scrutinize VM migration dynamics between servers in a cloud environment. By reinterpreting resource-centric migration mechanisms into a task-processing paradigm, we accommodate the stochastic nature of resource demands, characterized by random task arrivals and variable processing times. The model is specifically tailored to scenarios with two servers and three VMs. Through numerical examples, we elucidate several performance metrics: task blocking probability, average tasks processed by VMs, and average tasks managed by servers. Additionally, we examine the influence of task arrival rates and average task duration on these performance measures.
The paper proposes an approach to the joint use of statistical and machine learning (ML) models to solve the problems of the precise reconstruction of historical events, real-time detection of ongoing incidents, and the prediction of future quality of service -related occurrences for prospective development of the modern networks. For forecasting, a regression version of the deep Gaussian mixture model (DGMM) is introduced. First, the preliminary clustering based on the finite normal mixtures is performed. This information is then used as an input for some supervised ML algorithm. It is the basic concept of the probability -informed ML approach in the field of telecommunications networks. Using the real -world datasets from a Portuguese mobile operator as well as public cellular traffic data, the article compares this approach with methods such as random forests, support vector machine regression, gradient boosting and LSTM. Vector autoregression, informed by the parameters of the generalized gamma (GG) distribution, which has also been successfully used to reconstruct past traffic patterns, is also used as a benchmark. We demonstrate that DGMM-based regression is 6.82-22.8 times faster than LSTM for the dataset. Moreover, DGMM-based regression can achieve better results for the most important traffic characteristics (average and total traffic, the number of users). For metrics MAPE and RMSE, it surpasses the results of statistical methods up to 46.7% (RMSE) and 91.5% (MAPE) (median increases are 28.0% and 80.1%, respectively), as well as for ML methods up to 13.0% (RMSE) and 35.7% (MAPE) (median increases are 0.39% and 2.5%, respectively). Thus, the use of a probability -informed approach for telecommunication data seems optimal for the computational speed and accuracy trade-off. Also, we introduce a novel statistical hypothesis testing method based on GG distribution for detecting suspected anomalies in traffic.
This paper presents a new approach in the field of probability-informed machine learning (ML). It implies improving the results of ML algorithms and neural networks (NNs) by using probability models as a source of additional features in situations where it is impossible to increase the training datasets for various reasons. We introduce connected mixture components as a source of additional information that can be extracted from a mathematical model. These components are formed using probability mixture models and a special algorithm for merging parameters in the sliding window mode. This approach has been proven effective when applied to real-world time series data for short- and medium-term forecasting. In all cases, the models informed by the connected mixture components showed better results than those that did not use them, although different informed models may be effective for various datasets. The fundamental novelty of the research lies both in a new mathematical approach to informing ML models and in the demonstrated increase in forecasting accuracy in various applications. For geophysical spatiotemporal data, the decrease in Root Mean Square Error (RMSE) was up to 27.7%, and the reduction in Mean Absolute Percentage Error (MAPE) was up to 45.7% compared with ML models without probability informing. The best metrics values were obtained by an informed ensemble architecture that fuses the results of a Long Short-Term Memory (LSTM) network and a transformer. The Mean Squared Error (MSE) for the electricity transformer oil temperature from the ETDataset had improved by up to 10.0% compared with vanilla methods. The best MSE value was obtained by informed random forest. The introduced probability-informed approach allows us to outperform the results of both transformer NN architectures and classical statistical and machine learning methods.
This paper compares two statistical methods for parameter reconstruction (random drift and diffusion coefficients of the Itô stochastic differential equation, SDE) in the problem of stochastic modeling of air–sea heat flux increment evolution. The first method relates to a nonparametric estimation of the transition probabilities (wherein consistency is proven). The second approach is a semiparametric reconstruction based on the approximation of the SDE solution (in terms of distributions) by finite normal mixtures using the maximum likelihood estimates of the unknown parameters. This approach does not require any additional assumptions for the coefficients, with the exception of those guaranteeing the existence of the solution to the SDE itself. It is demonstrated that the corresponding conditions hold for the analyzed data. The comparison is carried out on the simulated samples, modeling the case where the SDE random coefficients are represented in trigonometric form, which is related to common climatic models, as well as on the ERA5 reanalysis data of the sensible and latent heat fluxes in the North Atlantic for 1979–2022. It is shown that the results of these two methods are close to each other in a quantitative sense, but differ somewhat in temporal variability and spatial localization. The differences during the observed period are analyzed, and their geophysical interpretations are presented. The semiparametric approach seems promising for physics-informed machine learning models.
One of the challenging tasks in 5G networks is to organize a joint URLLC (ultra-reliable and low-latency communication) and eMBB (enhanced mobile broadband) transmission in such a way that provides the priority to URLLC connections. The eMBB users suffer quality of service degradation, primarily bit rate degradation, as well as service interruption. In the paper, we provide a queuing model for analyzing this effect depending on several path loss models. The queuing system is of type resource queuing system, where the resource has three-dimensional structure – frequency bandwidth, radio frame length, and transmitted signal power. Due to different URLLC and eMBB bit rate requirements, we use weighted round robin (WRR) resource allocation scheme. The stationary probability distribution depends on the conditional probabilities of session acceptance. We provide the formulas for calculating eMBB metrics – average bit rate, interruption probability, and blocking probability. A numerical example illustrates the impact of two path loss models for macro- and microcells on eMBB metrics.
A dynamic stochastic model based on the Langevin stochastic differential equation is introduced for the reanalysis data of the ERA5 database to model and analyze the behavior of latent and sensible air–sea heat fluxes in the North Atlantic for the period 1979–2022. The point estimates of the random coefficients (the drift vector and the diffusion matrix) of this type of equation for the entire period under consideration are presented. The numerical methods and software tools for statistical analysis of time evolution of the coefficients as well as determination their relationships and the behavior of their maxima, averages and minima at various time intervals (days, months, years), are developed. A strong seasonality for the coefficients of the equation is demonstrated. The spatiotemporal variability of the dynamic and stochastic components of the coefficients of the Langevin equation and their relationship with jet streams of different regions of the North Atlantic is analyzed. The presence of non-trivial positive trends in the drift and diffusion coefficients, especially for the latent fluxes, within the interannual variability is demonstrated. One indicates a quantitative increase in the air–sea interaction on the interannual scale. Numerical estimation was carried out using high-performance computing cluster with software implementation in the Python programming language. The tools for dynamic visualization of various quantities on geographical maps of the region under consideration are also presented.
Statistical regularities of the intra- and interannual variability of sensible and latent heat fluxes in the North Atlantic, including those based on identifying regression dependencies with various averaging of time series, are investigated. Various characteristics of fluxes are estimated, such as maxima and minima over the water basin, mean values, and medians. Based on ERA5 reanalysis data in 1979–2021, the evolution of these values in the North Atlantic is studied and compared with the behavior of the heat fluxes, both from year to year and within a mean climatic year. It is shown that there is a positive trend in the fluxes; parameters of the fluxes are estimated. The spatiotemporal variability of the extreme characteristics of fluxes (maximum and minimum) over the computational domain at fixed times is analyzed.
Fifth-generation (5G) networks require efficient radio resource management (RRM) which should dynamically adapt to the current network load and user needs. Monitoring and forecasting network performance requirements and metrics helps with this task. One of the parameters that highly influences radio resource management is the profile of user traffic generated by various 5G applications. Forecasting such mobile network profiles helps with numerous RRM tasks such as network slicing and load balancing. In this paper, we analyze a dataset from a mobile network operator in Portugal that contains information about volumes of traffic in download and upload directions in one-hour time slots. We apply two statistical models for forecasting download and upload traffic profiles, namely, seasonal autoregressive integrated moving average (SARIMA) and Holt-Winters models. We demonstrate that both models are suitable for forecasting mobile network traffic. Nevertheless, the SARIMA model is more appropriate for download traffic (e.g., MAPE [mean absolute percentage error] of 11.2% vs. 15% for Holt-Winters), while the Holt-Winters model is better suited for upload traffic (e.g., MAPE of 4.17% vs. 9.9% for SARIMA and Holt-Winters, respectively).
The paper aims to identify hidden Markov model parameters. The unobservable state represents a finite-state Markov jump process. The observations contain Wiener noise with state-dependent intensity. The identified parameters include the transition intensity matrix of the system state, conditional drift and diffusion coefficients in the observations. We propose an iterative identification algorithm based on the fixed-interval smoothing of the Markov state. Using the calculated state estimates, we restore all required system parameters. The paper contains a detailed description of the numerical schemes of state estimation and parameter identification. The comprehensive numerical study confirms the high precision of the proposed identification estimates.
In the paper, justification is given for convergence of the median modification of the classical expectation-maximization (EM) algorithm for separation of finite mixtures of normal distributions. This modification is designed to overcome the instability of the classical EM algorithm with respect to initial data.
This paper presents a feature construction approach called Statistical Feature Construction (SFC) for time series prediction. Creation of new features is based on statistical characteristics of analyzed data series. First, the initial data are transformed into an array of short pseudo-stationary windows. For each window, a statistical model is created and characteristics of these models are later used as additional features for a single window or as time-dependent features for the entire time series. To demonstrate the effect of SFC, five plasma physics and six oceanographic time series were analyzed. For each window, unknown distribution parameters were estimated with the method of moving separation of finite normal mixtures. First four statistical moments of these mixtures for initial data and increments were used as additional data features. Multi-layer recurrent neural networks were trained to create short- and medium-term forecasts with a single window as input data; additional features were used to initialize the hidden state of recurrent layers. A hyperparameter grid-search was performed to compare fully-optimized neural networks for original and enriched data. A significant decrease in RMSE metric was observed with a median of 11.4%. There was no increase in RMSE metric in any of the analyzed time series. The experimental results have shown that SFC can be a valuable method for forecasting accuracy improvement.
The paper proposes the use of related components by the method of the moving separation of mixtures as nontrivial features to expand the feature space in problems of the learning of recurrent neural networks. These features are added based on the approximation of data increments using probabilistic models based on finite normal mixtures. To take into account relationships in the data as well as in related components, the article uses the long short-term memory variant of recurrent architectures. The proposed approach is used to build an automated trading strategy based on an ensemble of the long short-term memory networks for the three most commonly traded currency pairs: euro–US dollar, US dollar–Japanese yen, and euro–pound sterling, for which data are taken from January 2011 to the end of September 2021. It is shown that the profitability of the developed ensemble long short-term memory model using additional features, i.e., information on the probabilistic distribution of data increments, outperforms both the basic methods of algorithmic trading by financial indicators (advantage of up to 32.2% on test data) and well-known approaches based on long short-term memory networks without statistical expansion of the feature space (advantage of up to 23.3%). For the best models within the framework of model trading, the final and annual yields are found to be up to 99% and 54%, respectively.
At the specific power of electron cyclotron resonance (ECR) heating of 3.2 MW m–3 (plasma density of 2 × 1019 m–3, electron temperature of 0.6 keV), an increase in the plasma energy lifetime by not less than 30% is accompanied by a two-time-decrease in the level of short-wave turbulent density fluctuations. In such a shot, before the beginning of the quasi-stationary confinement stage, the turbulent state of density fluctuations is characterized by the stronger deviation from zero of the coefficient of excess of fluctuation increments than it is in shots without transport transitions. This indicates the stronger deviation of the probability distribution function of density fluctuation increments from the normal law in shots with transport transitions. Based on the analysis of increments of short-wave fluctuations using the special method for separating the continuous components in stochastic processes, a qualitative difference was established between the behaviors of the structural components forming the plasma turbulence in shots with and without transport transitions. In addition, for shots with transport transitions, a change in the shape of the approximating finite mixture of normal distributions and parameters of its component densities is demonstrated.
In this paper, statistical regularities of the intra-annual variability of heat fluxes in the North Atlantic during the ocean–atmosphere interaction are analyzed. A diffusion random process is considered as a mathematical model of the variability of heat fluxes. Parameters of this process, that is, the drift vector and the diffusion (or standard deviation) matrix, are estimated statistically using original methods. According to ERA-5 reanalysis data for 2011–2020, the evolution of these coefficients in the North Atlantic is studied and their behavior is compared with the behavior of the heat fluxes themselves. Zones of maximum, minimum, and average values of these flows are identified throughout the area under study with daily and 6-h averaging; their behavior and the behavior of their daily variability are described as random values throughout the year. Statistical fitting of parametric models of their distributions is implemented. Areas of the North Atlantic in which systematic factors are of decisive importance (the drift parameter exceeds the diffusion parameter) and vice versa are determined. This effect is discussed in terms of the behavior of the parameters of the probability distributions for increments of the processes under consideration. The spatiotemporal variability of the extreme characteristics of fluxes (maximum and minimum over the computational domain at a fixed time instant) is analyzed.
This paper is devoted to the research of quality improvement of medium-term data forecasting with neural networks within introduction of statistical models for observations. The main goal is to study the efficiency of nontrivial expansion of the feature space based on the characteristics of finite mixtures models. Such probabilistic models are successfully used as convenient approximations of plasma turbulence processes. In the paper, expectation, variance, skewness and kurtosis of the mixture models are introduced as additional features for machine learning algorithms. Comparison of medium-term forecasts with and without additional features is carried out, and various neural network architectures are investigated. The proposed methods are tested on the unique turbulent plasma ensembles obtained from the L-2M stellarator. It is demonstrated that the usage of the above-mentioned statistical characteristics can increase the accuracy of neural network forecasts in terms of such standard metrics as root-mean-square error and mean absolute errors. Hybrid high-performance computing cluster is used in order to increase the learning rate.
The paper presents comparing various machine learning algorithms in the problem of imputation of missing values in the spatiotemporal precipitation data. Using a special procedure to convert complete data to incomplete one, up to 40% of missing values are artificially placed into the datasets. Then, they are imputed in order to determine the most effective machine learning algorithms with the same hyperparameters for more than hundred worldwide weather stations. A two-step procedure, where the classification results are used to improve regression accuracy, are implemented using Python programming language. The efficiency of various combinations of methods including random forests, classic and extreme gradient boosting, support vector machine, EM algorithm are analyzed. It is demonstrated that the best classifier is extreme gradient boosting with average forecasting accuracy of 83.41%. Combination of such methods as XGBClass+XGBoost leads to the best quality of missing values imputation with the normalized RMSE equals from 0.01 to 0.07. All of the above-mentioned methods are tested for the same hyperparameter settings for all weather stations. The novelty of this paper is in the selection of the universal methods for imputation, the accuracies of which are sufficient for processing spatiotemporal meteorological data regardless of their geographic locations even without the fine-tuning. The results obtained allow us to implement methods of computational statistics for detecting extreme precipitation correctly. The presented approaches are also effective for a wider class of observations, for example, environmental data.
We present new mixture representations for the generalized Linnik distribution in terms of normal, Laplace, and generalized Mittag–Leffler laws. In particular, we prove that the generalized Linnik distribution is a normal scale mixture with the generalized Mittag–Leffler mixing distribution. Based on these representations, we prove some limit theorems for a wide class of statistics constructed from samples with random sized including, e.g., random sums of independent random variables with finite variances, in which the generalized Linnik distribution plays the role of the limit law. Thus we demonstrate that the scheme of geometric (or, in general, negative binomial) summation is by far not the only asymptotic setting (even for sums of independent random variables) in which the generalized Linnik law appears as the limit distribution.