
We propose a computational framework for reconstructing individual-level data from stratified summary statistics by fitting parametric distributions to reported summaries and their uncertainty. Parameters are estimated by minimizing a weighted loss function that aligns model-implied targets with reported values, where weights are derived from confidence intervals to reflect sampling uncertainty. This formulation enables the integration of heterogeneous summaries, while the uncertainty-weighted objective improves robustness to heterogeneous precision and constrained nonlinear optimization yields stable parametric fits. To approximate population-level distributions, synthetic microdata are generated within strata using the fitted models. Reconstruction accuracy is evaluated through Monte Carlo experiments in which ground-truth datasets are simulated, reduced to summary statistics, reconstructed, and compared with the originals via discrepancies in moments and percentiles across replications. Incorporating uncertainty-based weighting improves recovery of distributional features relative to unweighted fitting. As an illustration, the method is applied to published summary statistics of urinary creatinine concentrations from the Third National Health and Nutrition Examination Survey (NHANES III). The generalized gamma family reproduces reported quantities with high fidelity and yields distributionally coherent synthetic datasets consistent with published uncertainty. The proposed framework provides a reproducible and generalizable approach for recovering distributional structure and generating synthetic microdata from limited published information.
Regression with autoregressive errors is a critical area in time series analysis. This article develops a hybrid modeling framework combining classical regression with machine learning methods to explicitly handle autoregressive error structures. We investigate the integration of autoregressive components into penalized and tree-based models to capture both linearity and nonlinearity in predictors and temporal dependence in residuals. The performances of all these models are compared with one another using simulated and real data example.
Brain tumor classification plays a vital role in clinical diagnosis and treatment planning, yet existing deep-learning models often suffer from feature redundancy, noise sensitivity, and suboptimal feature selection, leading to limited diagnostic accuracy. To address these gaps, this study proposes a novel hybrid framework integrating Multi-Scale Adaptive Bilateral Filtering, Red Panda Integrated Snake (RPS) optimization, and an Attention-Based Dense Autoencoder Network (At_DAENet). The filtering stage enhances MRI quality by preserving edges while removing multi-scale artifacts. A comprehensive set of statistical features, including mean, entropy, kurtosis, and skewness, and textural features such as GLCM contrast, homogeneity, energy, correlation, and dissimilarity, is extracted and subsequently optimized using the RPS algorithm. The RPS method combines the global search capability of Red Panda foraging with the adaptive exploitation behavior of Snake Optimization (SOA), enabling superior convergence performance and reducing the risk of local-optima stagnation. These optimized features are further refined through At_DAENet, where dense connectivity and attention mechanisms prioritize discriminative tumor-specific patterns. Experimental results demonstrate that the proposed model achieves 0.998 accuracy, 0.991 precision, 1.0 recall, 0.9959 F-score, and 0.988 specificity on the BraTS dataset, and 0.972 accuracy, 0.969 precision, 1.0 recall, 0.968 F-score, and 0.919 specificity on the FigShare dataset, outperforming existing state-of-the-art approaches. These findings highlight the robustness of the proposed RPS-driven attention-enhanced classification pipeline and its potential value for reliable brain tumor diagnosis.
Traditional risk models often fail during market crises due to extreme volatility and fat-tail characteristics. To address this, this study proposes a "Tail-Aware RL" architecture that integrates Extreme Value Theory (EVT), Student-t copula, and the Soft Actor-Critic (SAC) algorithm. The framework models marginal tails via POT-EVT and dependence via t-copula, employing a hybrid training strategy (70% synthetic, 30% real data) to mitigate overfitting. Empirical analyses on S&P 100 and BIST 30 indices (2012-2025), covering the COVID-19 and 2023 Banking crises, demonstrate that the model significantly outperforms baseline RL approaches. Specifically, Maximum Drawdown rates improved from 34% to 19% for BIST 30 and from 29% to 14% for S&P100. Validated by Jobson-Korkie and Block-Bootstrap tests, these results indicate that the proposed architecture enhances capital protection and offers a robust alternative for dynamic portfolio management during periods of systemic risk.
Overlap measures are valuable tools for quantifying the degree of similarity between two probability distributions, with a wide range of applications across various fields of statistics. In this paper, we focus on estimating two important overlap measures, Matusita's measure and Weitzman's measure, for distributions belonging to the proportional hazards and proportional reversed hazards families. These families extend a baseline distribution by modifying its hazard or reversed hazard rate, making them particularly suitable for lifetime data modeling. We develop both maximum likelihood and Bayesian estimators for the overlap measures under these models, assuming parametric forms for the baseline distributions. Bayesian estimation is carried out using Lindley's approximation under the assumption of independent gamma priors, and the corresponding highest posterior density credible intervals are derived using the MCMC method. The performance of the proposed estimators is evaluated through Monte Carlo simulations, examining bias, mean squared error, and coverage probabilities. Applications to real-life data further demonstrate the practical utility of the methods.
This paper deals with the prediction analysis of municipality default in Italy using the hybridization of weighted machine learning algorithms with Stacking techniques for classification. Amongst the 32 variables considered in our study, the presence of mafia is identified as a key default predictor of municipality default. Mafia organizations increase municipality default by using violence against corruption of politicians and bureaucrats and then curbing the allocation of public funds aligned with their interests. In addition, mafia presence reduces tax autonomy and increases external dependence of municipalities. Indeed, the relevance of the variables in predicting municipality default is similar to that in predicting municipalities dissolved by mafia infiltration. Overall, the paper suggests that to improve efficiency of local public institutions and to reestablish democracy in mafia-dominated territories, it is necessary to enforce new rules that drastically reduce the mafia presence.
Proper inferences are of great interest to study investigators dealing with data from biased sampling designs. Within the framework of the logistic partially linear model, we address an inference problem for data collected using an innovative two-phase sampling design proposed by Wang et al., where certain covariates are only available in the second-phase. B-splines are used to approximate the nonparametric component, and a pseudo-likelihood method is employed to estimate the regression parameters. We derive the large sample properties of the proposed estimators, further illustrated through simulation studies. We apply the proposed procedure to analyze a real dataset from biomedical research to assess its practical application.
Cloud computing plays a vital role in providing data access and storage services, but it is vulnerable to malicious attacks and unauthorized access, compromising data confidentiality and integrity. Intrusion Detection Systems (IDS) are essential for detecting such threats; however, existing IDSs face challenges in handling large volumes of cloud traffic and often suffer from low accuracy and false negatives. The primary objective of this research is to develop a robust and accurate intrusion detection system called Quasi Recurrent Scaled Residual Network (QRSRes-Net) for cloud intrusion detection. Input network traffic is preprocessed using sigmoid normalization and missing value restoration, followed by Topsoe correlation (TopCorr) for feature selection. The Synthetic Minority Oversampling Technique (SMOTE) addresses class imbalance and improves minority attack representation. Intrusion detection is performed using QRSRes-Net, which integrates Scaling Wide Residual Networks (SWideResNet) and Quasi-Recurrent Neural Networks (QRNN) to capture spatial and temporal dependencies effectively. The model is trained and evaluated on the publicly available CIC QNSL dataset, containing diverse cloud traffic and multiple attack types. The proposed approach achieves 92.445% accuracy, 92.997% TPR, 92.309% TNR, 92.009% precision, and 92.500% F1-score, demonstrating its effectiveness and robustness for large-scale cloud intrusion detection.
Homoscedasticity in spatial data has always been an important assumption to simplify and speed up parameter estimation. However, the presence of heteroscedasticity is common in spatial data; this work presents a methodology based on maximum likelihood for the estimation and inference of SAR models with heteroscedasticity. The proposal is easy to implement and is not computationally intensive. Through simulation exercises, it was demonstrated that the proposed algorithm offers outstanding results in the recovery of both mean and variance parameters and greater precision compared to other methods. Applying the algorithm to regional wage information in Poland revealed subtle differences compared to previous estimates that did not consider heteroscedasticity and allowed the variance to be successfully modeled from the explanatory variables.
We introduce a new skew version of the well-known symmetric Cauchy probability distribution using quantile function approach. This approach is based on splitting the quantile function of the Cauchy distribution into two other quantile functions of some distributions. The original definition of the new distribution has an extra skewness parameter besides location and scale parameters. We also introduce an additional shape parameter to the distribution in order to enhance the flexibility of the distribution in data fitting. The proposed distribution is mathematically more tractable than the other forms defined in the literature and basic properties of the distribution are derived. Several parameter estimation methods are proposed. The results of a simulation study to compare the performances of the estimators are reported. Two real data fitting applications are also given. The results show that the newly defined distribution can be a useful alternative in modeling skew non-normal data.