
Receiver operating characteristic (ROC) and the area under the curve (AUC) are widely used to evaluate diagnostic classifiers, but global performance metrics may obscure variability across regions of the prediction space. This study assessed whether Z-score-based stratification provides a localized, distribution-aware evaluation of classifier performance. Classifier predictions were standardized and partitioned into central (|Z| <= 1), intermediate (1 < |Z| <= 2), and extreme (|Z| > 2) regions. ROC, detection error trade-off (DET), calibration, and threshold analyses were performed within each region using simulated and real-world datasets. Although global AUC values suggested good overall performance, substantial regional variability was observed. Central regions showed reduced discrimination, greater class overlap, poorer calibration, and less stable thresholds, whereas extreme regions demonstrated more favorable performance characteristics. Z-score-based stratification provides a simple extension to standard evaluation methods and may help identify regional differences in classifier behavior not apparent from global summaries alone.
We introduce here a new and general class of symmetric models, referred to as the q-symmetric family of distributions. The proposed framework includes several well-known symmetry structures, such as the classical location-symmetric, log-symmetric, and M & ouml;bius-symmetric families, as special cases. We then discuss the theoretical characterization of q-symmetric distributions and their relationships to existing symmetry concepts. Next, using some properties of order statistics, we characterize these distributions based on a range of information-theoretic measures associated with the probability density function. In particular, they include Shannon, R & eacute;nyi, and Tsallis entropies as well as their information generating functions, extropy and weighted extropy, to enunciate some links between symmetry, information measures, and properties of order statistics.
Poisson regression is widely used to model count data across various fields, including healthcare, economics, engineering, and environmental research. It is designed to describe the relationship between a numerical response variable and one or more explanatory variables. A maximum likelihood estimator is typically used to estimate the parameters of a Poisson regression model. This method is highly effective under normal conditions, but when the explanatory variables are highly correlated, it violates the model's assumption of independence. This problem is known as multicollinearity, which leads to the failure to identify individual effects, variance inflation, and an increased mean squared error, thus weakening the model's reliability. Some biased estimators are introduced as an alternative to the maximum likelihood estimator. In this study, we propose a novel, alternative, and effective estimator for estimating parameters in a Poisson regression model with multicollinearity. The statistical properties of the proposed estimator are presented, and to confirm its effectiveness, theoretical comparisons with existing estimators are used, along with an extended Monte Carlo simulation under various conditions. The results validate the superiority and effectiveness of the proposed estimator. To support these results, two real-world data applications are used, which confirm the findings of both the simulation and the theoretical comparisons, demonstrating the effectiveness and reliability of the proposed estimator in estimating the parameters of a Poisson regression model under multicollinear conditions.
In this paper, we study the density estimation for the edge frequency polygon under psi-mixing samples. Under appropriate conditions, the uniform strong consistency, the uniform asymptotic normality, and the convergence rate of uniform asymptotic normality are obtained. The convergence rate can reach $ O(n<^>{-1/6}) $ O(n-1/6) by choosing appropriate parameters. These results extend the existing results in the literature, and the validity of the theoretical results is verified by numerical simulations.
Real-world data typically have indeterminacy, uncertainty, and ambiguity, rendering classical statistical methods unreliable for inference. A flexible solution for addressing imprecise information is the neutrosophic framework. Estimation methods for crisp data are well-established, although neutrosophic estimators are still being developed. This paper introduces a novel exponential log-type ratio-cum-product estimator and a broader class of estimators for finite population mean estimation with uncertain and indeterminate data. Under the first-order approximation, we develop formulas for bias and mean squared error (MSE) and evaluate the suggested estimators' efficiency using percentage relative efficiency. Real temperature and medical data show that the method is effective in uncertain conditions. A thorough simulation investigation confirms the theoretical findings. Empirical and simulation results show that the suggested estimators have lower MSE and greater PRE than classical and existing neutrosophic estimators, establishing their superiority for real-world uncertain data analysis.
We propose a regression framework operating in the graph frequency domain to model relationships among variables observed on the vertices of a graph. The proposed model is constructed via graph filters and operates in the graph frequency domain. We derive ordinary least squares estimators for the model coefficients and establish their theoretical properties, including consistency and asymptotic normality. We further develop a statistical testing procedure for variable significance at each graph frequency and provide frequency-specific interpretations of the regression coefficients. The proposed framework serves as a foundational regression methodology for graph-indexed data. Its practical usefulness is demonstrated through a simulation study and a real data application to a trading network.
This study introduces a new transmuted-H (NT-H) family of probability distributions for modeling data on the bounded unit interval. The proposed family is generated through a flexible transmutation-based mechanism that enhances the shape adaptability of existing baseline models. Statistical inference for the proposed family is carried out using the maximum likelihood estimation (MLE) method. Several mathematical properties of the NT-H family are derived, with special emphasis on the new Transmuted Kumaraswamy (NTKw) distribution. Its structural and inferential properties are analytically investigated, and the performance of the MLEs is assessed through a Monte Carlo simulation study. To demonstrate practical applicability, the NTKw distribution is applied to two real datasets on bounded unit intervals. Its goodness-of-fit performance is compared with the classical Kumaraswamy and transmuted Kumaraswamy models. The results indicate that the NTKw model provides a comparatively better fit among the considered models, highlighting its flexibility and usefulness for modeling bounded data.
Traditional mean-variance portfolio optimization struggles with poor out-of-sample performance in large dimensions and fails to account for heterogeneous asset return distributions across time or markets. Existing methods often fall short, and simple stratification can paradoxically worsen performance by reducing effective sample size. We introduce COSPLAY (COhort and SParse-based two-LAYer oracle estimator), a novel double-penalized regression estimator. COSPLAY is specifically designed for heterogeneous and sparse large portfolios, extending the unconstrained regression framework. Its key innovations include a tailored clustering algorithm for robust cohort structure recovery, which boosts effective sample size and estimation precision, particularly for time-varying data. With established oracle properties, COSPLAY simultaneously achieves mean-variance efficiency and risk constraint satisfaction within identified groups, while effectively recovering both heterogeneity and sparsity. Extensive validation on real-world stock data confirms COSPLAY's consistent outperformance over conventional methods in Sharpe ratio and risk, demonstrating its superior capability in complex heterogeneous large portfolio management.
This research introduced a novel hybrid framework using a Fuzzy Convolutional-Super Resolution Generative Adversarial Network (FC-SRGAN) for the restoration of images. The input image is first obtained from a predefined dataset. The Deep Kronecker Network (DKN) is then employed to identify noise pixels, while a statistical model is used to eliminate the unwanted noise. Then image inpainting is accomplished using Context-Conditional Generative Adversarial Networks (CC-GAN) combined with an encoder optimized by the Jaya Waterwheel Plant Algorithm (JWWPA). The JWWPA is the fusion of Jaya optimization and the Waterwheel Plant Algorithm (WWPA). Finally, image restoration is done using the FC-SRGAN. The developed FC-SRGAN method obtained the highest value for Peak Signal-to-Noise Ratio (PSNR) as 38.16 dB, the Second-Derivative-like Measure of Enhancement (SDME) as 59.70 dB, the Structural Similarity Index (SSIM) as 0.799, Universal Quality Index (UQI) as 0.856, and the Figure of Merit (FOM) as 0.982.
Under the almost optimal moment conditions, the general results on complete convergence and complete moment convergence for randomly weighted sums of m-asymptotic negatively associated (m-ANA, for short) random variables are established in this paper. The main results obtained extend and improve the corresponding ones in the literature. As an application, the strong law of large numbers for the randomly weighted estimation of conditional Value-at-Risk (CVaR, for short) is given.
We propose a three-step kernel estimation method to estimate the cumulative distribution function of unobserved noises in autoregressive time series with an unknown trend. In the first step, we estimate the unknown trend function using a modified moving average method and then remove it. In the second step, we estimate the autoregressive coefficients using the Yule-Walker method based on the detrended time series, and compute the residuals. Finally, in the third step, we estimate the cumulative distribution function based on the residuals obtained in the second step. The proposed estimator is oracle efficient under mild assumptions, meaning that it is asymptotically as efficient as the infeasible estimator based on the unobserved noise sequence. We also construct a simultaneous confidence band based on the estimation. To validate our asymptotic theorems, we provide simulation examples. Additionally, we apply our proposed method to the annual global surface air temperatures from 1880 to 1985.
Risk analysis and risk management are becoming more data-driven disciplines. This evolution is posing significant challenges to risk analysts and risk managers, who now need to apply quantitative skills for increasing complexity. We cover in this paper an analogy between risk analysis and risk management and digital twins, which are now widely implemented in running complex systems. This analogy can prove mutually beneficial to system managers and risk managers in that it uncovers areas of overlap where combined efforts can prove synergistic. We dedicate a whole section to digital twins for risk management, called RMDT, and conclude with a discussion mapping possible paths where these disciplines can collaborate in the future.
In a linear regression setting, we rigorously explore tests for a possibly multidimensional suspected outlier under a very general framework that can account for a wide range of alternative hypotheses. We consider tests that are (a) based on the standard approach for testing linear hypotheses, and (b) motivated by Cook's distance. These are seen to cover some other intuitively appealing approaches as well. We also indicate how, in special cases, the tests in (a) or (b) relate to common procedures for outlier detection. Even though the test in (a) is dictated by the standard approach for testing linear hypotheses, we find that the test in (b) can have a larger power over an appreciable part of the parameter space under the alternative hypothesis. It is also observed that Cook's distance, as it stands, does not in general entail an F-test, so that a modification thereof is called for.
This paper introduces two innovative measures, namely Weighted Cumulative Past Extropy Inaccuracy and Weighted Dynamic Past Cumulative Extropy Inaccuracy. These measures serve as alternatives to traditional extropy-based measures based on the cumulative distribution function. We investigate the characterization of these measures within the framework of the proportional reversed hazard model, exploring their applications in identifying well-known distributions. Additionally, we discuss the stochastic ordering of the proposed measures. Nonparametric estimators for these measures are suggested using kernel and empirical methods. Finally, the application of these measures in model selection is comprehensively examined, with empirical results demonstrating their potential to enhance modelling accuracy.