The asymmetric Laplace (ALP) distribution is commonly used in quantile regression (QR) due to its analytical convenience. However, its moderately heavy tails make it unsuitable for data that strongly deviates from the Gaussian assumption. In this article, we reparameterize a skewed generalized normal (SGN) distribution, with a skewness parameter defined on the interval (0, 1), as a robust alternative to the ALP distribution within a likelihood-based QR framework. We present two stochastic representations of the SGN distribution: one for random sample generation and another that enables efficient maximum likelihood estimation via the Expectation-Maximization (EM) algorithm. By representing the SGN distribution as a scale mixture of skewed normal (SKN) distributions, we derive closed-form updates for the regression parameters. Building on the EM framework, we further develop a case-deletion diagnostics analysis for QR models. The effectiveness and robustness of the proposed method are demonstrated through extensive experiments on both simulated and real-world datasets.
The classical Mixture of Experts (MoE) model is a widely employed framework for modelling nonlinear regression relationships and addressing data heterogeneity, classification, and clustering. However, most existing linear MoE models assume that the expert components follow a Gaussian distribution, which makes them sensitive to outliers and potentially ill-suited for datasets containing groups with skewed or heavy-tailed distributions. In this paper, a general skewed generalized normal (GSGN) distribution to model asymmetric data is proposed. Expanding upon the classical MoE model, a novel robust non-normal mixture of experts model is established, termed MoE-GSGN, based on the GSGN distribution. An analytical Expectation-Maximization (EM) algorithm is derived to obtain the maximum likelihood estimates (MLEs) of the model parameters. Simulation studies are conducted to investigate the performance, effectiveness, and robustness of the proposed methodology. Finally, the practical applicability of the MoE-GSGN model is demonstrated through its application to a real tone perception dataset.
This paper develops an optimal model averaging approach for linear measurement error models, when the response variable is subject to random right censoring. Within this context, we propose a bias-corrected weighted least squares estimation for unknown regression parameters. A novel Mallows-type weight choice criterion that skillfully bypasses the unavailable true covariates is developed for allocating model weights. We provide two theoretical justifications for our proposal. First, we establish the asymptotic optimality property of the resulting model averaging estimator when all candidate models are misspecified. Second, we show our model averaging method estimates the model parameters with root-n rate when the true model is among the candidate models. Particularly, it is demonstrated that the optimality is still valid when a model screening strategy is conducted prior to model averaging. Simulation studies and a real data analysis highlight the superiority of the proposed method.
This paper considers estimation problem in multivariate regression models. Under this framework, we develop a novel two-stage model averaging procedure. In the first stage, we construct a scalable model averaging estimator which involves transforming the original model based on the singular value decomposition. When the dimension of the regressor vector is K, this approach enables us to average the estimators from the candidate model set of size K instead of size 2K. The second stage is to find the optimal weights for averaging by applying a weight choice criterion from Kullback-Leibler distance. We prove that the minimum weighted squared loss from the scalable model averaging is asymptotically the same as that from original model averaging, further demonstrate asymptotic optimality of the scalable model averaging estimator using Kullback-Leibler-distance-based weights, and derive the rate of the resulting weights tending to the risk-based optimal weights. In comparison with existing model averaging methods, the simulation results show that, in terms of weighted mean squared prediction error and computation time, our proposal is more efficient, especially under the situation where the number of candidate models is large and the sample size is small. Moreover, a real data analysis is provided to illustrate the application of our method in practice.
In this paper, we propose a flexible scale mixtures of the skewed generalized normal (FSMSGN) distributions, which encompasses several well-known asymmetric distributions as special cases. Within this general framework, we derive key distributional properties and introduce an efficient ECME-PLA algorithm that combines the Profile Likelihood Approach (PLA) with the classical Expectation/Conditional Maximization Either (ECME) algorithm to obtain maximum likelihood estimates (MLEs) of the model parameters, overcoming the convergence challenges faced by traditional methods. We further extend the method to linear regression models with FSMSGN-distributed errors and derive the corresponding updating equations for the regression coefficients. The asymptotic properties of the proposed estimators are also investigated. A comprehensive simulation study demonstrates the strong performance of the proposed approach, particularly in modeling skewed, heavy-tailed, and peaked data. The practical effectiveness of the method is illustrated through applications to three real-world datasets.
In this paper, we constructed a new type of skewed generalized normal distribution whose density function contains four parameters, which can give the distribution excellent flexibility and provide an accurate model for actual statistical data analysis. Several important properties of the distribution are presented, and the optimal confidence interval for the distribution parameters was constructed. Through simulation experiments, we verified that the confidence interval we constructed has high accuracy and reliability. In addition, the hypothesis testing problem for this distribution, including the analysis of power functions, has also been discussed in detail. The performance of testing methods under various situations is evaluated, and the optimal two-tailed testing method is proposed. The distribution and statistical analysis methods presented in this article can greatly improve the quality of data analysis.
The skew-t distribution, characterized by skewness and heavy tails, has been widely applied in modeling financial, economic, and stock market data. In this article, we proposed several goodness-of-fit test methods designed for the skew-t distribution. These tests were grounded in probabilistic principles that establish connections between the skew-t distribution, the t distribution, and the F distribution. We achieved this by transforming random samples into approximate t-distributed and F-distributed variables. Subsequently, we introduced the Anderson-Darling (AD) test statistic and the sample correlation coefficient test for each variable transformation. Critical values for various sample sizes and significance levels were determined using the bootstrap method. To assess the effectiveness of our proposed methods, we compared the test power of these new tests with that of the classical Anderson-Darling (AD) and Kolmogorov-Smirnov (KS) tests, performed before the variable transformation, across different sample sizes and under various alternative distributions. Simulation results indicated that the F-transform-based and classical untransformed AD tests generally performed competitively in terms of power against the considered alternative distributions.
In many fields, limited or censored data are often collected due to limitations of measurement equipment or experimental design. Commonly used censored linear regression models rely on the assumption of normality for the error terms. However, this approach has faced criticism in literature due to its sensitivity to deviations from the normality assumption. In this paper, we propose an extension of the CR model under the two-piece generalized t (TPGT)-error distribution, called TPGT-CR model. The TPGT-CR model offers greater flexibility in modeling data by accommodating skewness and heavy tails. We developed a modified maximum likelihood (MML) estimator for the proposed model and introduced the modified deviance residual to detect outliers. The developed MML estimator under the TPGT assumption possesses several appealing merits, including robustness against outliers, asymptotic equivalence to the maximum likelihood estimator, and explicit functions of sample observations. Simulation studies are conducted to examine the finite sample performance, robustness, and effectiveness of both the classical and proposed estimators. The results from both the simulated and real data illustrate the usefulness of the proposed method.
The proliferation of complex data across various domains has led to an increasing prevalence of skewed and heavy-tailed distributions, where traditional symmetric distributions become inadequate. While existing approaches like the skew-normal distribution have addressed this challenge, they remain constrained by limited ranges of skewness and kurtosis, particularly in handling extreme cases. In this study, we introduce the Skewed Generalized Normal (SGN) distribution, a novel and highly flexible distribution class capable of capturing an extensive range of skewness and kurtosis while maintaining mathematical tractability. We establish its theoretical foundation through comprehensive analysis of its properties and develop two robust parameter estimation methods: a classical Expectation-Maximization (EM) algorithm and an innovative deep learning-based approach called EstiFormer. The latter represents a significant advancement in parameter estimation, offering a pre-trained, adaptable model for rapid and precise estimation across diverse data scenarios. Through extensive simulation studies with varying sample sizes, we demonstrate that both estimation methods achieve high accuracy, with the SGN distribution significantly outperforming existing alternatives in fitting highly skewed and kurtotic data. Applications to real-world datasets further validate the SGN distribution’s superior flexibility and practical utility.
We propose a novel dynamic generalized Pareto distribution (GPD) framework for modeling the time-dependent behavior of the peak over threshold (POT) in extreme smog (PM2.5) time series. First, unlike static GPD, three dynamic autoregressive conditional generalized Pareto (ACP) models are introduced. Specifically, in these three dynamic models, the exceedances of air pollutant concentration are modeled by a GPD with time-dependent scale and shape parameters conditioned on past PM2.5 and other air quality factors (SO2, NO2, CO) and weather factors (daily average temperature, average relative humidity, average wind speed). Second, unlike the recent studies of ACP models, we impose a logistic function autoregressive structure on the scale and shape parameters of the ACP models, which has simple calculation and flexible modeling for the scale and shape parameters, since the logistic function is used to mean that the changes in the long memory parameter occur in a continuous manner and often applied in time series models. Third, the model averaging method is applied to improve predictive performance using AIC and BIC criteria to select combined weights of the three ACP models. In addition, based on goodness-of-fit tests, the thresholds of the three ACP models are chosen by eight automatic threshold selection procedures to avoid subjectively assigning a certain value as the threshold. Maximum likelihood estimation (MLE) is employed to estimate parameters of the ACP models and its statistical properties are investigated. Various simulation studies and an example of real data in PM2.5 time series demonstrate the superiority of the proposed ACP models and the stability of the MLE.
Partially linear models (PLMs) are widely employed in scientific research to analyze hybrid parametric-nonparametric relationships. However, their conventional reliance on symmetric error distributions severely limits applicability to real-world phenomena characterized by pronounced asymmetry and heavy-tailed behavior. To address this gap, we propose a novel PLM framework incorporating skewed generalized normal (SGN) distributed errors, which simultaneously accommodates extreme skewness and heavy-tailed attributes beyond the capabilities of symmetric or skew-normal (SN) specifications. Methodologically, we develop a penalized expectation-maximization (EM) algorithm with provable convergence guarantees and integrated adaptive smoothing selection, effectively resolving optimization instability in high-dimensional settings. Furthermore, we establish a unified diagnostic system that synergizes geometric leverage calculus with local influence analysis to systematically evaluate model robustness against perturbations and outliers. Extensive simulation studies demonstrate the framework's superior estimation accuracy compared to conventional symmetric and SN-based alternatives. Empirical validations based on real-world datasets reveal statistically significant improvements in model fit while maintaining interpretability.
This article is concerned with model averaging for de-noise linear models, in which some covariates are not observed, but their ancillary variables are available. The least-squares-based estimation procedure is used to estimate the unknown regression parameter in each candidate model after the calibrated error-prone covariates are obtained. Then a Mallows-type weight choice criterion is constructed. When all candidate models are misspecified, the model averaging estimator is asymptotically optimal in the sense that achieving the lowest possible squared error. On the other hand, when the true model is included in the set of candidate models, the model averaging estimator of the regression parameter is root n consistent. The finite sample performance of our model averaging estimator is evaluated by some simulation studies. The proposed procedure is further applied to real-data analysis.
In this paper, we introduce a family of distributions known as generalized scale mixtures of asymmetric generalized normal distributions (GSMAGN), characterized by remarkable flexibility in shape. We propose a novel finite mixture model based on this distribution family, offering an effective tool for modeling intricate data featuring skewness, heavy tails, and multi-modality. To facilitate parameter estimation for this model, we devise an ECM-PLA ensemble algorithm that combines the Profile Likelihood Approach (PLA) with the classical Expectation Conditional Maximization (ECM) algorithm. By incorporating analytical expressions in the E-step and manageable computations in the M-step, this approach significantly enhances computational speed and overall efficiency. Furthermore, we persent the closed-form expressions for the observed information matrix, which serves as an approximation for the asymptotic covariance matrix of the maximum likelihood estimates. Additionally, we expound upon the corresponding consistency characteristics inherent to this particular mixture model. The applicability of the proposed model is elucidated through several simulation studies and practical datasets.
Intuitionistic fuzzy sets provide a viable framework for modelling lifetime distribution characteristics, particularly in scenarios with measurement imprecision. This is accomplished by utilizing membership and non-membership degrees to accurately express the complexities of data uncertainty. Nonetheless, the complexities of some cases necessitate a more advanced approach of imprecise data, motivating the use of generalized intuitionistic fuzzy sets (GenIFSs). The use of GenIFSs represents a flexible modeling strategy that is characterized by the careful incorporation of an extra level of hesitancy, which effectively clarifies the underlying ambiguity and uncertainty present in reliability evaluations. The study employs a methodology based on generalized intuitionistic fuzzy distributions to thoroughly examine the uncertainty related to the parameters and reliability characteristics present in the Burr XII distribution. The goal is to provide a more accurate evaluation of reliability measurements by addressing the inherent ambiguity in the distribution’s shape parameter. Various reliability measurements, such as reliability, hazard rate, and conditional reliability functions, are derived for the Burr XII distribution. This extensive analysis is carried out within the context of the generalized intuitionistic fuzzy sets paradigm, improving the understanding of the Burr XII distribution’s reliability measurements and providing important insights into its performance for the study of various types of systems. To facilitate understanding and point to practical application, the findings are shown graphically and contrasted across various cut-set values using a valuable numerical example.
In this paper, a novel skewed generalized normal distribution (NSGN), which encompasses several well-known distributions, is a two-component mixture model based on the generalized normal distribution. The NSGN distribution offers flexibility in modeling skewed data characterized by high kurtosis and strong skewness, making it highly applicable in finance, medicine, and engineering. The primary characteristics and properties of this new distribution are introduced, and the explicit expressions for the moments of order statistics are provided using recursive relations. The maximum likelihood estimation based on a profile likelihood approach, L-moments estimation, and a two-step estimation method combining Bayesian posterior maximum estimation with moment estimates are presented to estimate the four parameters of NSGN distribution. A simulation study compares the performance of these methods across different sample sizes and different parameters. The appropriateness of the proposed distribution has been tested by comparing it with several skewed distributions. Additionally, applications to real datasets showcase the beneficial properties of the NSGN for applied statistical research.
In recent years, some interested data can be recorded only if the values fall within an interval range, and the responses are often subject to censoring. Attempting to perform effective statistical analysis with censored, especially heavy-tailed and asymmetric data, can be difficult. In this paper, we develop a novel linear regression model based on the proposed skewed generalized t distribution for censored data. The likelihood-based inference and diagnostic analysis are established using the Expectation/Conditional Maximization Either algorithm in conjunction with smoothing approximate functions. We derive relevant measures to perform global influence for this novel model and develop local influence analysis based on the conditional expectation of the complete-data log-likelihood function. Some useful perturbation schemes are discussed. We illustrate the finite sample performance and the robustness of the proposed method by simulation studies. The proposed model is compared with other procedures based on a real dataset, and a sensitivity analysis is also conducted.
Recently, a unique extension of fuzzy sets known as hesitant fuzzy sets has been established to address hesitant cases that previous methods were unable to manage adequately. In this paper, the triangular hesitant fuzzy sets approach has been employed to explore the inherent uncertainty within the parameters of the life distribution. Two essential reliability measures, triangular hesitant fuzzy reliability and the hazard rate function designed for the Pareto Type I life distribution, have been established. Moreover, the triangular hesitant fuzzy reliability measure is utilized to assess the reliability of series and parallel systems. Furthermore, the weighted averaging operator has been used on both the series and parallel systems, making them more reliable and giving much better results than hesitant fuzzy sets. Finally, a numerical example demonstrating the use of these techniques is provided, and the results are presented in tabular and graphical formats.
This paper is concerned with optimal model averaging procedure for semiparametric partially linear models where some covariates are subject to measurement error. We proposed a corrected semiparametric generalized least squares estimation for unknown parameters and nonparametric function, and developed a Mallows-type criterion for weight choice. The resulting model average estimator is shown to be asymptotically optimal in terms of achieving the smallest possible squared error under some regularity conditions. The simulation studies demonstrate that the proposed procedure is superior to traditional model selection and model averaging methods. Our approach is further applied to Ragweed Pollen Level data.
Manufacturers today prioritize increasing output quality while reducing per-unit costs to remain competitive.Statistical Process Control (SPC) is a useful tool for boosting quality and output. However, decision-makers in industries facedifficulties in cost and quality control due to uncertainty in data. Fuzziness can be applied to uncertain data to manage theeconomic design of a control chart, allowing for better control over the control chart's economic design. This study seeks toenhance the economic design of X control charts by using fuzzy set theory, specifically utilizing triangular fuzzy numbers forcost parameters and the signed distance approach for defuzzification. The objective is to enhance the adaptability of thesecharts in unpredictable situations. The proposed model enables the incorporation of cost parameters inside a fuzzy framework,with the objective of minimizing the control chart while adhering to the permissible limit. The applicability and improvedaccuracy of the model in optimizing control chart parameters for a glass bottle production process are exemplified by aneffective illustration. This study highlights the need of integrating fuzzy logic into SPC. It suggests a technique that improvesthe cost-effectiveness and operational effectiveness of control charts in situations when data is uncertain and imprecise.
The skewed generalized normal (SGN) distribution with four parameters is a versatile distribution that can effectively model data with skewness and heavy or light tails. In this paper, we conduct two classes of goodness of fit tests for the SGN distribution based on the empirical distribution function (edf) and the sample correlation coefficient. The first class involves transforming the sample into approximately mixed gamma observations, and then applying five classical parametric bootstrap edf-based goodness of fit tests. The second class is based on the inverse probability transformation and utilizes the sample correlation coefficient as the test statistic. We compare the finite sample performances of the proposed tests for different sample sizes and alternative distributions by extensive numerical studies. The simulation results demonstrate that the proposed tests provide a valid alternative to the standard tests using the original data, and the analysis of real data illustrates its application.