
center dot The family of Unified Skew Normal distributions is associated with probability distributions encountered in various problems, notably those of the selection of individuals in a normal population. In this work, we focus on the ordering of the real subfamily of this class of distributions with respect to its parameters and for some stochastic orders (usual order, increasing convex order, increasing concave order and likelihood ratio order). Our results are applied to a reliability problem as well as a selection problem.
In this paper, we present a Bayesian variable selection method for zero-inflated longitudinal data. For this purpose, we consider a zero-inflated power series random effects model that includes the zero-inflated Poisson and negative binomial random effects models. We propose using continuous spike and Dirac spike priors to simultaneously estimate the regression coefficients and select the important covariate variables. We apply the MCMC method using Gibbs sampling for posterior inference. Some simulation studies are performed to investigate the performance of the proposed approach, and it is also applied to analyze a real dataset from the RAND Health Insurance Experiment.
center dot Misuse of statistical significance continues to be prevalent in science. The absence of intuitive explanations of this concept often leads researchers to incorrect conclusions. For this reason, some statisticians suggest adopting S-values (surprisals) instead of P-values, as they relate the statistical relevance of an event to the number of consecutive heads when flipping an unbiased coin. This paper introduces the concept of surprisal intervals (S-intervals) as extensions of confidence/compatibility intervals. The proposed approach imposes the assessment of outcomes in terms of more and less surprising than some values, instead of statistically significant and statistically non-significant. Moreover, a novel methodology for presenting multiple consecutive S-intervals (or compatibility intervals as well) in order to evaluate the variation in surprise (or compatibility) with various target hypotheses is discussed.
This paper is the next step ahead in constructing probability distribution of changeable flatness of PDF that is expressed with well-known kurtosis measure. The distribution in question is named the Extended Easily Changeable Kurtosis and descends from the Easily Changeable Kurtosis. The paper covers PDF, CDF, modes and inflection points, quantiles, moments and Moors' measure and the Fisher Information Matrix. In addition generator of pseudo-random numbers that follow the Extended Easily Changeable Kurtosis is presented. Unknown parameters of the distribution are estimated with the maximum likelihood method. The paper ends with illustrative examples of applicability and flexibility of the new distribution. The most important R codes are presented in the supplementary material.
center dot Model specification and selection are important aspects of modeling exercises. In this context, the Fisher Information Matrix (FIM) plays an essential role. In this paper, we derive the Fisher Information Matrix (FIM) for the two way random effects panel data models in general as well as in some specific cases of heteroscedasticity. Some computational issues are then raised and discussed. In addition, some real data examples are reported and thoroughly discussed.
This paper describes the response surface methodology for mixed-level factors of the form s(1)(n1) x s (n2)(2) when experimental units experiences overlap effects from the adjacent neighbouring units. Conditions have been derived for the near orthogonal estimation of the parameters of the model and ensuring the constancy in the prediction variance. A method of construction of designs satisfying derived conditions has been developed. Some particular cases of s(1)(n1 )x s(2)(n2) has also described. An R package named rsdNE has also been developed for the generation of these designs.
center dot We address the estimation of extreme quantiles of Weibull tail-distributions. Since such quantiles are asymptotically larger than the sample maximum, their estimation requires extrapolation methods. In the case of Weibull tail-distributions, classical extreme-value estimators are numerically outperformed by estimators dedicated to this set of light-tailed distributions. The latter estimators are based on two key quantities: an order statistic to estimate an intermediate quantile and an estimator of the Weibull tail-coefficient used to extrapolate. The common practice is to select the same intermediate sequence for both estimators. We show how an adapted choice of two different intermediate sequences leads to a reduction of the asymptotic bias associated with the resulting refined estimator. This analysis is supported by an asymptotic normality result associated with the refined estimator. A data-driven method is introduced for the practical selection of the intermediate sequences and our approach is compared to three estimators of extreme quantiles on simulated data. An illustration on a real data set of daily wind measures is also provided.
Composite type distributions are being increasingly used to model insurance data. Yet no expressions seem to be available for observed information matrices. In this paper, we give expressions for the matrices for two-piece, three-piece and m-piece composite distributions in their most general forms. Expressions for a number of particular cases and a simulation study showing practical use are also given.
The negative binomial (NB) regression model is commonly used to model overdispersed count data. However, the NB regression model is not suitable for highly overdispersed data, for which the Poisson-inverse Gaussian (PIG) regression model is often used instead. The maximum likelihood (ML) estimator is typically used to estimate the coefficients of the PIG regression model. However, when multicollinearity exists among the explanatory variables, the ML estimator's variance can become inflated. To address this issue, we propose PIG ridge regression (PIGRR) and quantile-based ridge regression estimators for the PIG regression model. We also suggest using a Wald-type method to calculate the confidence interval on the mean response function of the PIGRR. To evaluate the performance of these proposed methods, we conducted a Monte Carlo simulation study, considering mean squared error and average confidence lengths as performance criteria. Additionally, we analyzed the traffic fatalities dataset to demonstrate the benefits of the proposed estimators for practitioners dealing with multicollinearity issues in real datasets.
In this article, the bivariate exponential distribution proposed by Downton (“Bivariate exponential distributions in reliability theory”, Journal of the Royal Statistical Society Series B, 1970) is extended to a bivariate Weibull distribution, and it is called the Downton’s bivariate Weibull (DBW) distribution. Statistical properties of the DBW distribution are explored and likelihood inference developed based on complete as well as right-censored bivariate data are discussed. Through extensive Monte Carlo simulations, performance of the point and interval estimates are evaluated. Two real datasets are analyzed for illustrative purposes. It is concluded that the DBW distribution is very useful to model bivariate data.
Zhang, Leng and Tang (2015) propose joint parametric modelling of the means, variances, and the correlations by decomposing the correlation matrix via hyperspherical co-ordinates and show that this results unconstrained parameterization, fast computation, easy interpretation of the parameters, and model parsimony. With unconstrained structures, they also suggest future research on modelling the mean, the variance, and the correlations non-parametrically and semiparametrically. In this paper we explore semiparametric modelling via simulations and data analysis. Extensive simulations show that the semiparametric modelling produces similar bias and efficiency properties of the parameter estimates as those by the parametric modelling. However, model selection, using the AIC and the BIC, through the analysis of two real biomedical data sets show significant improvement in model parsimony.
The distribution of the sum of two independent Lindley random variables having the same parameter defines the 2S-Lindley distribution recently introduced in the statistical literature. In this paper, we use it to create a new one-parameter discrete compound distribution called the Poisson 2S-Lindley distribution. More precisely, it is obtained by compounding the Poisson and 2S-Lindley distributions with the idea of combining their desirable functionalities. The mathematical and statistical properties of the proposed distribution are investigated and analyzed systematically. Subsequently, a statistical standpoint is adopted. Various estimation methods are considered, and an extensive simulation study is conducted to compare them. A count regression model as well as a first-order integer-valued autoregressive process based on the proposed model are constructed. In total, four real data sets in different fields are used to prove the empirical importance of the distribution.
Several works concerning the utilization of Partial Least Squares as a supervised dimension reduction technique have been developed over the years in the field of chemometrics, among others, for regression purposes. However, Partial Least Squares can be a challenging procedure especially in the case of multivariate multiple regression due to data characteristics and complexity. Thus, in this work we propose the use of Partial Least Squares method as a variable selection technique in linear regression tasks that involve high dimensional spectral data sets. More precisely, we suggest the exploitation of the regression coefficients that Partial Least Squares estimates in order to identify and eject the insignificant predictor variables from the analysis. In such manner we are able to remove the uninformative variables and obtain in most cases better results than the classical Partial Least Squares regression but with simpler structure. We compare our proposed technique with the classical Partial Least Squares and Principal Component Analysis in both univariate and multivariate regression.
This paper describes the response surface methodology for mixed-level factors of the form S1n1 x S2n2 when experimental units are affected by the adjacent neighbouring units. Conditions have been derived for the orthogonal estimation of the parameters of the model. A method of construction has been developed and the developed designs satisfy the rotatability conditions. Some particular cases of S1n1 x S2n2 are described. An R package named rsdNE developed for the generation of these designs has also been discussed.
We study the distributional properties of the product of random stochastic matrices by using the Dirichlet distribution. Our observations widely generalize some known results on the randomly weighted averages about multivariate Dirichlet distributions. We present a new method for approximating the distribution that can be used for the multivariate random variables with a slight change.
The skew-normal distribution and some of its extensions have been considered in the last two decades in view of distribution theory and the associated properties. However, less attention has been paid to other aspects of this family of distributions. In this paper, we focus on the information properties of this distribution and the distributions of order statistics of a simple random sample from the skew-normal distribution. The Shannon’s entropy as well as Kullback-Leibler divergence between the order statistics of two independent skew-normal distributions are studied. Some interesting properties of the information measures of different order statistics are presented.
Ignoring temporal dependence when modelling sequences of extreme observations yields underestimated standard errors which can lead to inaccurate risk assessment of extreme phenomena such as floods and economic crises. One remedy is to inflate standard errors with a sandwich or block bootstrap estimator. In this study, the performance of four such standard error estimators is investigated, through simulation, when modelling extremes from bivariate sequences. The results show that under strong temporal dependence, all considered estimators seriously underestimate standard errors, while under moderate to weak dependence both the sandwich and the bootstrap estimators can mitigate this underestimation.
The target of this paper is to generalize the Heat Equation, highly related with the Normal distribution. Therefore a generalization of the Normal distribution, the 7-order Normal distribution is introduced, which influences the entropy type information measures and offers the generalization of the Heat Equation.
A transformation of a density function is introduced to derive two families of continuous densities, the first symmetric and the second not-necessarily symmetric, exhibiting both unimodality and bimodality. Their respective density functions are provided in closed form, allowing us to simply obtain moments and related quantities. We focus on the case where the normal distribution is considered, although it can be applied to other models, such as the logistic and Cauchy distributions. This transformation is also extended to derive a family of asymmetric unimodal and bimodal distributions via Azzalini’s scheme. An example related to environmental science illustrate these models’ practical performance.
It is long familiar that the stratified ranked set sampling (SRSS) is more efficient than ranked set sampling (RSS) and stratified random sampling (StRS). The existence of missing values may alter the final inference of any study. This paper is a fundamental effort to suggest some combined and separate imputation methods in the presence of missing data under SRSS. The proposed imputation methods become superior than the mean imputation method, ratio imputation method, Diana and Perri (2010) type imputation method, and Sohail et al. (2018) type imputation methods. A simulation study is administered over two hypothetically drawn asymmetric populations.