
The formulation of generalized class of distributions for modeling and analyzing data in diverse domains is of enormous practical significance. Captivated by the need for greater flexibility and relevance when modeling data in practice, academics in probability and distribution theory have proposed, studied and implemented novel techniques of generating new distributions from existing models. The XGamma distribution has serious shortfalls when it comes to applications in real world data set. This has attracted researchers to develop more generalized XGamma distributions to provide the desired outcomes in applications. The focus of this curated compilation constitute freshly conceived illuminating extensions, generalizations and modifications of the XGamma distribution. This is expected to serve as a vibrant platform and future research direction for researchers in probability and distribution theory.
In this paper, some representations of the Bivariate Pareto Distributions were studied for their mathematical properties. A numerical study was undertaken to compare the performance of some Archimedean Copulas for these bivariate pareto representations. Details are provided in the main text of the paper.
Reviewer Acknowledgements for International Journal of Statistics and Probability, Vol. 15, No. 2
Sample size plays a fundamental role in statistical inference using normal and t-tests, yet inappropriate sample sizes can lead to serious inferential distortions. Large samples tend to produce artificially small p-values and overly narrow confidence intervals, resulting in test bias and potentially misleading conclusions. In contrast, small samples often yield unstable variance estimates, reducing the reliability and power of statistical tests. This paper addresses both extremes within a unified framework. For large samples, we develop the Random Group Method (RGM) and a correction factor approach to mitigate bias by stabilizing variability across subgroups. For small samples, we introduce pseudo-sample expansion techniques, emphasizing a mean-based method that preserves expectation while reducing variance inflation caused by duplication. A key contribution of this work is the formulation of a relative variance (RV) criterion for determining a “good” sample size. We show that the optimal sample size is not a single value but lies within an interval over which statistical inference remains stable. Theoretical results are supported by numerical examples illustrating the relationship between sample size, variance control, and hypothesis testing behavior. The proposed methods provide practical tools for improving statistical inference in both traditional settings and modern large-scale data applications, including artificial intelligence and survey analysis.
This study proposes a novel test statistic for the two-sample problem involving sub-mean vectors under a two-step monotone missing data structure. The proposed procedure is constructed based on the structure of Rao's U-statistic by combining a Hotelling's T^2-type statistic for monotone missing data with the standard Hotelling's T^2 statistic, thereby efficiently utilizing the available information in incomplete observations. We consider the problem of testing the equality of sub-mean vectors between two populations under the assumption that a subset of the mean components is common. The asymptotic expansion of the null distribution of the proposed statistic is derived, and its distribution function and approximate upper percentiles are obtained. To improve the accuracy of the chi-squared approximation in finite samples, Bartlett and Bartlett-type correction methods are also developed. The performance of the proposed approximations and correction procedures is investigated through extensive Monte Carlo simulations under various dimensional and sample size settings. A numerical example based on real data is presented to illustrate the applicability and practical usefulness of the proposed methodology.
Multicollinearity is a pervasive challenge in regression analysis that inflates variance estimates and obscures the interpretation of predictor effects. This study develops intuitive geometric approaches for addressing multicollinearity within an observation-axes framework. We examine the structure of multicollinearity and introduce three complementary geometric approaches. First, we present a simple and effective method for resolving multicollinearity when the causal ordering among predictors is known. Second, we refine and extend the geometric interpretation of Principal Component Analysis (PCA) proposed by Wickens (2014) by incorporating angular information between predictor variables. Third, motivated by Partial Least Squares (PLS) regression, we develop a geometry-based method that identifies directions jointly determined by predictor variance and alignment with the response. Together, these solutions demonstrate that geometric reasoning provides richer and more intuitive insights into variable relationships, offering substantial pedagogical and methodological value as a resource for beginners and as a foundation for future research in regression analysis.
This study analyzes community and ancestral mathematical knowledge about volume among students in Indigenous and general primary schools in the state of Hidalgo, Mexico. The objective was to estimate the expected probability of presence of this knowledge (Y2), considering contextual demographic and cultural factors. Eleven explanatory variables were considered (age, sex, community, municipality, economic community activities, parents and grandparents, parents' and grandparents' education, and school activities), and several counting models were adjusted and compared (Poisson, quasi-Poisson, and negative binomial). Finally, a negative binomial regression model with a log link function was selected. Parameters were estimated using maximum likelihood (Fisher-Scoring), and fit was assessed using deviance, AIC, and the D/df ratio. Only three predictors were statistically significant: age (β1=-0.09, p=0.07), community (β2=-0.61, p<0.001), and school activities (β3=+0.33, p<0.001). The adjusted deviation of 0.10006 and the lowest AIC (1299.9) indicates a good fit. When transforming the predictions to the original scale, no combination of variables sufficiently raises Y2 to be considered representative of the presence of knowledge in the communities studied. Although variables such as community and school activities had a significant influence, their practical impact is limited. Whereas the negative binomial model proved to be adequate for explaining the relationship between demographic and contextual variables and volume of knowledge, the coefficients suggest that these factors do not impact the practical representativeness of knowledge.
Analyzing marine animal movement data is essential for understanding at-sea behavior. This study introduces an autoregressive spectral analysis framework for assessing time-varying movement dynamics in northern fur seals. A time segmentation approach, based on an AR(3) model with Yule-Walker estimation, is employed to estimate movement parameters and characterize behavioral variability over time. The method captures temporal changes in movement persistence and oscillatory patterns, enabling the segmentation of behavioral states such as resting, exploratory diving, and active foraging. Using high-resolution vertical velocity data from eight northern fur seals tagged at the Pribilof Islands, Alaska, the analysis shows that a 26-minute window with 50% overlap achieves stationarity while preserving behavioral transitions. Spectral analysis identifies mid-frequency oscillations associated with active diving, low-frequency signals corresponding to exploratory diving, and spectral shifts indicative of behavioral transitions. Comparisons with non-parametric methods highlight the advantages of AR(3)-based spectral estimation in producing smooth and interpretable frequency-based insights. The framework provides new perspectives on fur seal foraging strategies and behavioral adaptations to environmental conditions. It also offers a computationally efficient alternative to state-space models, with potential for broader application in studying movement patterns of marine predators to support ecological and conservation research.
Data aggregation effects are examined in relation to both statistical and machine learning approaches to data modeling. It is shown that heavily data-centric artificial neural network and random forest methods are subject to aggregation effects similar to those affecting statistical methods. Several basic examples are discussed.
Competing risks analysis based on the pseudo-observations has been applied to medical studies in recent years. The analysis allows the direct evaluation of the effect of covariates on the cause-specific cumulative incidence function (CIF) using an estimating equation. In a case with a large the number of covariates, variable selection (selecting the covariates that affect the outcome variables) is critical. However, there are few studies addressing the problem of variable selection for the pseudo-observations technique. In this study, we propose two variable selection methods based on pseudo-observations. One is a method based on a criterion derived by an estimator of an expected pseudo quasi-likelihood. The other is a penalized estimating equation, inducing a sparse estimates for model parameters. The estimator given by the penalized estimating equation has so-called oracle properties under some appropriate penalty functions. When applying such a penalization method, it is essential to choose an optimal tuning parameter that determines the magnitude of the penalty. Then, we construct a BIC-type criterion for the ordinary penalized least squares and show it can consistently identify the true set of covariates as the sample size grows. Simulation studies show the performance of the proposed two variable selection methods.
When modeling more complex larger data sets, researchers/practitioners often use higher-order parametric distributions. However, such extensions pose various issues, prominently computational, and the significance of estimated parameters. For example, when estimating parameters of higher-order parametric distribution under the likelihood method, convergence problems and higher standard error of the estimated parameter can often be found. Therefore, considering these issues, we were motivated to look for a new avenue to remedy the situation. The parameter tuning method is introduced to reduce the number of parameters of higher-order parametric distribution with a minimal change of the likelihood value. Two different well-known actuarial data examples are used to demonstrate the procedure and illustrate the applicability and flexibility using well-known risk measures.
In this article, we construct an adaptive wavelet thresholding estimator for the density in a finite mixture model. Our sample is drawn from an almost periodically correlated process under a weak dependence assumption. We assess the asymptotic performance of our estimator by establishing an upper bound for the integrated mean squared error over a Besov ball.
This paper introduces a new methodology for detecting active effects in unreplicated two-level factorial experiments, which are of great importance in many scientific and practical fields. The proposed method aims to enhance the reliability and accuracy of detecting active effects compared to the popular method introduced by Lenth (1989). The new approach utilizes the F-distribution for significance testing, and it eliminates the needs of estimating the error variance or creating new critical value tables. A comprehensive simulation study was conducted using the Monte Carlo simulation method to evaluate the performance of the proposed method compared to Lenth's method in terms of the size and power of the test under different conditions. The results demonstrated the superiority of the proposed method. Additionally, three practical applications were analyzed using both methods to illustrate the practical effectiveness of the new approach. The simplicity and robustness of the proposed method make it a practical and effective method and a good choice for the analysis of unreplicated two-level factorial experiments.
Despite the significant consequences of delisting, there have been only a few studies of business failure prediction that concentrate on delisting, particularly for the U.S. market. This study aims to predict delisting using quarterly financial ratios of 8,870 companies ever listed in the major U.S. exchanges from 1970 to 2022, along with key economic indices. We construct our data set to build a delisting prediction model that is robust across different company phases and economic conditions. To enhance predictive performance, we build an ensemble model to predict delisting, integrating five widely used machine learning methods : Logistic Regression, Random Forest, Gradient Boosting, Support Vector Machine, and Neural Network. We compare the predictive performance of the five base learners and the ensemble model. The ensemble model achieves an accuracy of 0.836 and MCC of 0.457, demonstrating comparable accuracy and an improved MCC relative to the top-performing base learner, Random Forest. We leverage MCC, a reliable measure for imbalanced response data, to determine a classification threshold and comprehensively evaluate the prediction results. We find that the Price-Earnings ratio, profitability ratios and inflation are among the most informative predictors for forecasting, aligning with prior research.
In this paper, we consider a testing problem of multivariate normality (MVN). We deal with the kurtosis test statistic based on Mardia's multivariate kurtosis as an MVN test and propose a modified normalizing transformation (NT) statistic. The accuracy of the normal approximation of the proposed test statistic through a Monte Carlo simulation is investigated. The results of empirical power of a modified NT statistic are presented. Alternative distributions are chosen to represent different types of departure from multivariate normality. Moreover, to compare the empirical power of the modified NT statistic, we target a NT statistic, the improved Mardia's test statistic, and the Henzer-Zirkler test statistic. Finally, an example is provided.
In kinesiology and exercise science, researchers want to identify factors associated with players’ performance. In racquet sports, matches are played in a tournament format, and researchers often find observational data for their studies rather than collecting experimental data. Even when an experiment is conducted, which is very rare in literature, a randomized design has been considered to produce unbiased results. Given a small number of participants and limited time, in theory, an optimal experimental design can produce more statistical information about parameters of interest than a completely randomized design. This article includes simulations and a real application of an optimal experimental design to racquet sports research. In our research plan, there were some logistical considerations (e.g., no replicated matches, computation time), and simulations demonstrated that the c- and D-optimal designs result in higher statistical power for single- and multiple-parameter hypothesis testing, respectively, than the completely randomized design. For our research, the D-optimal design was applied to a doubles pickleball tournament with 16 subjects and one half of a day. Participants’ fitness and skill levels were measured in the morning, the doubles tournament was designed in real time on site using a pre-written algorithm, and the tournament design was executed on the same day. To make this computation accessible, an interactive applet is provided in the Appendix of this article with instructions.
We propose a novel method for testing the hypothesis of an eigenvector based on the exact distribution of the multiple correlationcoefficientunderanormalpopulation. Inparticular, wediscussbothnonsingularandsingularcases, addressing the relationship between sample size and the number of variables. The proposed test has the advantage of being invariant to the ordering of the target eigenvector, focusing only on whether the target vector is an eigenvector. The ordering of the eigenvector is determined by the minimum angle between the target vector and the sample eigenvector. Furthermore, we demonstrated that type I errors is exactly controlled at a particular significance level, and the power under the specified alternative hypothesis can be calculated by the Gauss hypergeometric function in the nonsingular case. Our simulation studies confirm that the empirical distribution of the test statistic is in agreement with theoretical distribution.
This article introduces an extension of the H theorem to an arbitrary order d ≥1 . The extended H theorem defines the extended entropy Hd for a given distribution and identifies the distribution function f_d(v) that maximizes entropy H_d . When d=1, H_1 corresponds to the original entropy defined by Boltzmann, and the distribution that maximizes it is the Boltzmann distribution. For d=2 , the maximizing distribution is the Rayleigh distribution. For d=3 , the maximizing distribution is the Maxwell distribution. Therefore, in the three-dimensional physical world, the correct speed distribution of gas particles is the Maxwell distribution, which maximizes the extended entropy H_3 rather than the original entropy H_1 . The result distribution f_d(v) that maximizes the extended entropy H_d for any d≥1 is derived and applied to gases that include rotational and strain energy in addition to translational energy. Numerical validation supporting these findings is also presented.
Quantile regression has become increasingly popular across various disciplines due to its robustness, offering an alternative to traditional mean-based regression. Unlike traditional linear regression, quantile regression estimates conditional quantiles, capturing the full complexity of the relationships between variables (specifically, the conditional dependence of lifetime on covariates in lifetime analysis). However, inference procedures in quantile regression often involve complex non-parametric methods, as the variance of an estimator typically depends on the unknown error density, making it difficult to estimate. In this paper, we present a bootstrap-type resampling method that simplifies the construction of the inference procedures using the censored median regression estimator originally proposed by Yang (1999). Numerical simulations are performed to validate the proposed procedures.
{n probability theory, one of the classical trends in probability theory is research related to the so-called Bernoulli sequences of random variables (r.v.). We are talking about independent random variables X1, X2,… taking value 1 with some probability p, 0