
In this article, we introduce a weighted version of the Xgamma exponential distribution, extending its utility in modeling lifetime data. We derive several important distributional properties of the proposed model, including moments, residual life functions, generating functions, stochastic ordering, aging intensity, and entropy. These properties provide deeper insights into the behavior and structure of the proposed distribution. To estimate the model parameters, we discuss the maximum likelihood estimation approach, focusing on complete sample data. To demonstrate the practical applicability of the proposed distribution, we analyze two real-world lifetime data sets. The performance of the weighted Xgamma exponential distribution is compared with several well-established one- and two-parameter lifetime distributions, along with their weighted versions. Additionally, comparisons are made with length-biased and area-biased lifetime distributions to further assess the robustness of the proposed model. The results of these comparisons indicate that the proposed weighted distribution offers a superior fit, particularly for data sets exhibiting an increasing failure rate. The model’s ability to outperform competing distributions highlights its potential as an effective alternative for analyzing lifetime data in reliability and survival studies.
The inaccuracy measure has recently become a valuable tool for detecting errors in experimental data. This measure applies only when random variables have density functions. To circumvent this constraint, the cumulative inaccuracy measure is a commonly used alternative measure of inaccuracy in the literature. When the observations generated by a stochastic process are recorded using a weight function, weighted distributions are established. Based on right-censored dependent data, we provide a nonparametric estimate for the weighted dynamic cumulative past inaccuracy measure in this study. The proposed estimator’s asymptotic characteristics have been examined, and its performance demonstrated through simulated and real-world data sets.
The game of cricket is enjoyed by millions of fans across the Globe. India is perusing the game like anything. Though, at present, many local tournaments like IPL were conducted in India, the ability of the batsmen and their patterns of scoring runswere reflected by the amount of runs they score in the international ODI’s played between Countries. The runs scored by batsmen with information on whether they are “batting-first” or “chasing” throws more light on the patterns of runs-scoring. Also, the order in which the batsmen bat is also taken into consideration in this study. Further, the performance of batsmen varies across different teams of the game. More uncertainty exists in the run-scoring pattern. All these make it difficult for predicting the runs scored by a batsman. Survival analysis comes handy in predicting the probabilities of such events. The study, in this perspective, considers nine batsmen selected from the one-day Indian world cup squad 2023. The information of these batsmen, particularly the runs scored by them against different countries were used to find the probabilities of their run-scoring pattern. The study uses Kaplan-Meier’s product limit estimator, Cox Proportional Hazard model and Accelerated Failure Time Parametric model for analysing the patterns. Log-rank test is used for comparing survival distributions. Also, the study compares the relative performance of selected batsmen. Data, updated as on 4th September 2023, were used for each of the batsman under consideration. The analysis has been carried out using R program.
This paper conducts a quantile-based income study of the Power-Pareto distribution. The major income inequality measures of the Power-Pareto distribution are derived, and the Lorenz ordering is studied. A simulation study is conducted to assess the performance of four competing estimation methods. The model is applied to a real income dataset, and both empirical and estimated income inequality measures are computed.
A cure rate model under a competing risk scenario where the number of competing causes follow a shifted binomial distribution with parameter p is proposed. Interestingly, the resulting distribution is the well-studied transmuted class of distribution. Few existing cure rate models are shown to be special cases of the proposed model. The identifiability issues of the model are studied in detail. Further properties of the model are investigated, and we discuss the maximum likelihood estimation of the parameter. The performance is confirmed through a simulation study using a defective Gompertz baseline and with competing causes. The Bayesian approach to the estimation of the parameter is adopted. The complexity of the likelihood function is handled through the Metropolis-Hastings algorithm. We analyse the data consisting of 8966 patients who have undergone bone marrow transplantation at the European Society for Blood and Marrow Transplantation (EBMT). The validation of the estimation algorithm is conformed using the bootstrap technique.
The advent of modern technology, permitting the measurement of thousands of variables simultaneously, has given rise to floods of data characterized by many large or even huge datasets. This new paradigm presents extraordinary challenges to data analysis and the question arises: how can conventional data analysis methods, devised for moderate or small datasets, cope with the complexities of modern data? The case of high-dimension low-sample size data is particularly revealing of some of the drawbacks. We look at the case where the number of variables measured in an object is at least the number of observed objects and conclude that (under the further assumptions that the data are observations from continuous random variables and that linear combinations of the variables are meaningful operations) this configuration leads to geometrical and mathematical oddities and is an insurmountable barrier for the direct application of traditional methodologies. If scientists are going to base their conclusions on high-dimension low-sample size data, ignoring fundamental mathematical results arrived at in this paper and blindly use software to analyze data, the results of their analyses may not be trustful, and the findings of their experiments may never be validated. That is why new methods together with the wise use of traditional approaches are essential to progress safely through the present reality.
We consider data modelling under one inflation for zero-truncated count data, as they typically arise in capture-recapture modelling. One-inflation in zero-truncated count data has recently found considerable attention. In this regard, zero-truncated New Discrete distribution and a distribution to a point mass at one are used to create a one-inflated model namely one-inflated zero-truncated New Discrete distribution. Its reliability characteristics, generating functions, and distributional properties are investigated in some detail. which includes survival function, hazard rate function, probability generating function, characteristic function, variance, skewness, and kurtosis. Monte Carlo Simulation have been undertaken to evaluate the effectiveness of the maximum likelihood estimators. To test the compatibility of our proposed model, the baseline model and the proposed model are distinguished by using the two different test procedures. The adaptability of the suggested model is demonstrated using two real-life datasets from separate domains by taking various performance measures into consideration.
Measures of income inequality are used for modelling and analysis of income data. In this paper, we present various income inequality measures in the quantile set up. We also introduce quantile version of well known dullness property. The interrelationships among these measures are investigated. The monotonic behaviour of income inequality measures are discussed. We also develop new quantile functions useful for income analysis. Various applications of the measures are discussed.
In this article, we propose non-parametric estimators for mean inactivity time function for complete and censored data. The asymptotic properties of the estimators are established using suitable regularity conditions. Monte Carlo simulation studies are used to study the efficiency of the estimators. Three real data sets are used to demonstrate the usefulness of the estimation procedure.
This additional material was presented by the author during the discussion meeting on the paper by Duembgen and Davies (2024), held on October 24, 2024, at the Department of Statistical Sciences, University of Bologna.
Shrinkage methods for estimating the parameters of a regression model with autoregressive integrated moving average (ARIMA) errors are presented when some regression parameters are restricted to a subspace. The estimates are obtained by maximizing the likelihood function with and without restrictions, yielding the unrestricted and restricted estimators, respectively. Shrinkage estimators optimally combine these two estimators. To demonstrate the optimality of these estimators, we use metrics such as asymptotic distributional bias (ADB) and asymptotic distributional risk (ADR), aiming to minimize both quantities. We show that the relative efficiency of the shrinkage estimator is superior to that of the unrestricted estimator when the shrinkage dimension exceeds two. Our large-sample theory and simulation study demonstrate that shrinkage estimators dominate the unrestricted estimator across the entire parameter space. An empirical example using Canadian crime rate data is also provided.
This paper introduces a new class of logistic distribution, namely Exponentiated logistic distribution, which is derived from type II logistic distribution. We have investigated its properties, discussed parameter estimation, and demonstrated its usefulness in analysing real-life medical data. The developed model provides researchers with valuable tools for accurately modelling and analysing medical phenomena, thereby contributing to advancements in healthcare research and decision-making.
In a regression setting with a response vector and given regressor vectors, a typical question is to what extent the response is related to these regressors, specifically, how well it can be approximated by a linear combination of the latter. Classical methods for this question are based on statistical models for the conditional distribution of the response, given the regressors. In the present paper it is shown that various p-values resulting from this model-based approach have also a purely data-analytic, model-free interpretation. This finding is derived in a rather general context. In addition, we introduce equivalence regions, a reinterpretation of confidence regions in the model-free context.
There are various fields where observations are taken on directions in three dimensions, e.g., sphere and torus. Herewe will introduce a very general family of distributions on sphere and torus by use of time series spectra, which includes a lot of proposed classical one as special cases. Because time series spectra can be described by a lot of famous parametric models, e.g., AR, ARMA etc., we can develop the systematic model selection in this field by use of AIC, BIC, etc. Applications are very wide.
Huang and Kotz (1984) proposed a two-parameter extension of the original Fralie-Gumble-Morgenstern (FGM) family to model the higher association between the random variables. In this problem, we develop an iterated FGM (IFGM) based dependent stress-strength reliability model using Lindley marginals. Some important statistical and reliability properties of the proposed distribution are also derived. The prime goal of this study is to investigate the effect of stress-strength reliability parameters with respect to the variation in the dependence parameters alpha and beta. Further, we compared the IFGM stress-strength reliability model with the original FGM using graphical representations to assess whether reliability was over or under-estimated. Finally, we investigated the performance of the proposed estimators through both Monte Carlo simulations as well as real data sets.
By using a sup-norm, sufficient conditions for the convergence of multivariate extremes and the potential limit types were fully identified by Barakat et al. (2020a). In this paper, we prove an intriguing result that by using the sup-norm, the weak convergence of multivariate extremes to the Fr & eacute;chet type implies the convergence of those multivariate extremes in an arbitrary D-norm to the same type-limit by using the same normalizing constants. As a result of this finding, the weak convergence to the Fr & eacute;chet type takes place by employing any logistic norm. Moreover, the two other possible limit types (max-Weibull and Gumbel types) are discussed. Similar findings are also demonstrated for multivariate record values. Finally, we demonstrate in a real-world scenario how to model multivariate extreme data sets utilizing the R-ordering principle and different norms.
In many cases involving hypothesis testing for parameters in multivariate Gaussian populations and certain other populations, likelihood ratio criteria, or their one-to-one functions, can be expressed in terms of the determinant of a real type-1 beta matrix. In geometrical probability problems, when the random points are type-1 beta distributed, the volume content of the parallellotope generated by these points is also associated with the determinant of a real type-1 beta matrix. These problems in the corresponding complex domain do not seem to have been discussed in the literature. It is well-known that the determinant of a real type-1 beta matrix can be written as a product of statistically independently distributed real scalar type-1 beta random variables. This paper addresses the general h-th moments of a scalar random variable, in either the real or complex domain, for any arbitrary h. The structure of these moments is quite general, and the paper provides exact distribution results, asymptotic gamma function results, and asymptotic normal results for both the real and complex domains.