
The Gini index is an important measure of income inequality, and calculating it under different conditions is highly significant. Left truncated and right censored data are common in economics, while existing literature on Gini index estimation often focuses on right censoring. We construct a nonparametric estimator for the Gini index when data are subject to random left truncated and right censored, and investigate its consistency and asymptotic normality. Simulations and a real data example demonstrate that it has certain advantages when compared to existing methods.
To achieve greater flexibility for modelling heavy-tailed responses on the unit interval, a beta scale mixture (BSM) regression model is proposed. The conditional response is assigned a mean-parameterised beta distribution whose variability parameter is scaled by a mixing random variable, taking values on all or part of the positive real line, whose distribution depends on parameters governing the tail behaviour of the resulting compound distribution. The conditional mean of the response variable is linked to covariates through a logit link, while the mixing mechanism offers greater flexibility towards the skewness and kurtosis than classical beta regression. To validate the effectiveness of the proposed regression model, we conduct a simulation study and illustrate its practical relevance through applications to real datasets. The results indicate that the BSM regression model consistently outperforms the classical beta and alternative competing regression models for responses on the bounded unit domain. The methodology is implemented in the open-source BSMreg package for R, available at https://github.com/arnootto/BSM .
Incomplete data, such as missing or censored values, frequently arise in matrix-variate datasets. At the same time, latent heterogeneity–manifested as subpopulations or clusters–is often present. To our knowledge, no finite mixture model has yet been proposed that can simultaneously handle matrix-variate data with both missing and censored values. To address this gap, we introduce a novel model based on the matrix-variate normal distribution that accommodates missingness and various censoring types while capturing latent group structure through clustering. The missing/censoring mechanism is assumed to be missing at random, implying a non-informative (ignorable) mechanism. An efficient expectation-conditional maximization algorithm is developed for maximum likelihood estimation. Simulations demonstrate accurate parameter recovery under a variety of censoring and missingness patterns. Applications of the method to environmental data from the Chesapeake Bay and clinical data from the Age-Related Eye Disease Study illustrate its ability to reveal ecologically and clinically meaningful structures in complex, incomplete datasets.
This paper describes an interesting class of discrete probability distributions that arise from the discretization of continuous distributions. They are sometimes also called telescopic distributions, a term inspired by the concept of “telescoping series" in mathematics. After some introductory remarks, we will focus primarily on deriving two-dimensional telescopic distributions, i.e., discretized two-dimensional continuous probability distributions, which are typically constructed using a suitable copula. We will then present examples using both simulated and real data to demonstrate the advantages and potential applications of these models. Among other things, we will derive and apply the “telescopic" form of the bivariate negative binomial distribution.
Statistical procedures are proposed in overdispersed generalized partial nonlinear models (OGPNLM). An adaptation of the iteratively re-weighted least squares procedure based on the backfitting algorithm is derived for the parameter estimation, under (penalized) maximum likelihood estimation. Derivations of effective degrees of freedom, standard errors of the estimates and some hypothesis tests are given. Diagnostic methods based on known techniques, such as leverage measures, residual analysis and local influence graphics, under different perturbation schemes, are proposed. Finally, in order to apply the proposed techniques to real-world studies, two motivating datasets are analyzed by the method developed in the paper.
Cancer survival analysis in population-based settings relies on the relative survival framework, which enables the estimation of survival specific to the disease of interest in the absence of information on the individual cause of death. Within this framework, the overall hazard is decomposed into the excess hazard due to cancer and hazard due to the other competing causes of death, the latter derived from routine sources such as national or regional life tables. The survival function associated only with the excess hazard corresponds to the estimand of interest, known as net survival. Non-parametric estimators of this quantity, including Ederer I, Ederer II, Hakulinen, and Pohar-Perme, have been widely used in cancer epidemiology, yet their application has been the subject of long-standing methodological debates regarding bias, variability, and consistency. In this review, we introduce the relative survival framework and outline the mathematical foundations of the non-parametric estimators used in this field, highlighting their key statistical properties, historical development, and continued relevance in population-based cancer research.
High rates of cardiovascular disease are prevalent among individuals with diabetes. Weight variability has been identified as a distinct risk factor for cardiovascular disease (CVD) in the general population, and there is preliminary evidence suggesting its significance extends to those with type 1 diabetes as well, regardless of gender. Utilising data from the Swedish National Diabetes Register, we investigate the association between visit-to-visit body weight fluctuations and the risk of cardiovascular diseases in males and females with type 1 diabetes who had no pre-existing cardiovascular conditions at the start of the study. To assess the relationship between cardiovascular risk, weight variability, and gender, we propose a generalisation of weighted log-ratio analysis for a three-way contingency table. Weighted log-ratio analysis offers several advantages over other categorical data analysis techniques. These include the computation and visual representation of the odds ratios (ORs) upon which the analysis relies. Interestingly, the association between CVD and weight variability differs between males and females. The three-way log-ratio analysis method presented in this paper gives a visual description of these conditional ORs in terms of point distances on a biplot, illustrating the association between types of CVD, weight variability and gender.
Parameter estimation in constrained spaces is a widely discussed topic. A common Bayesian approach incorporates boundary information by employing a uniform prior over the constrained parameter space. However, subsequent inference typically relies on the squared error loss function, which is often inadequate because it fails to sufficiently penalize estimates near the boundaries. This paper proposes an interval entropy loss function for parameters constrained to an interval (a, b). This loss function effectively penalizes estimates that lie near the boundaries, analogous to how the squared error loss behaves on the entire real line. Bayesian estimation of the generalized extreme value distribution shape parameter, constrained to the interval (-1,1) is considered under this interval entropy loss function. Through Monte Carlo simulations, the advantages of the interval entropy loss function are demonstrated by comparing its performance with that of two other interval-defined loss functions. Finally, the practical application of this approach is shown by applying it to the computation of return levels for extreme precipitation data.
Income is a key variable in many surveys. Nonresponse rates are increasing in surveys conducted by national statistical institutes and income is no exception. Ignoring nonresponse can lead to two major problems, namely an increase in variability and bias in the estimators. Dealing with nonresponse regarding income is particularly challenging because the likelihood of responding depends on income (non-ignorable nonresponse). In addition, the empirical distribution of income may be complicated, with multiple bumps, asymmetry, and long tails. We propose an income mean estimator that can be used when income can be modeled using a mixture distribution. Our approach is inspired by the distribution of income collected from the Swiss Statistics on Income and Living Conditions (SLIC) survey data of 2009 (Federal Statistical Office, Survey on Income and Living Conditions in Switzerland. Data set of the 2009 survey, 2009). The proposed procedure combines two ingredients: a model-based estimation and a nonresponse weight adjustment. The values of the income (the variable of interest) are predicted on the basis of a fitted mixture distribution on the respondents’ set. The response probabilities are estimated using generalized calibration with the collected income as a response model variable. The inverse estimated response probabilities provide a nonresponse weight adjustment that accounts for nonignorable nonresponse. We incorporate both the predicted values of the variable of interest and the estimated response probabilities into a Hájek-type mean estimator, and provide an estimator of its variance by parametric bootstrap. We conduct a simulation study to check the behavior of the proposed estimator and its variance estimator. Finally, we show that the proposed estimator outperforms some competitors used to estimate the mean population income for the SILC data.
This paper formalizes and extends Goodman’s (J Am Stat Assoc 70(352):755–768, 1975) all-nothing Latent Class Analysis (LCA) to accommodate data with substantial proportions of respondents that endorse either all or none of the binary indicator variables. A simulation study compares all-nothing LCA to unrestricted LCA in terms of recovering the true number of latent classes under varying sample distributions. We apply both LCA model specifications to two previously published empirical datasets having all-zero and all-one response patterns: cross-sectional study-skill indicators and binary indicators for post-retirement employment status at multiple measurement occasions. The longitudinal application of all-nothing LCA is a latent trajectory analysis that mimics the mover–stayer latent Markov model. This paper illustrates for which data patterns all-nothing constraints improve class recovery and interpretation. We conclude with recommendations and future research directions for applying all-nothing LCA in combination with unrestricted LCA.
In parametric statistical decision theory, the performance of a decision function is evaluated in terms of its risk function, a quantity that depends on the parameter of the model. The risk function is typically summarized by the Bayes risk, its expected value with respect to a prior distribution assigned to the parameter. However, the study of the whole distribution of the risk function and the derivation of additional summaries (such as quantiles and the mode) may provide a better insight into the overall risk associated to the decision. In this article we consider models that are scale-exponential families. In the case of point or interval estimation we provide general closed-form expressions for cumulative distribution and density functions of the random risk along with some relevant summaries. We show how this unifying approach can be specialized to specific relevant models and also employed for the definition of sample size determination criteria.
This article proposes an extension of the Spatial Moving Average model from a normal error specification to a skew normal error specification. The probability distribution of the observations induced by the skew normal error specification has been derived. The expression for the characteristic function has been obtained together with some lower order moments. Important characteristics of the proposed model are its homoscedasticity and interpretable correlation structure. The model is implemented on a real dataset through a full-likelihood function using the Differential Evolution algorithm. Confidence intervals for the model parameters are constructed using the bootstrap method. The obtained results are compared with some existing models.
We introduce a Bayesian quantile mixed-effects model for censored longitudinal outcomes based on the skew exponential power (SEP) error distribution. The SEP family separates tail behavior and skewness from the targeted quantile and includes the skew Laplace (SL) distribution as a special case. We derive analytic likelihood contributions for left, right, and interval censoring under the SEP model, so censored observations are handled within a single parametric framework without numerical integration in the likelihood. In simulation studies with varying censoring patterns and skewness profiles, the SEP-based quantile mixed-effects model maintains near-nominal bias and credible interval coverage for regression coefficients. In contrast, the SL-based model can exhibit bias and undercoverage when the data’s skewness conflicts with the skewness implied by the target quantile. In an HIV-1 RNA viral load case study with left censoring at the assay limit, bridge-sampled marginal likelihoods and simulation-based residual diagnostics favor the SEP specification across quantiles and yield more stable estimates of quantile-specific viral load trajectories than the SL benchmark.
We consider a bivariate, possibly non-homogeneous, finite-state Markov chain (X,U)={(X_t,U_t)}_t=1^n . We are interested in the marginal process X, which typically is not a Markov chain. The goal is to find a realization (path) x=(x_1,… ,x_n) with maximal probability P(X=x) . If X is Markov chain, then such path can be efficiently found using the celebrated Viterbi algorithm. However, when X is not Markovian, identifying the most probable path—hereafter referred to as the Viterbi path—becomes computationally expensive. In this paper, we explore the branch-and-bound method for finding Viterbi paths. The method is based on the lower and upper bounds on maximum probability max _x P(X=x) , and the objective of the paper is to exploit the joint Markov property of (X, Y) to calculate possibly good bounds in possibly cheap way. This research is motivated by decoding or segmentation problem in triplet Markov models. A triplet Markov model is trivariate homogeneous Markov process (X, Y, U). In decoding, a realization of one marginal process Y is observed (representing the data), while X and U are latent processes. The process U serves as a nuisance variable, whereas X is the process of primary interest. Decoding means estimating the hidden realization of X based solely on the observation Y. Conditional on Y, the latent processes (X, U) form a non-homogeneous Markov chain. In this context, the Viterbi path corresponds to the maximum a posteriori (MAP) estimate of X, making it a natural choice for signal reconstruction.
Multiple methods were developed to estimate saved time to better interpret benefits of drugs in trials for Alzheimer’s disease. These methods include the backward projection methods by using the placebo disease progression curve estimated with the linear interpolation method or the spline method that connects the estimated mean values between the outcome changes over time. In addition to these methods, the disease progression scale based on both progression curves and a cubic regression model may be used to estimate saved time and its confidence interval. In this article, we compared the performance of these interval methods with regard to coverage probability and interval width under various scenarios to mimic the treatment effect in real trials. The backward projection method often has the coverage probability above the nominal level when the placebo disease progression curve is relatively smooth and the treatment effect size is large, at the cost of wider intervals. The intervals based on the disease progression scale have the coverage probability very close to the nominal level with a shorter interval width than other methods in many cases. Data from the phase 2 donanemab trial were used to illustrate the application of these interval methods.
Cluster sampling is a widely used sampling design for survey data selection. The collected data provide information on both the individual units and the clusters. The parameter of interest is an unknown population total at the unit level. Due to the cluster sampling design, the model-assisted generalized regression estimator can be established either at the unit or at the cluster level. The unit-level estimator, however, ignores the cluster sampling design. The cluster-level estimator relies solely on the per-cluster aggregated information, which leads to the loss of individual patterns and the risk of ecological fallacy. As a remedy, in this paper we propose a hybrid generalized regression estimator as a new approach that balances between unit- and cluster-level modeling. It is implemented at the unit level like the study variable. However, in addition to the information of the units, it also takes into account the information of the other cluster members.
Capture-recapture methods for estimating the total size of elusive populations are widely-used, however, due to the choice of estimator impacting upon the results and conclusions made, the question of performance of each estimator is raised. Motivated by an application of the estimators which allow covariate information to meta-analytic data focused on the prevalence rate of completed suicide after bariatric surgery, where studies with no completed suicides did not occur, this paper explores the performance of the estimators through use of a simulation study. The simulation study addresses the performance of the Horvitz–Thompson, generalised Chao and generalised Zelterman estimators, and develops a novel, generalised, form of the modified Chao estimator to account for both covariate information and one-inflation. In addition, the performance of the analytical approach to variance computation is addressed. Given that the estimators vary in their dependence on distributional assumptions, additional simulations are utilised to address the question of the impact outliers have on performance and inference.
This paper considers the inference of unknown parameters for an extension of the new Pareto-type distribution based on progressive type-II censored data. First, the estimation of the model parameter using maximum likelihood and Bayesian methods has been discussed. The approximated confidence intervals and Bayesian credible intervals are discussed as well. We then establish a Bayesian optimal design with respect to variance minimization criteria. Monte Carlo simulations are implemented to compare different methods of estimation, and finally, two real data sets, where the first one represents the remission times (in months) of bladder cancer patients and the second one represents the repair times (in hours) for an airborne communication transceiver have been analyzed for illustrative purposes.
We demonstrate an alternative derivation of the distributional properties of the maximum likelihood estimators for the parameters in an inverse Gaussian distribution, requiring only the univariate transformation, the moment generating function technique, and the Basu’s theorem. This framework simplifies existing methods and extends to other parametric distributions, such as the negative exponential distribution.