
Robust inference for overdispersed count data is crucial in applications where outliers may substantially distort classical likelihood-based estimation of both the mean and dispersion. We develop robust estimation procedures for independent and identically distributed negative binomial data and provide practical guidelines on the choice of the estimator. We propose a combined robust M-estimator that jointly estimates the mean and dispersion parameter through an alternating updating scheme based on bounded score functions. The mean update relies on a bias-corrected Tukey-type M-estimator, extending the Poisson framework of Elsaied and Fried (2016) to the negative binomial setting, while the dispersion update follows a robust modification of score-based estimation in the spirit of Aeberhard et al. (2014). Under standard regularity conditions, we establish key theoretical guarantees for the proposed estimator, including consistency, local convergence of the alternating algorithm, asymptotic normality, and robustness to outliers through bounded influence. Extensive simulation studies compare the proposed method with maximum likelihood estimation, minimum disparity estimators, and weighted maximum likelihood approaches under both clean and contaminated sampling. The results show that classical likelihood procedures can be highly sensitive to additive contamination, particularly in dispersion estimation, whereas the proposed estimator achieves a favorable efficiency–robustness trade-off across a broad range of sample sizes and contamination regimes, while retaining high efficiency under the nominal model. An analysis of epileptic seizure count data further illustrates the practical relevance of the approach and the stability of the resulting inference without relying on ad hoc data cleaning.
Compositional data, which capture relative contributions of parts to a whole, are constrained by a constant sum. They arise in various fields such as geochemistry, soil science, and paleoecology, describing data typically affected by measurement errors. We discuss the kernel method for contaminated compositional data, using a deconvolution approach. Some numerical experiments are provided using both simulated and real data.
Simulation studies are often carried out to study theoretical properties of a statistical procedure that cannot be derived analytically. Oftentimes, simulation studies can be very extensive, including various factors with various levels. This makes the results of such studies often difficult to report, and difficult for readers to interpret the general picture in the overabundance of results. Additionally, in many cases such extensive studies may not even be necessary. This work aims to explore and discuss some general guidelines on when (not) to include a simulation factor. Depending on the research question, oftentimes varying only one or two factors may suffice for answering the research question. In the current paper two different types of simulation studies are defined, namely proof of principle studies and robustness studies. Next, for a number of quality measures it is argued which factors are and are not useful to vary, both in proof of principle studies and robustness studies. Finally, the results of a literature study are presented confirming the need for guidelines on what factors to vary in a simulation study. A general recommendation is to not vary too many factors, especially for proof of principle studies.
Most parametric statistical analyses focus on estimating or testing parameters within a known distribution, but there is growing interest in predictive inference for practical decision-making, such as setting insurance premiums or shaping public policy. This paper reviews three existing non-Bayesian methods for obtaining predictive densities. For models with a scalar parameter, we investigate predictive densities obtained from significance functions and derive explicit predictive distributions for several commonly used models. In the case of linear regression models, the proposed method can be directly applied because the significance functions of the scale parameter and the location parameters are explicitly available from the marginal analysis and conditional analysis, respectively. For more general models where a significance function is not defined for vector parameters, an approximate predictive density is constructed using the significance functions of individual parameters, which are derived from the asymptotic distribution of the maximum likelihood estimator. The methodology is demonstrated with real-world examples and evaluated through simulation studies against existing methods.
This paper presents a scalable and supervised Hidden Markov Model (HMM) for predicting the operational states of a tractor based on mixed sensor measurements. Accurately predicting the operational states of a tractor is essential for optimizing cost monitoring, resource utilization, and decision-making in modern agriculture. The data used in this study consists of various sensor measurements and environmental factors collected from a tractor. The dataset is compressed by summing the measurements corresponding to the same operational state, reducing its complexity while retaining essential information. The proposed supervised HMM specifies a Markovian state process and state-dependent emission distributions; parameters are estimated from labeled training sequences and prediction is performed on unlabeled test sequences. The performance of the scalable supervised HMM is compared with other machine learning algorithms commonly used in agriculture. This study aims to develop a predictive model that captures the relationship between the operational state and the mixed sensor/environmental variables, considering the complexity of the data. The proposed methodology facilitates a deeper understanding of tractor operational dynamics and supports data-driven optimization of farm processes.
This study investigates attitudes towards migrants by analysing responses collected through three measurement instruments: an ad hoc Stereotyped Beliefs about Immigrants Scale (SBIS), the Osgood semantic differential and the Bogardus social distance scale. Item Response Theory (IRT) models are applied to the structured instruments, while Structural Topic Modelling (STM) is used to analyse open-ended responses. This dual-method approach integrates quantitative and qualitative data, offering a comprehensive examination of prejudice towards marginalised groups across both structured scales and spontaneous narratives. IRT is employed to identify and compare latent traits underlying attitudes across the three instruments, providing insight into the strength and coherence of these beliefs. Additionally, STM is used to uncover dominant topics in open-ended responses and assess their association with varying levels of prejudicial disposition. The findings contribute to a deeper understanding of the multidimensional nature of attitudes towards migrants, with implications for addressing bias and prejudice in public discourse.
We propose a flexible and interpretable framework for circular data, motivated by wind-direction regimes. The model is a covariate-dependent finite mixture of Kato–Jones (KJ) components, reparameterized so that each component is a convex combination of a wrapped–Cauchy core and a uniform background (a WC–uniform representation). This boundary–concentration form separates regime shape from diffuse transitional mass and, at the mixture level, induces a wrapped–Cauchy mixture with an explicit uniform background. Covariates such as hour of day and wind speed enter only through a multinomial-logit link on the mixture weights, so regime incidence varies with conditions while component locations and concentrations remain globally interpretable. On the theoretical side, we establish identifiability of the covariate-dependent WC–uniform mixture via a trigonometric-moment (Toeplitz/Carathéodory–Fejér) argument under mild ordering and compactness conditions, and we derive a likelihood-ratio test for the presence of a uniform background with a chi-bar-square limit, calibrated in practice by a parametric bootstrap. For estimation we develop a block-separable sieve EM algorithm that alternates a strictly concave multinomial-logit update for the weights with closed-form moment updates for the wrapped–Cauchy parameters, guaranteeing monotone ascent of the observed likelihood and convergence to stationary points. Simulation studies with and without covariates demonstrate accurate recovery of skewed, overlapping, and switching regimes, and highlight the robustness of the WC–uniform representation to diffuse noise. Applications to NOAA buoy wind directions show that the proposed framework provides interpretable regime and uniform-background diagnostics in multi-regime, transition-rich coastal settings, while also identifying cases in which simpler von Mises mixtures or kernel-based methods provide stronger predictive likelihood.
A major concern with the use of non-probability samples is their lack of representativeness that, if not accounted for properly, may lead to large bias in survey estimates. Non-probability samples involve subjective methods for sample selection, so that inclusion probabilities are unknown and it is not possible to apply the traditional randomization theory for inference on the population parameters. In this paper, the uncertainty in survey estimates resulting from the non-identifiability of the sampling design acting in the non-probability sample, as well as its reduction due to availability of extra-sample information, is discussed. Next, the effect of non-identifiability on survey estimates accuracy is evaluated. Finally, an application to real enterprise data from Italy is performed.
We introduce and study the logratio Student’s t distribution as a robust and flexible alternative to the classical logratio normal model for compositional data. This distribution preserves the geometric coherence of the Aitchison simplex framework while accommodating heavier tails, enabling improved detection of outliers — an essential feature in many real-world applications. Our main contribution is to formally embed the t distribution within the logratio framework and to demonstrate its equivalence across logratio representations for the purposes of estimation and outlier detection, while clarifying the distinct mathematical properties of each representation beyond distribution fitting. Monte Carlo simulations under scale contamination show that our model, combined with a Leave-One-Out procedure, outperforms traditional robust methods (MCD, COMCoDa) by ensuring higher AUC and near-perfect specificity, even in high dimensions. A real-world application illustrates the model’s superior ability to uncover structure and identify outliers in multivariate compositional data. These results position the logratio t distribution as a theoretically sound and practically powerful tool for robust inference in the simplex.
In this paper, the mathematical properties of monotonicity and distribution discriminatory, which have been shown to be requisite properties when comparing transition probability matrices in credit risk modelling, are proved to hold for the risk-adjusted difference indices. This is a key result, and it implies that the risk-adjusted difference indices are suitable for comparing transition probability matrices; as they will be able to capture key properties that are peculiar to transition probability matrices in credit risk modelling. To the best of the researchers' knowledge, for the first time an index that has been proved to satisfy the requisite properties is used to compare the performance of the embeddability problem solving algorithms. The results of the non-parametric Kruskal-Wallis test and the Wilcoxon rank sum test show that the differences in the performances of the algorithms are statistically significant.
This article addresses the problem of nonparametric deconvolution of a univariate cumulative distribution function with respect to the Wasserstein-Kantorovich distance. We consider a deconvolution model in which the observations consist of a signal contaminated by an independent measurement error. The error distribution is assumed to be known and ordinary smooth. The signal sequence is assumed to be strictly stationary and either ρ -mixing or φ -mixing (both implying α -mixing), a framework that encompasses many Markov models and other dependent data structures. For instance, classical ARMA processes are α -mixing (or strongly mixing) with coefficients decaying to zero at an exponential rate. We establish non-asymptotic upper bounds on the Wasserstein risk for an approximate minimum L^1 -distance kernel-based estimator of the signal marginal cumulative distribution function, under the assumption that the signal has finite first absolute moment. Two distinct scenarios are analyzed: first, when no additional regularity is imposed on the signal distribution, and second, when the signal distribution has a Lebesgue density in a Sobolev-type class. In both settings, we provide matching lower bounds for the minimax risk, thereby establishing the minimax optimality of the derived convergence rates. Notably, these rates coincide with those recently established for the i.i.d. setting, indicating that weak dependence does not increase the minimax complexity of the problem. Our results complement those for the i.i.d. case and represent a first step toward establishing minimax-optimal convergence rates for cumulative distribution function deconvolution under more general dependent signal processes.
The statistical modeling of consensus in Delphi data faces a critical bottleneck: the high dimensionality of questionnaire items relative to the limited sample size of expert panels. This rank deficiency leads traditional latent variable models, such as Principal Component Analysis, to be structurally unstable and prone to overfitting. Addressing this methodological gap, this study proposes a transition from variable-centric covariance models to network-centric connectivity models. By mapping item correlations onto a weighted graph topology, we present a simulation-based benchmark that utilizes community detection algorithms to identify latent thematic structures, effectively addressing the spectral instability and rank deficiency typical of high-dimensional, low-sample-size regimes. The research systematically evaluates the robustness of topological approaches based on structural density, information flow, and spectral partitioning against synthetic datasets designed to replicate the pathological conditions of consensus data, including ordinal scales and systemic noise. The central methodological contribution lies in demonstrating that collinearity among expert judgments - traditionally treated as statistical redundancy to be regularized - can be effectively reinterpreted as a topological signal of cohesion. This framework provides researchers with a structured and automated procedure for dimensionality reduction, ensuring structural stability and psychometric consistency even in small-sample regimes where standard factor analysis breaks down.
We investigate whether the impact of latent market expectations, models and current holdings (collectively termed market positioning) has an impact on the reaction of U.S. Treasury bond markets to macroeconomic surprise. Our initial findings indicate a positive linear relationship between market reaction and the surprise content of the announcement. A positive linear relationship is also found between the surprise content and the market reaction post announcement. However, we demonstrate that the transitivity of these relationships fails to hold, in that no significant linear relationship could be found between pre-announcement reaction and post-announcement reaction. Thus, we hypothesise that latent market positioning is distorting the linearity between pre and post reaction, and that the market response can be regime dependent. Motivated by these findings, we introduce a threshold auto-regressive model to study the reactions to price movements in the lead-up to macroeconomic announcements. The model leverages multiple regimes which are determined by combinations of the observed macroeconomic surprise and by the latent market positioning, therefore capturing the non-linearity of the process. One critical novelty of our work is that the market positioning is latent and thus inferred from the data. We adopt a combination of particle swarm optimisation and genetic algorithms to estimate the model and to perform model selection: we discover three market positioning states to be optimal. These states are highly interpretable and the transition probabilities between states are determined based on the surprise of previous events. As an illustrative example, we use these findings to sketch a general playbook for how to trade these events, thus providing asymmetric risk reward opportunities.
The logistic-normal multinomial distribution has been used for modelling microbiome data obtained from high-throughput sequencing technologies, which are compositional in nature. A logistic-normal multinomial distribution is a hierarchical multinomial distribution that assumes the latent variable which are the additive log-ratio (ALR) transformed proportions in a multinomial distribution follows a Gaussian distribution. Model-based clustering algorithms have also been developed for clustering microbiome data based on the logistic-normal models. However, the Gaussian assumption may violated when the ALR transformed variable exhibit heavy-tailed distributions or has outliers. Our study introduces a novel mixture of logistic-t multinomial models that effectively address these challenges. Utilizing a multivariate t distribution for ALR transformed latent variables, the proposed approach provides the flexibility to capture heavy tails and thus better accommodates the outliers when clustering microbiome compositional data. We incorporate a variational Expectation-Maximization (EM) algorithm to facilitate efficient parameter estimation for intractable posterior distributions. Our model demonstrates competitive performance in terms of clustering accuracy and parameter recovery as compared to existing approaches through simulation studies and real data analysis, hence offering a robust tool for exploring the complex structures and heterogeneity of microbiome compositional data.
The paper describes statistical methods for the regression analysis of spatial data when some spatial information is missing. Our numerical example comes from the results of the 1988 carbon dating of the Shroud of Turin (TS). The statistical problem is that the dating was performed at three laboratories, which each received a sample from the Shroud. The physical location on the TS of these samples is known. However, the data consist of readings on several subsamples, the locations of which within the samples are not known. This is then a problem of spatial regression with partially labelled regressors. In 2019 the laboratory data (the ‘raw’ data) from which the original dating was derived became available. We analyse the raw data by considering all 165,888 permutations of observations to subsamples. We first use modern graphical methods to interpret the output from these permutations, with a focus on the age of the TS. Further exploration of the data requires the selection of a single fitted model. We select a model by reference to the distribution of R^2 , the selected model being used, via robust regression, to screen the data for outliers. Simpler versions of this problem occur in commercial ovens in food and antibiotic preparation, when records on shelf location are not available.
This paper introduces a novel heavy-tailed and alternative t type distribution within a bounded interval, arising as a scale mixture of the logit-normal distribution. Characterized by its prominent heavy-tails, excess kurtosis, and bounded characteristics, the proposed model offers a versatile and fitting solution for various computer vision and pattern recognition challenges. The ECME algorithm is considered for the precise estimation of model parameters and its finite mixtures. To demonstrate the efficacy of our approach, we conduct experiments using authentic natural and medical images. Our numerical findings underscore the enhanced robustness and accuracy of the proposed model in image segmentation when compared to traditional mixture models.
The interrater agreement of ratings on a nominal scale regarding a group of targets (individuals, objects, etc.) is usually computed by kappa-type indices. Recently, a homogeneity index for a qualitative variable was considered to define a single target measure of agreement between raters for each target, and to propose a global measure of agreement for the whole group of targets allowing to overcome limitations affecting traditional kappa-type indices. In this paper, the sampling properties and the asymptotic distribution of the proposed index are investigated. Finally, a simulation study is performed to demonstrate the accuracy of the index and an application to real educational data is provided.
This study proposes a methodological approach to investigate gender disparities in education, particularly focusing on the schooling phase, which has a crucial influence on career trajectories. The research employs multilevel linear models to analyze student performance concerning various factors, with a particular emphasis on gender-specific outcomes. The study aims to identify and test context-specific independencies that may reflect educational disparities between genders. The methodology includes the introduction of supplementary parameters in multilevel models to capture and examine these independencies. Furthermore, the research proposes encoding these novel relationships in graphical models, specifically stratified chain graph models, to visualize and generalize the complex dependencies among covariates, random effects, and gender influences on educational outcomes.
Clustering long term follow-up data has broad applications in supervised and unsupervised learning. This study proposes to accumulate the dissimilarity measure across the study interval to provide an overall index for clustering. The data, typically non-Gaussian, are assumed to be collected at regular time points, and dissimilarity is calculated as a rank-based noncentrality. To initiate the clustering, the K-means and a discretized K-center method of functional principal component analysis is used. We propose a fast computational algorithm to speed up the clustering process with tolerable sacrifice in correct rates. The proposed algorithm is illustrated through simulation studies. For real-world data analysis, a year of fine particulate matter (PM _2.5 ) monitoring data are analyzed for exploring the spatial closeness between and within clusters.
Understanding the nature of the data, dealing with outliers and redundant information are key issues when designing a proper metric for clustering and classification. Distance-based generalized linear models are prediction tools which can be applied to any kind of data whenever a distance measure can be computed among units. In this work, robust ad-hoc metrics are proposed to be used in the predictors’ space of these models, incorporating more flexibility to this tool. Their performance is evaluated by means of an extensive simulation study and compared to those based on Gower’s and the Euclidean distances through several data sets of multivariate heterogeneous data with the presence of anomalous observations. Accuracy, precision, recall, F1 score, Auc Roc and Log Loss measures are used to evaluate the effectiveness in the prediction of responses. Applications on real data are provided in order to illustrate the predictive power of these models, which appear to be competitive with state-of-the-art machine learning approaches, such as random forests or neural networks. Computations are made using the dbstats package for R.