Traditional statistical methods need to be updated to work with modern distributed data storage paradigms. A common approach is the split-and-conquer framework, which involves learning models on local machines and averaging their parameter estimates. However, this does not work for the important problem of learning finite mixture models, because subpopulation indices on each local machine may be arbitrarily permuted (the "label switching problem"). Zhang and Chen (2022) proposed Mixture Reduction (MR) to address this issue, but MR remains vulnerable to Byzantine failure, whereby a fraction of local machines may transmit arbitrarily erroneous information. This paper introduces Distance Filtered Mixture Reduction (DFMR), a Byzantine tolerant adaptation of MR that is both computationally efficient and statistically sound. DFMR leverages the densities of local estimates to construct a robust filtering mechanism. By analysing the pairwise L2 distances between local estimates, DFMR identifies and removes severely corrupted local estimates while retaining the majority of uncorrupted ones. We provide theoretical justification for DFMR, proving its optimal convergence rate and asymptotic equivalence to the global maximum likelihood estimate under standard assumptions. Numerical experiments on simulated and real-world data validate the effectiveness of DFMR in achieving robust and accurate aggregation in the presence of Byzantine failure.
Finite mixtures of beta distributions provide a flexible framework for modelling bounded data, such as proportions or rates, and are widely used in applications in genomics and biostatistics. However, estimation is complicated by challenges such as explosive likelihood and non-identifiability. We develop new methodologies to address these issues. We propose a penalized maximum likelihood estimator that stabilizes estimation, along with a score-adjusted moment-based alternative that improves computational efficiency and serves as an effective initialization strategy for the EM algorithm. For order selection under non-identifiable models, we propose a beta mixture reduction method based on composite transportation divergence, which provides an efficient computational tool for generating lower-order mixtures in order selection workflows. The proposed methods are evaluated through simulation studies and applied to DNA methylation data. All implementations are available in the MixtureInf2.0 R package on GitHub.
We establish the validity of bootstrap methods for empirical likelihood (EL) inference under the density ratio model (DRM). In particular, we prove that the bootstrap maximum EL estimators share the same limiting distribution as their population counterparts, both at the parameter level and for distribution functionals. Our results extend existing pointwise convergence theory to weak convergence of processes, which in turn justifies bootstrap inference for quantiles and dominance indices within the DRM framework. These theoretical guarantees close an important gap in the literature, providing rigorous foundations for resampling-based confidence intervals and hypothesis tests. Simulation studies further demonstrate the accuracy and practical value of the proposed approach.
Gaussian mixtures are widely used for approximating density functions in various applications such as density estimation, belief propagation, and Bayesian filtering. These applications often utilize Gaussian mixtures as initial approximations that are updated recursively. A key challenge in these recursive processes stems from the exponential increase in the mixture's order, resulting in intractable inference. To overcome the difficulty, the Gaussian mixture reduction (GMR), which approximates a high order Gaussian mixture by one with a lower order, can be used. Although existing clustering-based methods are known for their satisfactory performance and computational efficiency, their convergence properties and optimal targets remain unknown. In this paper, we propose a novel optimization-based GMR method based on composite transportation divergence (CTD). We develop a majorization-minimization algorithm for computing the reduced mixture and establish its theoretical convergence under general conditions. Furthermore, we demonstrate that many existing clustering-based methods are special cases of ours, effectively bridging the gap between optimization-based and clustering-based techniques. Our unified framework empowers users to select the most appropriate cost function in CTD to achieve superior performance in their specific applications. Through extensive empirical experiments, we demonstrate the efficiency and effectiveness of our proposed method, showcasing its potential in various domains.
We study the first-order stochastic dominance (SD) test in the context of two independent random samples. We introduce several test statistics that effectively capture violations of the dominance relationship, particularly in the tail regions. Additionally, we develop a resampling procedure to compute the p-values or critical values for these tests. The proposed tests have asymptotic type I error rates for frontal configurations equal to the nominal level alpha. Furthermore, their powers approach 1 for any fixed alternatives. Through simulation experiments, we demonstrate that our SD tests outperform the recentring test proposed by Donald and Hsu (2016) as well as the integral-type test presented by Linton et al. (2010) in various scenarios discussed in existing literature. We also employ the proposed tests to analyze changes in the distribution of household income in the United Kingdom over time. The proposed tests offer some insights into potential dominance relationships within this context. Ce travail porte sur le test de dominance stochastique (SD) du premier ordre dans le cas de deux & eacute;chantillons al & eacute;atoires ind & eacute;pendants. Ses auteurs introduisent plusieurs statistiques de test capables de d & eacute;tecter efficacement les violations de la relation de dominance, en particulier dans les queues de distribution. Ils d & eacute;veloppent & eacute;galement une proc & eacute;dure de r & eacute;& eacute;chantillonnage permettant de calculer les p-valeurs ou valeurs critiques associ & eacute;es & agrave; ces tests. Asymptotiquement, les tests propos & eacute;s contr & ocirc;lent le risque de premi & egrave;re esp & egrave;ce au niveau nominal alpha dans les configurations frontales. De plus, leur puissance tend vers 1 pour toute alternative fix & eacute;e. Des simulations montrent que ces tests SD surpassent celui de recentrage de Donald & Hsu (2016) ainsi que le test int & eacute;gral de Linton et al. (2010), dans divers sc & eacute;narios & eacute;tudi & eacute;s dans la litt & eacute;rature. Enfin, les auteurs analysent l'& eacute;volution de la distribution des revenus des m & eacute;nages au Royaume-Uni & agrave; l'aide de leurs tests, offrant un & eacute;clairage sur les potentielles relations de dominance en jeu.
Mixture of regression model is widely used to cluster subjects from a suspected heterogeneous population due to differential relationships between response and covariates over unobserved subpopulations. In such applications, statistical evidence pertaining to the significance of a hypothesis is important yet missing to substantiate the findings. In this case, one may wish to test hypotheses regarding the effect of a covariate such as its overall significance. If confirmed, a further test of whether its effects are different in different subpopulations might be performed. This paper is motivated by the analysis of Chiroptera dataset, in which, we are interested in knowing how forearm length development of bat species is influenced by precipitation within their habitats and living regions using finite Gaussian mixture regression (GMR) model. Since precipitation may have different effects on the evolutionary development of the forearm across the underlying subpopulations among bat species worldwide, we propose several testing procedures for hypotheses regarding the effect of precipitation on forearm length under finite GMR models. In addition to the real analysis of Chiroptera data, through simulation studies, we examine the performances of these testing procedures on their type I error rate, power, and consequently, the accuracy of clustering analysis.
The quest for effective tests of homogeneity within mixture models has a long history. An illustrative example provided by Hartigan shows that the likelihood ratio statistic, unlike its counterpart in regular models, tends to diverge to infinity even in the context of an extremely simplified normal mixture model. While imposing compact restrictions on the subpopulation parameter space and a separation condition can prevent this divergence and land on a limiting distribution, it does not lead to a practical testing procedure. In contrast, the C( $$\alpha $$ ) test, which is a modification of the popular score test, exhibits a simple limiting distribution and proves to be an effective tool for homogeneity testing in mixture models with single-parameter subpopulation distributions. Chapter 9 is dedicated to introducing the C( $$\alpha $$ ) test and providing specific expressions for its application within NEF-VEF mixtures.
While achieving a profound understanding of the limiting distribution of the likelihood ratio test under mixture models is indeed a grand success, translating this knowledge into concrete data analysis procedures can be challenging. One potential approach is to adhere to the fundamental principle of the likelihood ratio test but make slight adjustments to develop effective procedures. The modified likelihood ratio test for homogeneity is a product of this approach, grounded in a deep comprehension of two forms of problematic partial non-identifiability within the finite mixture model. By introducing a penalty term to the log likelihood function, the modified likelihood ratio test mitigates one of these issues, restoring a degree of regularity. Chapter 11 is dedicated to providing a detailed analysis of the limiting distribution for the modified likelihood ratio test for homogeneity.
The modified likelihood ratio test represents a significant advancement over the limitations of conventional likelihood ratio tests. It is applicable to a specific class of finite mixture models, although its success remains somewhat constrained. An interesting variation of this test is the EM-test, which may be seen as a derivative of the modified likelihood ratio test. In contrast to comparing the maximum possible modified likelihood values under null and alternative hypotheses concerning the order of the finite mixture model, the EM-test focuses on how quickly the likelihood increases when using EM-iterations from the best-fitted null model in a specific manner. This chapter introduces the concept and the conclusions related to the simple homogeneity case, with the more complex scenarios addressed in the subsequent chapter. It unveils the somewhat more intricate technical advantages of this approach.
The application of finite Gamma mixture models is prevalent when working with strictly positive observations. Surprisingly, the likelihood function within this framework is unbounded, theoretically leading to an inconsistent maximum likelihood estimator. However, it is worth noting that, in practice, this inconsistency rarely manifests in data analysis. Nevertheless, this chapter provides an in-depth exploration of the development of a consistent penalized maximum likelihood estimator for Gamma mixture models. In addition to this, it enriches the discussion with valuable insights gained through simulation experiments and real-world data examples.
Despite the successful application of the C( $$\alpha $$ ) test, statisticians continue to explore the utility of the likelihood ratio test for homogeneity under finite mixture models. For those who opt for the likelihood ratio test, a crucial undertaking involves determining the distribution or the limiting distribution of this test. This task becomes exceedingly challenging when the parameter space for subpopulations is unbounded. One solution to this challenge was proposed, albeit it required a seemingly unnecessary separation condition. However, after numerous dedicated efforts, a comprehensive and satisfying solution emerged. Chapter 10 does not go over the grand solution, but is dedicated to providing answers to specific scenarios where concrete results are attainable. This chapter serves to demystify the complexities associated with this formidable task.
The fundamental requirement in data analysis is the consistent estimation of a parameter. As the sample size increases, the precision of the estimator naturally improves, following a rate of $$n^{-1/2}$$ for parameters under regular statistical models. However, when dealing with finite mixture models, it becomes evident that this rate is strongly influenced by how excessively the order of the mixture is specified. Chapter 8 sheds light on the development of the claimed best possible rate, which was $$n^{-1/4}$$ when the order is over-specified. It is now widely recognized that the minimax rate is $$n^{-1/6}$$ even when the order is merely over-specified by one. The overall scenario is considerably more intricate than initially expected, and this chapter is dedicated to deepening the understanding of the optimal rate of convergence and its implications under finite mixture models.
The order selection problem, in contrast to determining whether a lower-ordered finite mixture model should be rejected in favor of a higher-ordered model, seeks to answer the question of what constitutes the most appropriate order for the finite mixture model. The challenge here lies in the fact that there are as many criteria for ”most suitable” as there are statisticians, making it a far more diverse problem than a simple hypothesis test. Moreover, each of these criteria often cannot be directly evaluated but requires approximations using complex asymptotic tools. The already intricate nature of finite mixture models further compounds this challenge. In Chap. 16, we offer a brief overview of several order selection procedures for finite mixture models. These include the transplanted Akaike Information Criterion (AIC), Bayes Information Criterion (BIC), as well as somewhat tailored Widely Applicable Bayesian Information Criterion (WBIC) and Singular Bayesian Information Criterion (sBIC), in addition to regularization-based techniques. This chapter serves as a limited introduction to these various methods without endorsing any particular approach.
The latent structure inherent in mixture models offers a fitting scenario for the renowned EM algorithm. In Chap. 7, we begin by offering a general overview and then get into comprehensive explanations of the EM algorithm’s application in computing the maximum likelihood estimate for finite mixture models. This chapter also addresses the critical matter of algorithm convergence, taking into consideration the global convergence theorem. Additionally, it provides specific insights into the workings of the EM algorithm.
While the success of the EM-test in the previous two chapters was confined to finite mixture models with subpopulation distributions belonging to a one-parameter distribution family, the underlying principle of the EM-test is generally applicable. This principle involves comparing the degree of improvement achieved after several EM-iterations starting from a fitted null model. However, this approach does not always yield a clear limiting distribution for the resulting test statistic, which is essential for practical testing procedures. A remarkable revelation occurs when the EM-test is applied to finite Gaussian mixture models. The meticulously designed EM-test for determining the order of the finite Gaussian mixture model leads to elegant and unexpected chi-square limiting distributions. These distributions align with the striking results seen in the likelihood ratio test for regular models. Chapter 15 offers a detailed account of this remarkable achievement.
The materials presented in this book are undeniably technical in nature. We established numerous theoretical conclusions by relying on classical probability theory results. Although these conclusions are well-known, many of us often struggle to recall the exact details. Notably, these conclusions are sometimes quoted from papers or books without specifying their applicability to a given context, making it challenging to ensure that the required conditions are met. Chapter 17 compiles some of these cited conclusions from the previous chapters to facilitate reference and ensure clarity regarding their specific contexts.
In Chap. 1, we introduce the fundamental concept of a statistical model and provide a detailed definition of both mixture models and finite mixture models. We uncover the latent structure inherent in mixture models, address the issue of identifiability, and explore various commonly utilized mixture models. Additionally, we highlight the relationship between mixture models and the models applied in the context of over-dispersed populations.
The finite normal mixture model stands out as the most frequently employed model in statistical applications. Nevertheless, it exhibits some peculiar characteristics. Notably, the general finite normal mixture model boasts an unbounded likelihood function, rendering the straightforward maximum likelihood estimator of the mixing distribution inconsistent. However, in the case of the univariate normal mixture with a structured scale parameter, the maximum likelihood estimator is consistent. To attain consistency in the likelihood-based approach, one can apply a penalty function to the likelihood. Chapter 4 charges into these issues and more, offering a comprehensive examination of the most crucial asymptotic properties.
Chapter 14 serves as a natural extension of the preceding chapter, offering a comprehensive examination of the EM-test when applied to high-order null hypotheses within the finite mixture model. While certain aspects of the results align with those of the modified likelihood ratio test, the unique design of the EM-test enables a clear presentation of the limiting distribution. Importantly, this presentation comes with fewer restrictions on the finite mixture model, making it applicable to a wider range of scenarios, including general orders.
In many statistical and econometric applications, we gather individual samples from various interconnected populations that undeniably exhibit common latent structures. Utilizing a model that incorporates these latent structures for such data enhances the efficiency of inferences. Recently, many researchers have been adopting the semiparametric density ratio model (DRM) to address the presence of latent structures. The DRM enables estimation of each population distribution using pooled data, resulting in statistically more efficient estimations in contrast to nonparametric methods that analyze each sample in isolation. In this article, we investigate the limit of the efficiency improvement attainable through the DRM. We focus on situations where one population's sample size significantly exceeds those of the other populations. In such scenarios, we demonstrate that the DRM-based inferences for populations with smaller sample sizes achieve the highest attainable asymptotic efficiency as if a parametric model is assumed. The estimands we consider include the model parameters, distribution functions, and quantiles. We use simulation experiments to support the theoretical findings with a specific focus on quantile estimation. Additionally, we provide an analysis of real revenue data from U.S. collegiate sports to illustrate the efficacy of our contribution.