To deal with superfluous clusters and to reduce the number of clusters as often desired in practice, a modal cluster merging procedure is proposed. Based on some new properties established in this paper for Morse functions, the procedure merges clusters in a sequential manner without causing unnecessary density distortion. Each cluster is evaluated for its significance relative to the other clusters, using the Kullback-Leibler divergence or its log-likelihood approximation, by truncating the density for the cluster at an appropriate level. The least significant cluster is then merged into one of its adjacent clusters, using the novel concept of cluster adjacency defined in this paper. The resulting hierarchical clustering tree is useful for determining the number of clusters, as may be preferred by a specific user or in a general, meaningful manner. Numerical studies show that the new procedure deals well with difficult clustering problems and often produces intuitively appealing and numerically more accurate clustering results, as compared with several other popular clustering methods in the literature.
The conditional independence assumption has recently appeared in a growing body of literature on the estimation of multivariate mixtures. We consider here conditionally independent multivariate mixtures of power series distributions with infinite support, to which belong Poisson, Geometric or Negative Binomial mixtures. We show that for all these mixtures, the non-parametric maximum likelihood estimator converges to the truth at the rate (ln(nd))(1+d/2)n(-1/2) in the Hellinger distance, where n denotes the size of the observed sample and d represents the dimension of the mixture. Using this result, we then construct a new non-parametric estimator based on the maximum likelihood estimator that converges with the parametric rate n(-1/2) in all l(p)-distances, for p >= 1. These convergences rates are supported by simulations and the theory is illustrated using the famous V & eacute;lib dataset of the bike sharing system of Paris. We also introduce a testing procedure for whether the conditional independence assumption is satisfied for a given sample. This testing procedure is applied for several multivariate mixtures, with varying levels of dependence, and is thereby shown to distinguish well between conditionally independent and dependent mixtures. Finally, we use this testing procedure to investigate whether conditional independence holds for V & eacute;lib dataset.
We consider the problem of estimating a mixture of power series distributions with infinite support, to which belong very well-known models such as Poisson, Geometric, Logarithmic or Negative Binomial probability mass functions. We consider the nonparametric maximum likelihood estimator (NPMLE) and show that, under very mild assumptions, it converges to the true mixture distribution r0 at a rate no slower than (log n)3/2n-1/2 in the Hellinger distance. Recent work on minimax lower bounds suggests that the logarithmic factor in the obtained Hellinger rate of convergence can not be improved, at least for mixtures of Poisson distributions. Furthermore, we construct nonparametric estimators that are based on the NPMLE and show that they converge to r0 at the parametric rate n-1/2 in the [p-norm (p E [1, 00] or p E [2, 00]): The weighted least squares and hybrid estimators. Simulations and a real data application are considered to assess the performance of all estimators we study in this paper and illustrate the practical aspect of the theory. The simulations results show that the NPMLE has the best performance in the Hellinger, [1 and [2 distances in all scenarios. Finally, to construct confidence intervals of the true mixture probability mass function, both the nonparametric and parametric bootstrap procedures are considered. Their performances are compared with respect to the coverage and length of the resulting intervals.
Compositional data, representing proportions constrained to the simplex, arise in diverse fields such as geosciences, ecology, genomics, and microbiome research. Existing nonparametric density estimation methods often rely on transformations, which may induce substantial bias near the simplex boundary. We propose a nonparametric mixture-based framework for density estimation on compositions. Nonparametric Dirichlet mixtures are employed to naturally accommodate boundary values, thereby avoiding the transformation or zero-replacement, while also identifying components supported on the boundary, providing reliable estimates for data with zero or near-zero values. Bandwidth selection and initialization schemes are addressed. For comparison, nonparametric Gaussian mixtures, coupled with log-ratio transformations, are also considered. Extensive simulations show that the proposed estimators outperform existing approaches. Three real data applications, including GDP data analysis, handwritten digit recognition, and skin detection, demonstrate the usefulness of nonparametric Dirichlet mixtures in practice.
A new method for estimating the proportion of null effects is proposed for solving large-scale multiple comparison problems. It utilises maximum likelihood estimation of nonparametric mixtures, which also provides a density estimate of the test statistics. It overcomes the problem of the usual nonparametric maximum likelihood estimator that cannot produce a positive probability at the location of null effects in the process of estimating nonparametrically a mixing distribution. The profile likelihood is further used to help produce a range of null proportion values, corresponding to which the density estimates are all consistent. With a proper choice of a threshold function on the profile likelihood ratio, the upper endpoint of this range can be shown to be a consistent estimator of the null proportion. Numerical studies show that the proposed method has an apparently convergent trend in all cases studied and performs favourably when compared with existing methods in the literature.
Toroidal data is an extension of circular data on a torus and plays a critical part in various scientific fields. This article studies the density estimation of multivariate toroidal data based on semiparametric mixtures. One of the major challenges of semiparametric mixture modelling in a multi-dimensional space is that one can not directly maximize the likelihood over the unrestricted component density as it will result in a degenerate estimate with an unbounded likelihood. To overcome this problem, we propose to fix the maximum of the component density, which subsequently bounds the maximum of the mixture and its likelihood function, hence providing a satisfactory density estimate. The product of univariate circular distributions are utilized to form multivariate toroidal densities as candidates for mixture components. Numerical studies show that the mixture-based density estimator is superior in general to the kernel density estimator.
Nonparametric density estimation is studied for spherical data that may arise in many scientific and practical fields. In particular, nonparametric mixture models based on likelihood maximization are used. A nonparametric mixture has component distributions mixed together with a mixing distribution that is completely unspecified and needs to be determined from data. For mixture components, a two-parameter distribution family can be used, with one parameter as the mixing variable and the other to control the smoothness of the density estimator. For example, the popular von Mises-Fisher distributions can be readily used for this purpose. Numerical studies with various spherical data sets show that the resultant mixture-based density estimators are strong competitors with the best of the other density estimators.
Modal clustering has a clear population goal, where density estimation plays a critical role. In this paper, we study how to provide better density estimation so as to serve the objective of modal clustering. In particular, we use semiparametric mixtures for density estimation, aided with a novel mode-flattening technique. The use of semiparametric mixtures helps to produce better density estimates, especially in the multivariate situation, and the mode-flattening technique is intended to identify and smooth out spurious and minor modes. With mode flattening, the number of clusters can be sequentially reduced until there is only one mode left. In addition, we adopt the likelihood function in a coherent manner to measure the relative importance of a mode and let the current least important mode disappear in each step. For both simulated and real-world data sets, the proposed method performs very well, as compared with some well-known clustering methods in the literature, and can successfully solve some fairly difficult clustering problems.
Data visualization is important for statistical analysis, as it helps convey information efficiently and shed lights on the hidden patterns behind data in a visual context. It is particularly helpful to display circular data in a two-dimensional space to accommodate its nonlinear support space and reveal the underlying circular structure which is otherwise not obvious in one-dimension. In this article, we first formally categorize circular plots into two types, either height- or area-proportional, and then describe a new general methodology that can be used to produce circular plots, particularly in the area-proportional manner, which in our opinion is the more appropriate choice. Formulas are given that are fairly simple yet effective to produce various circular plots, such as smooth density curves, histograms, rose diagrams, dot plots, and plots for multiclass data. Supplemental materials for this article are available online.
A new density-based classification method that uses semiparametric mixtures is proposed. Like other density-based classifiers, it first estimates the probability density function for the observations in each class, with a semiparametric mixture, and then classifies a new observation by the highest posterior probability. By making a proper use of a multivariate nonparametric density estimator that has been developed recently, it is able to produce adaptively smooth and complicated decision boundaries in a high-dimensional space and can thus work well in such cases. Issues specific to classification are studied and discussed. Numerical studies using simulated and real-world data show that the new classifier performs very well as compared with other commonly used classification methods.
针对B样条表达的空间自由曲线形状的相似性评价问题,提出了基于曲率特征的相似性评价算法.首先在自由曲线上等弧长均匀取点,并计算各点的曲率,得到曲线的曲率分布;然后以曲率分布作为描述空间曲线的几何特征,采用EMD (Earth Mover's Distance)算法计算两个曲率分布的距离,衡量对应曲线的相似程度.算法以提取曲率特征的方法将三维空间曲线的相似性比较问题转化为二维分布的距离计算,简单有效地度量空间曲线形状的相似性.为检验算法的效果,以VS2010为集成开发环境对不同类型的B样条空间曲线进行大量的试验,结果表明,提出的相似性比较算法可行有效,且对旋转和取点顺序有较好的鲁棒性,能很好地反映空间曲线的相似程度.
Algorithms for computing the nonparametric maximum likelihood estimate of a univariate log‐concave density are briefly described, and some relevant issues discussed. The fast few algorithms are further numerically compared in a small‐scale simulation study. This article is categorized under: Statistical and Graphical Methods of Data Analysis > Markov Chain Monte Carlo (MCMC) Statistical and Graphical Methods of Data Analysis > Density Estimation Statistical and Graphical Methods of Data Analysis > Nonparametric Methods Algorithms and Computational Methods > Algorithms
A new fast algorithm for computing the nonparametric maximum likelihood estimate of a univariate log-concave density is proposed and studied. It is an extension of the constrained Newton method for nonparametric mixture estimation. In each iteration, the newly extended algorithm includes, if necessary, new knots that are located via a special directional derivative function. The algorithm renews the changes of slope at all knots via a quadratically convergent method and removes the knots at which the changes of slope become zero. Theoretically, the characterisation of the nonparametric maximum likelihood estimate is studied and the algorithm is guaranteed to converge to the unique maximum likelihood estimate. Numerical studies show that it outperforms other algorithms that are available in the literature. Applications to some real-world financial data are also given.
A new algorithm is presented and studied in this paper for fast computation of the nonparametric maximum likelihood estimate of a U-shaped hazard function. It successfully overcomes a difficulty when computing a U-shaped hazard function, which is only properly defined by knowing its anti-mode, and the anti-mode itself has to be found during the computation. Specifically, the new algorithm maintains the constant hazard segment, regardless of its length being zero or positive. The length varies naturally, according to what mass values are allocated to their associated knots after each updating. Being an appropriate extension of the constrained Newton method, the new algorithm also inherits its advantage of fast convergence, as demonstrated by some real-world data examples. The algorithm works not only for exact observations, but also for purely interval-censored data, and for data mixed with exact and interval-censored observations.
Nonparametric mixture models are commonly used to estimate the number of unobserved species for their robustness of modelling heterogeneity. In particular, Poisson mixtures are popular in this regard, but they are also known to have the boundary problem that causes unstable estimation. A family of shape-restricted distributions known as discrete k-monotone distributions is proposed for species richness estimation. These distributions are in fact also nonparametric mixtures and can thus be fitted rapidly via some algorithms that are made available recently. As shown by empirical studies, as compared with Poisson mixtures, their use avoids the boundary problem and gives more stable and more accurate estimates.
A new method is proposed for nonparametric multivariate density estimation, which extends a general framework that has been recently developed in the univariate case based on nonparametric and semiparametric mixture distributions. The major challenge to a multivariate extension is the dilemma that one can not maximise directly the likelihood function with respect to the whole component covariance matrix, since the likelihood is unbounded for a singular covariance matrix, and that one can not leave the covariance matrix or a large part of it to be determined by a selection method, since it would be computationally infeasible. To overcome it, we consider using a volume parameter h to enforce a minimal restriction on the covariance matrix so that, with h fixed, the likelihood function is bounded and its maximisation can be successfully carried out with respect to all the remaining parameters. The role played here by the scalar h is just the same as by the bandwidth in the univariate case and its value can be determined by a model selection criterion, such as the Akaike information criterion. New efficient algorithms are also described for finding the maximum likelihood estimates of these mixtures under various restrictions on the covariance matrix. Empirical studies using simulated and real-world data show that the new multivariate mixture-based density estimator performs remarkably better than kernel-based density estimators.
The fact that a k-monotone density can be defined by means of a mixing distribution makes its estimation feasible within the framework of mixture models. It turns the problem naturally into estimating a mixing distribution, nonparametrically. This paper studies the least squares approach to solving this problem and presents two algorithms for computing the estimate. The resulting mixture density is hence just the least squares estimate of the k-monotone density. Through simulated and real data examples, the usefulness of the least squares density estimator is demonstrated.
We propose a new algorithm for computing the maximum likelihood estimate of a nonparametric survival function for interval-censored data, by extending the recently-proposed constrained Newton method in a hierarchical fashion. The new algorithm makes use of the fact that a mixture distribution can be recursively written as a mixture of mixtures, and takes a divide-and-conquer approach to break down a large-scale constrained optimization problem into many small-scale ones, which can be solved rapidly. During the course of optimization, the new algorithm, which we call the hierarchical constrained Newton method, can efficiently reallocate the probability mass, both locally and globally, among potential support intervals. Its convergence is theoretically established based on an equilibrium analysis. Numerical study results suggest that the new algorithm is the best choice for data sets of any size and for solutions with any number of support intervals.