
We provide necessary and sufficient conditions for the uniqueness of the k-means set of a probability distribution on a separable Banach space. This property is closely related to the choice of the parameter k, since some values may lead to multiple optimal clusterings, thereby complicating interpretation and affecting algorithmic stability. Within a risk minimization framework and under standard Donsker-type conditions, we derive the asymptotic distribution of the empirical within-cluster sum of squares (WCSS). This yields a probabilistic characterization of k-means uniqueness in terms of the limiting behavior of the empirical WCSS. Building on this characterization, we construct a bootstrap test for assessing k-means uniqueness. The finite-sample performance of the procedure is investigated through a simulation study.
The analysis of non-Euclidean data has gained increasing importance in modern statistics due to the growing prevalence of complex data structures across diverse domains. However, regression methods for predictors taking values in a broad class of metric spaces satisfying mild regularity conditions remain limited. We introduce a novel kernel regression framework for random object predictors in metric spaces, building on the concepts of the metric distribution function and the associated metric probability density function. Both local and global estimators are proposed, and their theoretical properties are established, including conditional L2 convergence rates and asymptotic normality under suitable regularity conditions. For practical implementation, we develop data-driven bandwidth selectors based on least-squares metric cross-validation. The proposed methods are validated through simulations on both manifold and non-manifold metric spaces, together with a real-data application involving stock indices and their constituent stocks, demonstrating favorable finite-sample performance and broad applicability.
We consider parametric inference in the drift term of a system of N identically distributed diffusion paths where a common noise is added to each path. This results in correlated identically distributed diffusions paths. The diffusion matrix of the system is supposed to be known. The purpose of this paper is to analyze the effects of a common noise on the statistical inference. The sample paths are continuously observed on a finite time interval [0,T]. The asymptotic framework is N tends to infinity. Having proved that the diffusion matrix of the system is invertible with explicit inverse, we study the exact likelihood. We prove that the statistical model has a limiting random information matrix (LAMN property). The exact maximum likelihood estimator is proved to be consistent and to converge with rate N to a mixture of Gaussian distributions. We compare the results to those obtained when there is no common noise (N i.i.d. sample paths). Various examples with constant and non constant diffusion terms including the Cox-Ingersoll-Ross and the bilinear processes are provided.
We study nonparametric drift estimation for a diffusion observed under continuous-time right censoring. The latent process X satisfies dXt=b(Xt)dt+σ(Xt)dWt, while the observer records only Yt=min(Xt,Ct) together with the indicator δt=1{Xt
In this paper, we consider a new class of semi-parametric regression models called the generalized partially linear spatially varying coefficient model (GPLSVCM). We propose using the bivariate penalized spline over triangulation (BPST) method to approximate the coefficient functions and employing a quasi-likelihood maximization to obtain model estimators. The proposed method can handle data distributed over arbitrarily shaped domains with complex boundaries and interior holes. We prove the consistency of the varying coefficient estimators and establish the asymptotic normality of the linear coefficients under some regularity conditions. Additionally, we propose a model selection procedure via BIC that can accurately identify the covariates with constant and varying effects. The proposed methods are evaluated with several benchmarks through simulation studies and a real data analysis.
Suppose (X-i is an element of R-P, y(i)) comes from a two-class Gaussian mixture model, where Chi(i) similar to Nu(mu(eta), Sigma(yi)) and y(i) is an element of {0,1}, 1 <= i <= n The goal is to estimate the class label Y for a new observation X. When the mean vector mu(yi), and the covariance matrix Sigma(yi), depend on the class label (yi), the quadratic discriminant analysis (QDA) algorithm leverages both the covariance and mean to achieve better classification results. However, QDA faces both theoretical and computational challenges in the high-dimensional setting that p >> n. In this paper, we propose two high-dimensional QDA algorithms: QDAw and QDAfs. When the non-zeros in & micro;(0) - mu(1 )are weak and moderately dense, the QDAw algorithm with an unbiased factor in the classification rule works well. When the non-zeros in & micro;(0) - mu(1) are sparse and moderately strong, the QDAfs algorithm with a thresholding step should be used. By modeling the sparsity and strength of non-zero entries in Sigma(-1)(0) - Sigma(-1)(1) and (0) - mu(1). the upper bounds of QDAw and QDAfs meet the lower bounds of the classification problem on the 4-dimensional parameter space, which proves the optimality of the algorithms. Our theory is supported by real data analysis results.
Two-sample hypothesis testing-determining whether two sets of data are drawn from the same distribution-is a fundamental problem in statistics and machine learning with broad scientific applications. In the context of nonparametric testing, maximum mean discrepancy (MMD) has gained popularity as a test statistic due to its flexibility and strong theoretical foundations. However, its use in large-scale scenarios is plagued by high computational costs. In this work, we use a Nystr & ouml;m approximation of the MMD to design a computationally efficient and practical testing algorithm while preserving statistical guarantees. Our main result is a finite-sample bound on the power of the proposed test for distributions that are sufficiently separated with respect to the MMD. The derived separation rate matches the known minimax optimal rate in this setting. We support our findings with a series of numerical experiments, emphasizing applicability to realistic scientific data.
As network data has grown in popularity and ubiquity, collections of networks have become a common object of study, with appropriate data analysis typically requiring the identification of structural similarities and differences across networks and at different scales within them. We describe a statistically principled, scalable "omnibus embedding", in which multiple graphs on the same vertex set are jointly mapped into a single space with a distinct representation for each graph. This embedding streamlines graph comparison and comes with performance guarantees, including consistency and a central limit theorem, for the estimation of underlying network parameters. The joint embedding and accompanying central limit theorem provide solutions to several multiscale graph inference questions, such as the identification of graph-wide and vertex-specific differences across networks. We show that in simulated data, the omnibus embedding exhibits near-optimal estimation accuracy when the networks have the same generative structure, while still preserving discriminative power when the networks are different. We also analyze neuroscientific data collected from human subjects, and use the omnibus embedding to successfully identify specific brain regions associated with structural differences and markers of pathology.
We study a simple & ell;1-regularized generalized least-squares (GLS) estimator for high-dimensional regressions with autocorrelated errors. The estimation procedure consists of three steps: performing a LASSO regression, fitting an autoregressive model to the realized residuals, and then running a second-stage LASSO regression on the rotated (whitened) data. We examine the theoretical performance of the method in a sub-Gaussian random-design setting, in particular assessing the impact of the rotation on the design matrix and how this impacts the estimation error of the procedure. We show that the GLS (and a feasible variant) maintains a smaller estimation error than an unadjusted LASSO regression when the errors are driven by an autoregressive process. A simulation study verifies the performance of the proposed method, demonstrating that the penalized (feasible) GLS-LASSO estimator performs on par with the LASSO in the case of white noise errors, whilst outperforming when the errors exhibit significant autocorrelation.
We propose a deep generative approach to semi-supervised classification. The proposed approach is based on the idea that learning a conditional class probability function is equivalent to learning a conditional generator function for the conditional class probability. This idea leads to the construction of an objective function for learning a classifier and a conditional class generator that naturally combines the information from both labeled and unlabeled observations. We take advantage of the approximation power of deep neural networks to approximate the classifier and the conditional generator nonparametrically. We establish consistency and convergence properties of the resulting estimators with respect to certain metric under suitable conditions. Moreover, we conduct numerical studies with simulated and real data to evaluate the proposed method and illustrate that the proposed method outperforms several existing semi-supervised classification methods as well as supervised methods trained with only labeled samples, especially when the size of the labeled sample is small.
This paper introduces a rigorous approach to establish the sharp minimax optimalities of both sparse group LASSO and SLOPE, where parameters exhibit both element-wise and group-wise sparsity. We derive an original upper bound for Gaussian complexity based on a weighted [1 + [1,2 norm. Based on this result, we establish the sharp optimal results for both sparse group LASSO and SLOPE under restricted eigenvalue (RE) type conditions. To the best of our knowledge, the assumptions employed in this paper are the weakest among existing analyses. We find the astonishing result that the intersection of sparse restricted eigenvalue cones for element-wise sparsity and for group-wise sparsity corresponds to the solution of weighted [1 + [1,2 penalty. Furthermore, based on the proof techniques, we can provide the optimal sample complexity under the framework of random design, where both weak moment assumption and sub-Gaussian random design are considered. Numerical experiments also confirm that, under the double sparsity setting, our algorithm achieves superior performance compared to methods based on only single type of sparsity assumption. Our analysis technique can also be extended to the theoretical analysis of other general structured sparsity problems.
In this paper, we study high-dimensional quantile linear models with a scalar response and non-Euclidean covariates taking values in Riemannian Hilbert manifolds. The conditional quantile function is constructed using Hilbert-Schmidt operators and then reformulated in terms of real-valued scores obtained from spectral decomposition. To overcome the non-smoothness and lack of strong convexity of the traditional quantile loss, we adopt a convolution smoothing technique. The resulting penalized optimization problem is solved by a groupwise iteratively locally adaptive majorize-minimization algorithm. We derive C2-and C1-type error bounds for the initial LASSO estimator and prove the contraction property for the sequence of iteratively updated estimators. Furthermore, when employing penalty functions with the vanishing-gradient property, we establish the strong oracle property. Numerical simulations and a real data example demonstrate the practical effectiveness of the proposed method.
Causal implications can be undermined by control and treatment groups that are unbalanced by external confounders. Furthermore, the interventions on the treatment group can target unknown nodes of a Directed Acyclic Graph (DAG), with unknown alterations on the structure of dependencies. We propose a Bayesian methodology on a graph-driven multivariate Gaussian potential outcome model that identifies the unknown target nodes, quantifies the related causal effects and learns the network of dependencies pre-and post-interventions, accounting for observed confounders. For the purpose, we extend the DAG-Wishart prior to a Normal-DAG-Wishart prior in presence of covariates, whose conjugacy and marginal likelihood are derived. We first study the asymptotic properties of the propensity score Bayesian estimator, used to balance control and treatment groups from confounding effects. We then show the posterior ratio consistency of treated and untreated graphs, and the limiting distribution of the Average Treatment Effect estimator, under different asymptotic scenarios in terms of treated and control sample sizes. The theoretical results are validated on simulated data by an appropriately developed MCMC posterior sampler, and then implemented on Acute Myeloid Leukemia malignancies, to evaluate network dependencies, targets and causal effects of Histone Deacetylases inhibitor treatments.
For datasets with unknown but stationary serial dependence, a robust long run variance estimator is essential to handle diverse scenarios. Spectral variance estimators are commonly used but tend to exhibit significant negative bias in the presence of positive correlation. To overcome this, zero lugsail estimators have been introduced, offering zero asymptotic bias regardless of the correlation structure. However, there are currently no guidelines for selecting the optimal bandwidth for lugsail estimators, a critical component in the estimation process. We propose an inference optimal bandwidth rule for lugsail estimators, based on nonstandard fixed-smoothing limiting distributions developed in our study. This approach significantly improves bias correction, accounts for variability, and provides an estimator optimized for robust inference. Our theoretical findings are supported by a simulation study.
Tests of independence are an important tool in applications, specifically in connection with the detection of a relationship between variables; they also have initiated many developments in statistical theory. In the present paper we build upon and extend a recently established link to Discrete Mathematics and Theoretical Computer Science, exemplified by the appearance of copulas in connection with limits of permutation sequences, and by the connection between quasi-randomness and consistency of pattern-based tests of independence. The latter include classical procedures, such as Kendall's tau, which uses patterns of length two. Longer patterns lead to tests that are consistent against large classes of alternatives, as first shown by Hoeffding (1948) with patterns of length five, and by Yanagimoto (1970) and Bergsma and Dassios (2014) for patterns of length four. More recently Chan et al. (2020) characterized quasi-randomness for sets of patterns of length four, which leads to several new consistent pattern-based test for independence. We give a detailed and complete description of the respective limiting null distributions. In connection with the power performance of the tests, which is of interest for practical purposes, we provide results on their (local) asymptotic relative efficiencies. We also include a small simulation study that supports our theoretical findings.
Motivated by a recent method for approximate solution of Fred-holm equations of the first kind, we develop a corresponding method for a class of Fredholm equations of the second kind. In particular, we consider the class of equations for which the solution is a probability measure. The approach centres around specifying a functional whose gradient flow admits a minimizer corresponding to a regularized version of the solution of the underlying equation and using a mean-field particle system to approximately simulate that flow. Theoretical and numerical support for the method is presented.
The Gaussian Kinematic Formula (GKF) is a powerful and computationally efficient tool to perform statistical inference on random fields and became a well-established tool in the analysis of neuroimaging data. Using realistic error models, [10] recently showed that GKF based methods for voxelwise inference lead to conservative control of the familywise error rate (FWER) and to inflated false positive rates for cluster-size inference. In a series of two articles we identify and resolve the main causes of these shortcomings in the traditional usage of the GKF for voxelwise inference. In particular applications of Random Field Theory have depended on the good lattice assumption, which is only reasonable when the data is sufficiently smooth. We address this here, allowing for valid inference under arbitrary applied smoothness and non-stationarity, for Gaussian and Gaussian related fields. We address the assumption of Gaussianity in the follow-up article, where we also demonstrate that our GKF based methodology for voxelwise inference is non-conservative under realistic error models.
We study the problem of robust estimation and inference for contaminated data in distributed learning systems. Existing methods, such as those designed for Byzantine failures, typically assume that contamination is limited to a small subset of machines experiencing complete failure. In contrast, we address distributed partial contamination, where all machines may handle datasets containing some corrupted observations. By generalizing Huber's ϵ -contamination model to distributed settings, we propose a robust framework in which each machine's dataset may be contaminated by a proportion ϵ of observations from an arbitrary distribution. In a linear model setting, we develop a communication-efficient M-estimator that achieves the optimal convergence rate of centralized data. For robust inference, we introduce a distributed multiplier bootstrap method that requires no additional communication post-estimation while maintaining efficiency. To handle high contamination proportions, we present a debiasing procedure to mitigate bias. Extensive simulations demonstrate the robustness and scalability of our methods across diverse contamination scenarios.
This paper studies the asymptotic properties for the counts of successive hitting times of path-dependent functionals of continuous It & ocirc; semimartingales in a fixed time span. We prove that for a wide class of functionals, the counts of hitting times can be used to construct consistent and asymptotically normal estimators of the quadratic variation of the underlying continuous It & ocirc; semimartingale, which can achieve a sizeable variance reduction compared to the existing realized variance-type methods under a common sampling frequency.
This work is concerned with the Backfitting (BF) algorithm for the nonparametric additive quantile regression (AQR) model. We establish a strong uniform consistency rate for the Bahadur representation of the BF estimators, which includes the two stage estimator of [11] as a special case. These results are fundamental for statistical inference in AQR models and for applications that involve plugging such estimators into other functionals where some control over higher order terms is required. Two examples of applications are provided: one is on the root-n consistency of the parameter estimates for the partially linear AQR model, the other concerns structure recovery for the nonparametric AQR model. MSC2020 subject classifications: Primary 62G08, 62H12; secondary 62G20.