
ABSTRACT Measuring the correlation between two random vectors is required and plays an important role in almost every scientific field and correspondingly many criteria have been proposed in the literature. In this paper, we investigate a new correlation measure and propose an empirical estimator for interval‐censored failure time data. The theoretical guarantee of the new measure and the asymptotic properties of the estimator are established and in particular, the estimator is shown to be consistent and asymptotically normally distributed. One main advantage of the new correlation is that it significantly reduces the computational burden compared to some existing criteria for high‐dimensional data. By using the new measure, a model‐free feature screening procedure is proposed and both sure screening property and rank consistency of the proposed approach are established. An extensive simulation study indicates that the screening procedure works well in practical situations and an illustration is provided.
ABSTRACT We consider the problem of pairwise Kalman filter (PKF) robustness in the context of Gaussian homogeneous Markov pairwise models (GH‐PMMs). We provide a fast and accurate sequential calculation of the mean square error increase for the pairwise Kalman filter when the parameters of the model used deviate from the actual parameters. We illustrate the relevance of the theoretical results by studying filter deterioration in a specific situation, where one replaces the real PMM model by the classical homogeneous Gaussian state‐space model (GH‐SSM). We also present an application to the real‐world problem of estimating soil moisture by filtering the temperature sequence.
ABSTRACT This article considers maximum spacing (MSP) estimation for multivariate observations under model misspecification. A broad class of MSP estimators corresponding to different information‐type divergence measures is studied. This class also includes the MSP estimator corresponding to the Kullback‐Leibler information measure, obtained using the logarithmic function. It is shown that, under model misspecification, the MSP estimator converges to a well‐defined limit that depends on the chosen divergence measure; this establishes consistency. The behavior of MSP estimators under different divergence measures is explored through simulation studies. The asymptotic properties of MSP estimators and spacing functions are exploited to illustrate tools for detecting model misspecification in large samples.
ABSTRACT The problem of measuring the similarity or homology of (finite alphabet) strings appears in many areas of applications. A most common similarity measure is the length of the longest common subsequence (LCS). To every common subsequence corresponds at least one alignment, the alignments corresponding to LCS are called optimal. The set of optimal alignments is not unique and in [1] it was proposed to use the diversity of optimal alignments as another measure of similarity. This diversity was quantified via the distance between extremal alignments. In [1], it was shown that for strongly related sequences, this distance increases logarithmically and the more related the sequences, the slower is the growth. In the present paper, we consider another model—a pairwise Markov model—and show that for strongly related sequences the logarithmic growth still holds. Based on this logarithmic growth, a rate of convergence for the mean length of LCS is obtained. This rate is of order , which is faster than the best current rate for the iid independent case. Moreover, we show that extremal alignments are a useful technical tool for studying properties of LCS. In particular, a bound to the unknown Chvatal–Sankoff constant is derived.
Generalized factor models are gaining traction for multivariate data dimension reduction due to their flexibility. This paper investigates the sparse estimation of generalized factor models with weaker loadings. We introduce a group-wise penalized estimation approach, which results in a sparse loading matrix. This sparsity not only facilitates variable selection but also enhances the interpretability of the reduced-dimensional results. To tackle computational challenges, we develop a projected alternating maximization algorithm, achieving simultaneous parameter estimation and variable selection. Considering the importance of determining the number of factors, we propose a sparsity information criterion for this purpose. Under the weak loadings assumption and other mild conditions, we establish upper and lower error bounds for the overall parameter estimates, derive the convergence rates for the loading matrix and factor scores, and demonstrate the consistency of both variable selection and the number of factors. Furthermore, we extend the model to accommodate missing data, providing corresponding theoretical guarantees. The efficacy of our proposed method is validated through extensive simulations and applications to two real-world datasets.
We look at Monte Carlo numerical integration from a stochastic geometry point of view. While crude Monte Carlo estimators relate to linear statistics of a homogeneous Poisson point process (PPP), linear statistics of more regularly spread point processes can yield unbiased estimators with faster-decaying variance, and thus lower integration error. Following this intuition, we introduce a Coulomb repulsion operator, which reduces clustering by slightly pushing the points of a configuration away from each other. Our empirical findings show that applying the repulsion operator to a PPP as well as, intriguingly, to more regular point processes, preserves unbiasedness while reducing the variance of the corresponding Monte Carlo estimator, thus enhancing the method. We prove this variance reduction when the initial point process is a PPP. On the computational side, the complexity of the operator is quadratic, and the corresponding algorithm can be parallelized without communication across tasks.
Modern biomedical research increasingly relies on longitudinal studies with repeated measurements and complex within-subject dependence. In this paper, we develop a new framework for automatic variable selection in longitudinal quantile regression, motivated in part by applications to Alzheimer's disease research. Our approach combines the quadratic inference function (QIF) methodology with smooth-threshold estimating equations (SEEs) to accommodate within-subject correlation while enabling computationally efficient estimation and automatic variable selection in settings where the number of covariates may diverge. We establish variable selection consistency of the proposed method and show that the Bayesian information criterion can be used to select tuning parameters in a principled manner. Simulation studies demonstrate strong performance in both estimation accuracy and selection reliability. Finally, we apply the proposed procedure to data from the Alzheimer's Disease Neuroimaging Initiative (ADNI), identifying biomarkers and risk factors associated with extreme cognitive decline.
Recent interest has emerged in community detection for dynamic networks, which are observed along a trajectory of points in time. In this paper, we present a time-varying (t-v) degree-corrected stochastic block model to fit a dynamic network which allows evolving heterogeneity in the degrees of nodes within a community over time. In order to aggregate network information over time via a sliding window, we propose a smoothing-based method that simultaneously allows to recover the: (i) t-v node connection probabilities; (ii) t-v community memberships; and (iii) t-v node degree parameters. As a novelty compared to existing literature, we provide asymptotic theory for all of these three goals, that is, we derive rates of consistency of our smooth estimators for degree parameters and communities using a time-localized profile-likelihood approach as well as strong consistency of community membership estimation at every time point. Extensive simulation studies and applications to two different real data sets complete our work.
In high-dimensional regression, covariates often have a natural grouping structure, and it is crucial to perform analyses at the group level rather than the individual covariate level. Group regression is an effective approach for leveraging inherent grouping structures among covariates, enhancing interpretability, and improving analyses. Penalized group regression methods such as group lasso impose structured sparsity or group-level penalties to obtain robust and interpretable results. However, there are two major limitations with such existing methods: first, they are developed based on a one-penalty-type-fits-all approach, which can be restrictive and suboptimal in practice, and second they only work if groups of covariates are specified in advance. To address these issues, we provide a general framework for high-dimensional group regression and propose methods that allow different types of penalties for groups depending on group structures. We develop a novel group correlation learning method to identify groups of covariates in situations where no groups are prespecified. We provide extensive theoretical results for our methods, including upper bounds for estimation and prediction errors along with asymptotic rates. We also obtain the exact distributions of our group-penalized estimators for statistical inference. We provide an R package called HDGR for the implementation of the proposed methods.
In this paper, we study the generalized empirical likelihood inference for a partially linear model with measurement errors in both the covariates of parametric and nonparametric parts, on the basis of longitudinal data. The proposed generalized empirical likelihood approach considers within-group correlations to estimate regression coefficients and does not involve direct estimation of nuisance parameters in the correlation matrix. At the true parameter, the empirical -likelihood ratio is proved to be asymptotically chi-square distributed under regularity conditions, and the corresponding confidence intervals are constructed. Next, the profile EL ratios for the parameter are given, and it is verified that these ratios follow the chi-square distribution. Furthermore, the strong consistency and the convergence rate for the estimator of a nonparametric function are obtained, and the -consistency of the estimator of is verified. The performance of the proposed method is verified by numerical simulation and real data analysis.
In modern statistics, testing for dependence is often essential. The complex relationships between covariates and response variables present significant analytical challenges. While kernel-based independence analysis has become a powerful alternative to tackle these issues, there is currently no universally effective kernel-based test available. We introduce a series of kernel-based independence tests within the framework of reproducing kernel Hilbert spaces (RKHS). This work includes explicit sample-level expressions for these tests, as well as their asymptotic null distributions. Additionally, we develop two specific tests: The Maximal Kernel-based Independence Test (MKIT) and the Maximin Efficient Robust Test (MERT), both of which are derived from the broader category of kernel-based independence tests. Theoretically, we prove that MKIT and MERT asymptotically conform to the extreme-value type I-Gumbel distribution and the normal distribution under certain regular conditions, respectively, and analyze the powers of MKIT and MERT. We conduct extensive simulations that show that MKIT and MERT outperform numerous existing methods across various scenarios. Applications to heterogeneous stock mice data and human connectome project data further highlight the superior performance of the proposed test methods.
Models with latent factors have recently attracted considerable attention. However, most existing studies focus on linear regression models and therefore fail to capture potential nonlinear structures. To address this limitation, we consider the factor augmented single-index model. We first examine whether the inclusion of the augmented component is necessary by introducing a score-type test statistic. Unlike existing test statistics, the proposed one does not require estimating high-dimensional regression coefficients or precision matrices, making it computationally simpler and more stable. To determine the critical value, we employ a Gaussian multiplier bootstrap, whose theoretical validity is established under mild regularity conditions. We further investigate the penalized estimation of the regression model. With estimated latent factors, we establish the error bounds of the estimators. In addition, we construct confidence intervals for individual coefficients based on a debiased estimator. Importantly, our method does not impose moment conditions on the error distribution, allowing it to perform well even when the errors are heavy-tailed or contaminated by outliers. Comprehensive simulation studies and an application to a gene expression dataset demonstrate the effectiveness and robustness of the proposed procedure.
We propose to deal with high-dimensional event-based data using Hawkes processes. Focusing on the Multivariate Hawkes Processes (MHP) in high dimension, an estimation task, followed by a classification task, are addressed in this article. In both cases, we assume to have access to a large number of repeated observations of the process over the same short time interval. MHPs form a versatile class of point processes that model interactions among connected individuals within a network. In this work, we allow the network dimension to be large relative to the number of observations, which necessitates a sparsity assumption on the adjacency matrix. Furthermore, we assume that the observations belong to different classes, discriminated by both the exogenous intensity vector and the adjacency matrix, which encodes the strength of interactions. Specifically, the observed training data consist of labeled, repeated, and independent realizations over a fixed time interval. In this context, we propose a novel methodology comprising an initial interaction recovery step, conducted per class, followed by a refitting step guided by a suitable classification criterion. To recover the support of the adjacency matrix in each class, we introduce a Lasso-type estimator and prove the consistency of the estimated supports under appropriate assumptions on the processes. Leveraging the estimated supports, we then construct a classification procedure based on empirical error minimization. Notably, we provide convergence rates for our classifier. An in-depth numerical study, using both synthetic and real-world datasets, supports our theoretical findings, both for support recovery and for supervised classification.
The estimation of points at which regression coefficients change has attracted considerable interest across many research fields. As in other regression contexts, missing data are ubiquitous. However, while much of the existing literature on missing data focuses on estimating regression coefficients, very few studies address missing data in the context of change point estimation. Linear spline models are powerful tools for studying change points; however, as we demonstrate in simulations, improperly handled missing outcomes can lead to biased estimates of change points. To address this, we propose two novel estimators for change points in linear spline models with potentially missing outcomes: an inverse probability weighting (IPW) estimator and a doubly robust augmented inverse probability weighting (DR-AIPW) estimator, and study an imputation-based outcome regression (OR) estimator. We establish the consistency and asymptotic normality of these estimators. Moreover, the DR-AIPW estimator provides dual protections for consistency, and its optimality is also verified using semi-parametric theory. We develop two-step IPW, DR-AIPW, and OR algorithms for implementation. Simulation studies and real-data applications demonstrate the strong numerical performance of the proposed estimators.
This paper is motivated by the problem of optimal allocation of trials in multi-environment crop variety testing with a large number of varieties. Optimizing the allocation of trials results in the minimization of a design criterion with a Kronecker product structure in the information matrix. We consider the Kronecker-Bayesian linear criterion, which generalizes this design problem and has the form of the trace of the inverse of a sum of two Kronecker products. We derive a new general formula for the inverse of the sum of two Kronecker products, and we use this result to rewrite the Kronecker-Bayesian criterion in the form of the compound Bayes risk criterion, which can be recognized as a sum of Bayesian linear criteria with the same moment matrix. Based on the convexity and differentiability of the Kronecker-Bayesian linear criterion, we establish optimality conditions for approximate designs. We also propose a dimension reduction approach that provides highly efficient approximations for optimal designs. The proposed method allows for the preselection of an upper bound on the efficiency loss, which is independent of the true optimal design. Optimal or highly efficient designs can be computed under any kind of additional linear constraints, such as cost constraints. We apply our results to the problem of optimizing the allocation of trials in multi-environment crop variety testing, and we illustrate the behavior of the optimal designs by real data examples. Finally, we consider further applications of the general formula for computing the inverse of the sum of two Kronecker products in control theory or multivariate time series analysis.
In the era of big data, data privacy has attracted increasing attention. Differential privacy is a state-of-the-art framework for formal privacy guarantees. Many privacy-preserving inference methods have been developed for releasing information from a wide range of data analyses in the differential privacy framework. However, differential privacy statistical inference methods for streaming data, which represent a common type of big data, are still lacking. In this paper, we propose a computationally efficient privacy-preserving method for online updating and inference of linear regression models that is differentially private. We derive regression parameter estimates in the differential privacy framework, along with the covariance estimates based on which privacy-preserving confidence intervals for the parameters are constructed. We provide theoretical support for the proposed differentially private method, and numerical results demonstrate the good performance of our approach.
Many incomplete-data statistical inference procedures are developed under the missing at random (MAR) assumption. However, the MAR assumption has been criticized as being overly strong for real-data problems, and is unverifiable by using observed data. To handle data that are missing not at random (MNAR), sensitivity analysis has been proposed to investigate how conclusions are perturbed if the unverifiable MAR assumption is violated to a certain degree. This article proposes a new framework called multiple sensitivity models (MSMs) for performing general parameter estimation with the generalized estimating equation (GEE) method. Given user-specified sensitivity parameters, a range of estimators is derived by solving the roots of the bounds of MSM-assisted GEEs. Furthermore, we derive a representation for the proposed estimator so that it can be decomposed into several simpler estimators. It allows us to investigate the impact of different missing patterns. An asymptotically valid percentile bootstrap confidence region (CR) is also proposed. Theoretical justification is provided together with empirical evidence, which verifies the usefulness of the proposal's sensitivity analysis.
Nonresponse in probability sampling presents a long-standing challenge in survey sampling, often necessitating simultaneous adjustments to address sampling and selection biases. We develop a statistical framework that explicitly models sampling weights as random variables and establish the semiparametric efficiency bound for the parameter of interest under nonresponse. This study investigates strategies for eliminating bias and effectively utilizing available information, extending beyond nonresponse issues to data integration with external summary statistics. The proposed estimators are characterized by their efficiency and double robustness. However, realizing full efficiency hinges on the accurate specification of underlying models. To enhance robustness against potential model misspecification, we expand double robustness to multiple robustness through a novel two-step empirical likelihood approach. A numerical study evaluates the finite-sample performance of our methods. Additionally, we apply these methods to a dataset from the National Health and Nutrition Examination Survey, effectively integrating summary statistics from the National Health Interview Survey.
In high dimensions, the relationship between covariates and a response variable becomes increasingly intricate, with different covariate components often displaying varying degrees of variability. This complex interplay of dependence, heterogeneity, and high dimensionality presents a significant challenge when investigating the effects of covariates on the response variable. To address this, we propose a novel marginal testing procedure based on kernel-based conditional mean dependence, which can be implemented without requiring model assumptions. Theoretically, we establish the limiting normal distributions of the test statistic under both null hypotheses and local alternatives by asymptotically approximating a class of quadratic forms. We also examine the asymptotic relative efficiency of the proposed test against several state-of-the-art alternatives. Our theoretical evaluations are conducted from two perspectives: a detailed analysis within a linear model framework and a comparison within a fully non-parametric setting. The effectiveness and applicability of the proposed method are demonstrated through both simulation studies and real data analysis.
Although ubiquitous in many areas, missing data problems become challenging when the missingness of the outcome depends on itself, which means the data are non-ignorable missing. To alleviate the risk of model misspecification and balance interpretation and efficiency, we model the non-missingness probability by a logistic model with a general additive covariate effect under shape restrictions. Each additive component is assumed to satisfy certain shape restrictions, such as monotone increasing/decreasing, convexity/concavity, or a combination of these. We develop a shape-restricted and tuning-parameter-free estimator for the population outcome mean with the help of an instrument variable. We systematically establish the consistency, convergence rates, and asymptotic normalities of the proposed estimators. Our numerical results indicate that the proposed shape-restricted estimator has comparable performance to competing estimators with parametric models when the parametric models are correct, and outperforms them when the parametric models are misspecified. Finally, our method is applied to two real datasets providing more interpretable results than its competitors.