Tensor regression has attracted significant attention in statistical research. This study tackles the challenge of handling covariates with smooth varying structures. We introduce a novel framework, termed functional tensor regression, which incorporates both the tensor and functional aspects of the covariate. To address the high dimensionality and functional continuity of the regression coefficient, we employ a low Tucker rank decomposition along with smooth regularization for the functional mode. We develop a functional Riemannian Gauss--Newton algorithm that demonstrates a provable quadratic convergence rate, while the estimation error bound is based on the tensor covariate dimension. Simulations and a neuroimaging analysis illustrate the finite sample performance of the proposed method.
We propose a new perspective to conduct robust functional data analysis of discretely observed functional data ranging from sparse to dense sampling designs. This analysis caters to processes with various distributions, including heavy-tailed, skewed, or contaminated distributions. We study the robust functional mean (M-location) and introduce a robust dimension reduction method via principal component analysis. Theoretical outcomes for the robust functional mean and eigenfunction estimates, derived from pooled discretely observed data, are elucidated, matching their non-robust counterparts. The established convergence rates for these estimated eigenfunctions, with indices increasing with sample size, pave the way for further modeling and analysis.
Structural changes often arise in real-world dynamic systems due to external interventions or environmental shifts, such as policy changes in epidemiology or climate forcing in environmental science. In this paper, we propose a unified framework for detecting and localizing structural changes in dynamic systems governed by ordinary differential equations. Unlike existing methods that assume mean or linear trend changes, our approach accommodates complex, nonlinear dynamics with both stable and diverging trajectories. We develop a new test statistic that combines residual-based discrepancy and normalized parameter contrast, capturing evidence for structural changes from both model fit and parameter shifts. Candidate structural changes are efficiently screened using a multiscale seeded-narrowest-over-threshold algorithm with a data-driven thresholding strategy. To refine selections and control false discoveries, we introduce a false discovery rate control procedure that leverages order-preserved sample splitting and symmetric contrast calibration. Theoretical guarantees are established, including detection consistency, near-minimax localization accuracy, and valid FDR control under weak dependence. Extensive simulations demonstrate superior performance over existing methods in both accuracy and FDR control. Applications to real-world data sets, including COVID-19 dynamics and global temperature trends, highlight the practical relevance and broad applicability of our method.
This article addresses Bayesian inference related to partial differential equations (PDEs), particularly nonparametric regression constrained by PDEs. To effectively encode prior information, we propose a novel framework that learns a prediction function of the prior distribution from historical training datasets. We introduce hyper-prior and hyper-posterior distributions and derive a generalization error estimate, which accommodates data-dependent priors by extending the concept of differential privacy. Some mild conditions are given to validate the error estimate, where various typical PDEs such as diffusion and Darcy flow equations can be integrated. We thus formulate an infinite-dimensional optimization problem to obtain the point estimate of the hyper-posterior. Numerical examples demonstrate the performance of our proposed method in learning the prediction function of priors. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.
Unmeasured confounders are a major source of bias in regression-based effect estimation and causal inference. In this paper, we advocate a new profiled transfer learning framework, ProTrans, to address confounding effects in the target dataset, when additional source datasets that possess similar confounding structures are available. We introduce the concept of profiled residuals to characterize the shared confounding patterns between source and target datasets. By incorporating these profiled residuals into the target debiasing step, we effectively mitigates the latent confounding effects. We also propose a source selection strategy to enhance robustness of ProTrans against noninformative sources. As a byproduct, ProTrans can also be utilized to estimate treatment effects when potential confounders exist, without the use of auxiliary features such as instrumental or proxy variables, which are often challenging to select in practice. Theoretically, we prove that the resulting estimated model shift from sources to target is confounding-free without any assumptions imposed on the true confounding structure, and that the target parameter estimation achieves the minimax optimal rate under mild conditions. Simulated and real-world experiments validate the effectiveness of ProTrans and support the theoretical findings.
We develop a framework of canonical correlation analysis for distribution-valued functional data within the geometry of Wasserstein spaces.Specifically, we formulate an intrinsic concept of correlation between random distributions, propose estimation methods based on functional principal component analysis and Tikhonov regularization, respectively, for the correlation and its corresponding weight functions, and establish the minimax convergence rates of the estimators.In order to overcome the challenge raised by nonlinearity of Wasserstein spaces, the key idea is to adopt tensor Hilbert spaces to distribution-valued functional data.The finite-sample performance of the proposed estimators is illustrated via simulation studies, and the practical merit is demonstrated via a study on the association of distributions of brain activities between brain regions.
Online high-dimensional regression has gained increasing attention in recent years, yet existing methods typically assume that all candidate features, including important ones, are observed from the outset of data collection. This assumption is often violated in real-world scenarios, where new variables become available gradually as data accumulate. To address this gap, we introduce a novel framework, Recurrent Adaptive Variable Selection (RAVAS), for online regression with expanding observability. RAVAS employs a recurrent procedure that dynamically updates feature selection as both the sample size and the observable feature set grow. The algorithm is designed to be computationally efficient and memory-light, relying only on low-dimensional sufficient statistics that are updated online. A key advantage of the method lies in its ability to detect and incorporate important variables that emerge later, thereby mitigating the effect of early-stage missingness. We establish theoretical guarantees on model selection, estimation error, and feature coverage, and develop an adaptive online tuning strategy. Extensive simulations and real-world experiments verify the effectiveness of RAVAS for high-dimensional streaming data.
Frechet random forests extend the power of classical random forests to general metric spaces, offering promising advantages over traditional methods, especially in high-dimensional settings. While forests have empirically outperformed individual trees in Frechet regression, the theoretical basis for this improvement remains largely unexplored. This paper fills this gap by establishing non-asymptotic upper bounds for the prediction risk of Frechet Mondrian forests, complementing the existing literature that primarily focuses on asymptotic analysis. We demonstrate that, under suitable regularity conditions, Frechet Mondrian forests attain convergence rates comparable to their Euclidean counterparts. Moreover, under higher-order smoothness assumptions and with a sufficient number of trees, Frechet forests achieve faster convergence than individual Frechet trees, thereby providing a rigorous theoretical justification for the benefit of ensembles in non-Euclidean regression problems. The effectiveness of the proposed method is further corroborated through simulation studies across diverse settings, including probability distributions, symmetric positive-definite matrices, and spherical data.
Functional data analysis is an important statistical field that treats data as random functions. In practice, the random functions are often not fully observed but instead measured at discrete times. While simpler problems, such as mean and covariance estimation, have been widely studied for discretely observed data, optimal estimation of linear regression for this data type has remained unsolved for over two decades. To tackle this fundamental challenge, we propose a novel approach, referred to as pooling ridge estimation, which combines the advantages of pooling strategy and RKHS-based method by incorporating the unbiased estimation of operators based on discretely observed measurements from all subjects. This unified estimation framework enables us to achieve minimax optimality in prediction risk in arbitrary sampling schemes ranging from sparse to dense designs, for both scalar-on-function and function-on-function regression models. Such methodological and theoretical advances are obtained for the first time and accurately reveal the influence of discrete sampling. For scalar-on-function regression, the phase transition occurs once, separating the convergence behavior into two distinct regimes. Remarkably, for function-on-function regression, up to three phase transitions may occur, determined by the sampling frequencies of the predictor/response functions. Finally, simulation experiments and two real data examples provide empirical support for the proposed methods.
The exploration of dynamic systems governed by ordinary differential equations (ODEs) holds great interest in the field of statistics. Existing research mainly focuses on a single function. This study generalizes the scope to analyse a collection of functions observed at discretized times, with sampling frequencies varying from sparse to dense designs. The range of ODE models studied caters to diverse dynamic systems, and includes the complex nonlinear and non-Lipschitz scenarios. We introduce a new concept named functional moment method, a novel approach for parameter estimation within these ODE models and facilitating the recovery of curves for the discretely observed functions. Our numerical analysis underscores the method’s applicability across various application fields, including sociology, physics, and epidemiology.
This article studies the benefits of using spatially randomized experimental designs which partition the experimental area into distinct, non-overlapping units with treatments assigned randomly. Such designs offer improved policy evaluation in online experiments by providing more precise policy value estimators and more effective A/B testing algorithms than traditional global designs, which apply the same treatment across all units simultaneously. We examine both parametric and nonparametric methods for estimating and inferring policy values based on this randomized approach. Our analysis includes evaluating the mean squared error of the treatment effect estimator and the statistical power of the associated tests. Additionally, we extend our findings to experiments with spatio-temporal dependencies, where treatments are allocated sequentially over time, and account for potential temporal carryover effects. Our theoretical insights are supported by comprehensive numerical experiments.
Modeling interference effects in high-dimensional settings presents significant challenges due to the complexity of underlying dependence structures. Existing approaches often rely on explicit and homogeneous assumptions, limiting their applicability in real-world scenarios. In this paper, we introduce a novel low-rank and sparse treatment effect model that leverages high-dimensional techniques to identify interference structures without restrictive parametric assumptions. We propose an efficient profiling algorithm for estimating model coefficients and develop statistical methodologies for both global testing of interference existence and local detection of interference-affected units. Theoretical guarantees are established, including non-asymptotic error bounds for estimation and accuracy guarantees for detection based on the Jaccard index. Through numerical experiments, we demonstrate the effectiveness of our method in high-dimensional settings, highlighting its advantages in capturing complex interference patterns.
In many scientific fields, the generation and evolution of data are governed by partial differential equations (PDEs) which are typically informed by established physical laws at the macroscopic level to describe general and predictable dynamics. However, some complex influences may not be fully captured by these laws at the microscopic level due to limited scientific understanding. This work proposes a unified framework to model, estimate, and infer the mechanisms underlying data dynamics. We introduce a general semiparametric PDE (SemiPDE) model that combines interpretable mechanisms based on physical laws with flexible data-driven components to account for unknown effects. The physical mechanisms enhance the SemiPDE model's stability and interpretability, while the data-driven components improve adaptivity to complex real-world scenarios. A deep profiling M-estimation approach is proposed to decouple the solutions of PDEs in the estimation procedure, leveraging both the accuracy of numerical methods for solving PDEs and the expressive power of neural networks. For the first time, we establish a semiparametric inference method and theory for deep M-estimation, considering both training dynamics and complex PDE models. We analyze how the PDE structure affects the convergence rate of the nonparametric estimator, and consequently, the parametric efficiency and inference procedure enable the identification of interpretable mechanisms governing data dynamics. Simulated and real-world examples demonstrate the effectiveness of the proposed methodology and support the theoretical findings.
Nonparametric mean function regression with repeated measurements serves as a cornerstone for many statistical branches, such as longitudinal/panel/functional data analysis. In this work, we investigate this problem using fully connected deep neural network (DNN) estimators with flexible shapes. A novel theoretical framework allowing arbitrary sampling frequency is established by adopting empirical process techniques to tackle clustered dependence. We then consider the DNN estimators for Holder target function and illustrate a key phenomenon, the phase transition in the convergence rate, inherent to repeated measurements and its connection to the curse of dimensionality. Furthermore, we study several examples with low intrinsic dimensions, including the hierarchical composition model, low-dimensional support set and anisotropic Holder smoothness. We also obtain new approximation results and matching lower bounds to demonstrate the adaptivity of the DNN estimators for circumventing the curse of dimensionality. Simulations and real data examples are provided to support our theoretical findings and practical implications. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.
Functional data analysis is an important research field in statistics which treats data as random functions drawn from some infinite-dimensional functional space, and functional principal component analysis (FPCA) based on eigen-decomposition plays a central role for data reduction and representation. After nearly three decades of research, there remains a key problem unsolved, namely, the perturbation analysis of covariance operator for diverging number of eigencomponents obtained from noisy and discretely observed data. This is fundamental for studying models and methods based on FPCA, while there has not been substantial progress since Hall, M & uuml;ller and Wang (Ann. Statist. 34 (2006) 1493-1517)'s result for a fixed number of eigenfunction estimates. In this work, we aim to establish a unified theory for this problem, obtaining upper bounds for eigenfunctions with diverging indices in both the 2 pound and supremum norms, and deriving the asymptotic distributions of eigenvalues for a wide range of sampling schemes. Our results provide insight into the phenomenon when the 2 pound bound of eigenfunction estimates with diverging indices is minimax optimal as if the curves are fully observed, and reveal the transition of convergence rates from nonparametric to parametric regimes in connection to sparse or dense sampling. We also develop a double truncation technique to handle the uniform convergence of estimated covariance and eigenfunctions. The technical arguments in this work are useful for handling the perturbation series with noisy and discretely observed functional data and can be applied in models or those involving inverse problems based on FPCA as regularization, such as functional linear regression.
The identification of genetic signal regions in the human genome is critical for understanding the genetic architecture of complex traits and diseases. Numerous methods based on scan algorithms (i.e. QSCAN, SCANG, SCANG-STARR) have been developed to allow dynamic window sizes in whole-genome association studies. Beyond scan algorithms, we have recently developed the binary and re-search (BiRS) algorithm, which is more computationally efficient than scan-based methods and exhibits superior statistical power. However, the BiRS algorithm is based on two-sample mean test for binary traits, not accounting for multidimensional covariates or handling test statistics for non-binary outcomes. In this work, we present a distributed version of the BiRS algorithm (dBiRS) that incorporate a new infinity-norm test statistic based on summary statistics computed from a generalized linear model. The dBiRS algorithm accommodates regression-based statistics, allowing for the adjustment of covariates and the testing of both continuous and binary outcomes. This new framework enables parallel computing of block-wise results by aggregation through a central machine to ensure both detection accuracy and computational efficiency, and has theoretical guarantees for controlling family-wise error rates and false discovery rates while maintaining the power advantages of the original algorithm. Applying dBiRS to detect genetic regions associated with fluid intelligence and prospective memory using whole-exome sequencing data from the UK Biobank, we validate previous findings and identify numerous novel rare variants near newly implicated genes. These discoveries offer valuable insights into the genetic basis of cognitive performance and neurodegenerative disorders, highlighting the potential of dBiRS as a scalable and powerful tool for whole-genome signal region detection.
Principal component analysis has been one of prominent techniques in multi-dimensional statistics, and the core is to estimate the eigenvalues in descending ordering with their corresponding eigenvectors. However, when data are collected longitudinally in many scientific applications, the eigenvalues become dynamic over time, and the ordering of them may not be well-defined, for instance, does not coincide with the pointwise ordering. To deal with this issue, we propose a new framework, namely the dynamic principal component analysis. This addresses the identifiability of principal components from a global perspective, and transforms the problem into a regression model for data situated on the orthogonal matrix group space. The one-step unrolling method is exploited to solve the regression problem with a suitably constructed regular base curve. The minimax rate of the proposed estimators is established through theoretical analysis of the one-step unrolling and its connection to smoothing spline in the context of manifold-valued data.
Research on the localization of the genetic basis associated with diseases or traits has been widely conducted in the last a few decades. Scan methods have been developed for region-based analysis in whole-genome association studies, helping us better understand how genetics influences human diseases or traits, especially when the aggregated effects of multiple causal variants are present. In this paper, we propose a fast and effective algorithm coupling with high-dimensional test for simultaneously detecting multiple signal regions, which is distinct from existing methods using scan or knockoff statistics. The idea is to conduct binary splitting with re-search and arrangement based on a sequence of dynamic critical values to increase detection accuracy and reduce computation. Theoretical and empirical studies demonstrate that our approach enjoys favorable theoretical guarantees with fewer restrictions and exhibits superior numerical performance with faster computation. Utilizing the UK Biobank data to identify the genetic regions related to breast cancer, we confirm previous findings and meanwhile, identify a number of new regions which suggest strong association with risk of breast cancer and deserve further investigation.
Spectral analysis plays a crucial role in high-dimensional statistics, where determining the asymptotic distribution of various spectral statistics remains a challenging task. Due to the difficulties of deriving the analytic form, recent advances have explored data-driven bootstrap methods for this purpose. However, widely used Gaussian approximation-based bootstrap methods, such as the empirical bootstrap and multiplier bootstrap, have been shown to be inconsistent in approximating the distributions of spectral statistics in high-dimensional settings. To address this issue, we propose a universal bootstrap procedure based on the concept of universality from random matrix theory. Our method consistently approximates a broad class of spectral statistics across both high- and ultra-high-dimensional regimes, accommodating scenarios where the dimension-to-sample-size ratio $p/n$ converges to a nonzero constant or diverges to infinity without requiring structural assumptions on the population covariance matrix, such as eigenvalue decay or low effective rank. We showcase this universal bootstrap method for high-dimensional covariance inference. Extensive simulations and a real-world data study support our findings, highlighting the favorable finite sample performance of the proposed universal bootstrap procedure.