
This work discusses a theory of functional spaces over metric graphs that permits the definition of penalized likelihood methods for data observed over spatial supports that are graphs. Within the considered mathematical framework, we recover classical results in functional analysis, such as a Poincar & eacute;-type inequality. This, in turn, enables us to uplift, to the present setting, the theory of some fundamental penalized likelihood methods. Specifically, we present two important classes of statistical models: nonparametric regression and nonparametric density estimation, here defined for data observed over graphs. We derive theoretical results regarding the well-posedness of the associated estimation problems and the consistency of the estimators. We also demonstrate the performance of the defined estimators with respect to state-of-the-art alternatives.
Summary This paper addresses a long-standing open problem in crossed random effect models under unbalanced designs: how to find an analytic expression for the inverse of V, the covariance matrix of the observed response. For unbalanced crossed designs, V is dense and the lack of a closed-form representation for V−1, until now, has made using likelihood-based methods computationally challenging and difficult to analyse mathematically. We use the Khatri–Rao product to represent V and then construct a modified covariance matrix whose inverse admits an exact spectral decomposition. Building on this construction, we obtain an elegant and simple approximation to V−1 for asymptotic unbalanced designs. For non-asymptotic settings, we derive an accurate and interpretable approximation under mildly unbalanced data and establish an exact inverse representation as a low-rank correction to this approximation, applicable to arbitrary degrees of unbalance. Simulations demonstrate the framework’s accuracy, stability, and tractability.
Summary Learning causal relationships between pairs of complex traits from observational studies is of great interest in many scientific fields. However, most existing methods assume the absence of unmeasured confounding and restrict causal relationships between two traits to be unidirectional, assumptions that may be violated in real-world systems. In this paper, we address the problem of bivariate causal discovery in the presence of unmeasured confounding and potential feedback loops, leveraging possibly invalid instrumental variables. We propose a novel doubly robust identification framework that guarantees identifiability of the causal direction as long as one of two complementary assumptions holds, without requiring knowledge of which assumption is correct. Moreover, the causal effects are point identified in the unidirectional case and identifiable up to two candidate solutions in the bidirectional case, unless additional information is available. Building on this framework, we develop a finite-sample procedure to detect the causal direction and conduct inference on causal effects. We prove that our method consistently recovers the true causal direction and yields valid confidence intervals for causal effects. Extensive simulations demonstrate its superior performance compared to existing methods. We finally apply our approach to analyse real datasets from the UK Biobank.
Summary We present machine learning estimators for causal and predictive parameters under covariate shift, where covariate distributions differ between training and target populations. One such parameter is the average effect of a policy that alters the covariate distribution, such as a treatment that modifies surrogate covariates used to predict long-term outcomes. Another example is the average treatment effect for a population with a shifted covariate distribution. We propose a debiased machine learning method to estimate a broad class of these parameters in a statistically reliable and automatic manner. Our method eliminates regularization bias arising from the use of machine learning tools in high-dimensional settings, and relies solely on the parameter’s defining formula. It employs data fusion by combining samples of target and training data to eliminate bias. We give asymptotic theory that allows the sample sizes of the training and target datato grow at different rates. Computational experiments and an empirical study of the impact ofminimum-wage increases on teen employment, using the difference-in-differences framework with unconfoundedness, demonstrate the effectiveness of our method.
Summary We consider the problem of identifying the pattern of latent variables in high-dimensional linear latent variable models, which can also be interpreted as determining the source of spiked singular values in the data matrix. Specifically, we test whether the latent variables are continuous or categorical, a distinction which is crucial for data interpretation but challenging in the high-dimensional regime. To address this inference problem, we analyze the asymptotic behavior of empirical measures associated with singular vectors corresponding to large spiked singular values. Leveraging these insights,we propose novel test statistics based on the eigenvector quantile differences and establish their theoretical performance under the null hypothesis. Simulation studies and real data analyses demonstrate the effectiveness and practical utility of our method.
The average treatment effect is commonly used to quantify the main effect of a binary treatment on an outcome. Extensions to continuous treatments are usually based on the dose-response curve or shift interventions, but both require strong overlap conditions, and the resulting curves may be difficult to summarize. This article focuses instead on average derivative effects, which are scalar estimands related to infinitesimal shift interventions requiring only local overlap assumptions. Average derivative effects, however, are rarely used in practice because their estimation usually requires estimating conditional density functions. By characterizing the Riesz representers of weighted average derivative effects, we propose a new class of estimands that provides a unified view of weighted average derivative (respectively, treatment) effects when the treatment is continuous (respectively, binary). We derive the estimand in our class that minimizes the nonparametric efficiency bound, thereby extending optimal weighting results from the binary treatment literature to the continuous setting. We develop efficient estimators for two weighted average derivative effects that avoid density estimation and are amenable to modern machine learning methods, and we evaluate their performance in simulations and an applied analysis of Warfarin dosage effects.
This article concerns the estimation and inference of treatment effects in panel data settings when treatments change dynamically over time. We propose a balancing method that allows for (i) treatments to be assigned dynamically over time based on high-dimensional covariates, past outcomes and treatments; (ii) outcomes and time-varying covariates to depend on the trajectory of all past treatments; and (iii) heterogeneity of treatment effects. Our approach recursively projects potential outcomes' expectations on past histories. It then controls the bias arising from the nonexperimental and sequential nature of this setting by balancing dynamically observable characteristics over time. We establish inferential guarantees for the proposed method even in cases where the number of observable characteristics greatly exceeds the sample size. We study numerical properties of the estimator and illustrate the advantages of the procedure in an empirical application.
We consider the problem of estimating ratios of means of a multivariate outcome across covariates when the data are observed with unknown sample-specific and category-specific perturbations. Our model admits a partially identifiable estimand, and we establish full identifiability by imposing interpretable parameter constraints. To reduce bias and guarantee the existence of estimators in the presence of sparse observations, we apply an asymptotically negligible and constraint-invariant penalty to the loss function. We develop a fast coordinate-descent algorithm for estimation, and an augmented Lagrangian algorithm for estimation under null hypotheses. We construct a model-robust score test and demonstrate valid inference even for small sample sizes and under violated distributional assumptions. The flexibility of the approach, and comparisons with related methods, are illustrated through a simulation study and a meta-analysis of microbial associations with colorectal cancer.
Marginal structural models are a popular method for estimating causal effects in the presence of time-varying exposures. In spite of their popularity, no scalable nonparametric estimator exists for marginal structural models with multi-valued or continuous time-varying treatments. In this paper, we combine flexible, data-adaptive regression methods, including ensemble learning techniques, with recent developments in semiparametric efficiency theory for longitudinal studies to propose such an estimator. The proposed estimator is based on a study of the nonparametric identifying functional, including first-order von Mises expansions, as well as the efficient influence function and the efficiency bound. We show conditions under which the proposed estimators are efficient, asymptotically normal and sequentially doubly robust. We perform a simulation study to illustrate the properties of the estimators, and present the results of our motivating study on a COVID-19 dataset, studying the impact of mobility on the cumulative number of observed cases.
Comparing outcomes across treatments is essential in medicine and public policy. To do so, researchers typically estimate a set of parameters, possibly counterfactual, each targeting adifferent treatment. Treatment-specific means are commonly used, but their identification requires a positivity assumption: every subject has a nonzero probability of receiving each treatment. This assumption is often implausible, especially when treatment can take many values. Causal parameters based on dynamic stochastic interventions offer robustness to positivity violations. However, comparing these parameters may fail to reflect the effects of the underlying target treatments because the parameters can depend on outcomes under nontarget treatments. To clarify when two parameters targeting different treatments yield a useful comparison of treatment efficacy, we propose a comparability criterion: if the conditional treatment-specific mean for one treatment is greater than that for another, then the corresponding causal parameter should also be greater. Many standard parameters fail to satisfy this criterion, but we show that only a mild positivity assumption is needed to identify parameters that yield useful comparisons. We then provide two simple examples that satisfy this criterion and are identifiable under the milder positivity assumption: trimmed and smooth-trimmed treatment-specific means with multivalued treatments. For smooth-trimmed treatment-specific means, we develop doubly robust-style estimators that attain parametric convergence rates under nonparametric conditions. We illustrate our methods with an analysis of dialysis providers in New York State.
The martingale posterior framework replaces the elicitation of the likelihood and prior with that of a sequence of one-step-ahead predictive densities for Bayesian inference. Posterior sampling then involves the imputation of unobserved quantities and can then be carried out in an expedient and parallelizable manner using predictive resampling, without requiring Markov chain Monte Carlo. Recent work has investigated the use of plug-in parametric predictive densities, combined with stochastic gradient descent, to specify a parametric martingale posterior. This paper investigates the asymptotic properties of this class of parametric martingale posteriors. In particular, two central limit theorems based on martingale limit theory are introduced and applied. The first is a predictive central limit theorem, which enables a significant acceleration of the predictive resampling scheme through a hybrid sampling algorithm based on a normal approximation. The second is a Bernstein-von Mises result, which is novel for martingale posteriors, and provides methodological guidance on attaining desirable frequentist properties. We demonstrate the utility of the theoretical results through simulations and a real data example.
It is often of interest to infer lower-dimensional structure underlying complex data. As a flexible class of nonlinear structures, it is common to focus on Riemannian manifolds. Most existing manifold-learning algorithms replace the original data with lower-dimensional coordinates without providing an estimate of the manifold or using it to denoise the original data. This article proposes a new methodology to address these issues, allowing interpolation of the estimated manifold between the fitted data points. The proposed approach is motivated by the novel theoretical properties of local covariance matrices constructed from samples near a manifold. Our results enable the transformation of a global manifold-reconstruction problem into a local regression problem, allowing the application of Gaussian processes for probabilistic manifold reconstruction. In addition to the theory justifying our methodology, we provide simulated and real data examples to illustrate its performance.
We introduce a new class of conditional autoregressive models for spatially dependent functional data, formulated through conditional means given neighboring functional observations and characterized by a covariance operator and a spatial dependence parameter. Our estimation strategy consists of three components: (i) estimating the covariance operator using conditionally centered data, (ii) estimating the spatial dependence parameter by maximizing the likelihood of projected observations, and (iii) applying a novel profile-based approach to obtain the final estimators. Under an expanding lattice framework, we establish two key theoretical results. First, we establish the consistency of the proposed covariance estimator, which is not attainable using naive methods based on marginally centered data. Second, we prove that the spatial dependence parameter estimator is superconsistent and asymptotically normal, where the latter property enables statistical inference for spatial dependence in functional data – a contribution that is novel in the existing literature. Numerical studies support the theoretical results and demonstrate the computational efficiency of our method. Finally, we illustrate its practical utility by analyzing weekly PM_2.5 concentration trajectories in 2019 across counties in the Midwestern United States.
Motivated by the challenge of estimating effects of DNA methylation on 3D genomic contacts captured by multimodal single-cell Hi-C data, we consider the tensor-response partial least squares model with $ \mathcal{Y}=\mathcal{B}\times_{1}X+\mathcal{F} $, where the correlated high-dimensional predictors $ X\in\mathbb{R}<^>{n\times d_{1}} $ and the sparse and noisy high-dimensional responses $ \mathcal{Y}\in\mathbb{R}<^>{n\times\prod_{m}d_{m}} $ are observed, the low-rank and sparse partial least squares coefficient tensor $ \mathcal{B}\in\mathbb{R}<^>{\prod_{m}d_{m}} $ is unknown, and $ \mathcal{F}\in\mathbb{R}<^>{n\times\prod_{m}d_{m}} $ is the noise tensor. In this work, we study the problem of estimating the partial least squares coefficient $ \mathcal{B} $ and identifying its active entries in the tensor partial least squares framework. We show that the consistency of the existing tensor partial least squares estimator (Zhao et al., 2012) cannot be guaranteed under a high-dimensional regime in which both the number of predictors and the response tensor dimensions grow faster than the sample size. To address this high dimensionality, we propose the sparse higher-order partial least squares estimator and an accompanying algorithm that can simultaneously perform variable selection, dimension reduction and tensor-response denoising. We establish asymptotic guarantees for the proposed estimator in the high-dimensional regime and validate these results through comprehensive simulation studies, demonstrating our method's advantages over baseline approaches. Finally, application of the proposed estimator to the motivating multimodal single-cell Hi-C data provides novel biological insights into gene regulation by multiple regulatory elements.
We consider the problem of estimating the number of significant components in high-dimensional principal component analysis. We propose a new penalized approach using the explained variance ratio and the rigidity of the nonspiked sample eigenvalues of sample covariance matrices of $ p $ variables. Compared with methods in the existing literature, the consistency of the proposed estimator holds, not only for independent data, but also for some times series data when the dimension $ p $ and the sample size $ n $ both tend to infinity. Even for independent data our estimator works under weaker conditions than existing approaches such as the aic and bic, including allowing heterogeneity in the bulk of the population eigenvalues. Simulation studies are conducted to illustrate the performance of the proposed estimator.
This paper introduces a novel extension of Fr & eacute;chet means, referred to as generalized Fr & eacute;chet means, as a comprehensive framework for describing the characteristics of random elements. The generalized Fr & eacute;chet mean is defined as the minimizer of a cost function, and the framework encompasses various extensions of Fr & eacute;chet means that have appeared in the literature. The most distinctive feature of the proposed framework is that it allows the domain of minimization for the empirical generalized Fr & eacute;chet means to be random and different from that of its population counterpart. This flexibility broadens the applicability of the Fr & eacute;chet mean framework to various statistical scenarios, including sequential dimension reduction for non-Euclidean data. We establish a strong consistency theorem for generalized Fr & eacute;chet means and demonstrate the utility of the proposed framework by verifying the consistency of principal geodesic analysis on the hypersphere.
The estimation of regression parameters in spatially referenced data plays a crucial role across various scientific domains. A common approach involves employing an additive regression model to capture the relationship between observations and covariates, accounting for spatial variability not explained by the covariates through a Gaussian random field. We study the effect of misspecified covariates, in particular when the misspecification changes the smoothness. We analyse the theoretical properties of the generalized least-squares estimator under infill asymptotics, and show that the estimator can have counter-intuitive properties. In particular, the estimated regression coefficients can converge to zero as the number of observations increases if the covariates are too rough, despite high correlations between observations and covariates. This has important implications for practical applications as the importance of rough covariates can be severely underestimated, leading to incorrect scientific conclusions. We also show that the estimates can diverge to infinity under certain conditions, which can also lead to incorrect conclusions in practical applications. Through an application to temperature and precipitation data, we show that both behaviours can be observed for real data. Finally, we propose adding a smoothing step in the regression and show both theoretically and practically that this can solve the problem.
A new formulation of the additive interaction model is introduced. In contrast to existing approaches, the new formulation separates well the joint effects of covariates that cannot be accounted for by individual main effects. The new approach enables correct interpretation of interaction effects by making them orthogonal to the associated main effects in the $ L<^>{2} $ sense. A new method is developed to estimate the resulting main and interaction effects. Asymptotic $ L<^>{2} $ error rates are derived for the estimators under mild technical conditions. Numerical evidence is provided via simulation studies and real-data examples.
Factorial experiments are ubiquitous in the social and biomedical sciences, but when units fail to comply with each assigned factor, identification and estimation of the average treatment effects become impossible without strong assumptions. Leveraging an instrumental variables approach, previous studies have shown how to define and estimate the causal effect of treatment uptake among respondents who comply with treatment. A major caveat is that these results rely on strong assumptions on the effect of randomization on treatment uptake. This work shows how to bound these complier average treatment effects for bounded outcomes under milder assumptions on noncompliance.
We propose a method for comparing survival data based on higher criticism of $ p $-values obtained from many exact hypergeometric tests. The method accommodates noninformative right-censorship and is sensitive to hazard differences in unknown and relatively rare time intervals. It attains much better power against such differences than the log-rank test and its variants. We demonstrate the usefulness of our method in detecting rare and weak non-proportional hazard differences compared to existing tests using simulations and gene expression data. Additionally, we analyse the asymptotic power of our method and other tests under a theoretical framework describing two groups experiencing failure rates that are usually identical over time, except in a few unknown instances where one group's failure rate is higher. Our test's power experiences a phase transition across the plane of rarity and intensity parameters that mirrors the phase transition of higher criticism in two-sample rare and weak normal and Poisson means settings. The region of the plane in which our method has asymptotically full power is larger than the corresponding region for the log-rank test.