Predicting outputs that are located in non-Euclidean spaces, such as probability distributions, networks, and symmetric positive-definite matrices, is becoming increasingly important in modern data analysis, particularly when inputs are high-dimensional. We propose DeSI (Deep Single-Index Fréchet Regression), a semiparametric framework for regression with metric space-valued outputs and multivariate inputs that assumes a single-index structure for the conditional Fréchet mean. DeSI estimates an interpretable index direction, which quantifies the relative importance of inputs, using a deep neural network, and performs Fréchet regression along the resulting one-dimensional index in the target metric space. This structure mitigates the curse of dimensionality while retaining interpretability, which stands in contrast to standard deep neural networks. We establish theoretical guarantees for DeSI, including consistency and convergence rates, and demonstrate its strong predictive performance through simulations on distributions, networks, and symmetric positive-definite matrices, as well as an application to compositional mood data from New Jersey.
Linear regression is widely used to model relationships between responses and predictors. In modern applications, one encounters data where the responses are non-Euclidean random objects situated in a metric space, paired with Euclidean predictors. Global Fréchet regression generalizes linear regression to such general settings, however statistical inference has remained largely unexplored. We develop a significance test for the null hypothesis that the Fréchet regression function does not depend on the predictors, addressing the challenge of an absence of linear operations in metric spaces. We also develop a test for the partial effect of a subset of the predictors in analogy to, but quite different from, the partial F-tests commonly used in classical linear regression under Gaussian assumptions. Key ideas are to employ random multipliers to obtain non-degenerate null distributions for the proposed test statistics and the Cauchy combination method. We obtain consistency and convergence results under the null hypothesis and contiguous alternatives and demonstrate the finite sample performance of the proposed tests through simulations on network data represented by graph Laplacians and spherical data with geodesic distances. We further illustrate our method using transport networks arising from New York City taxi trip data and U.S. energy source compositional data.
We develop a unified framework for testing independence and quantifying association between random objects that are located in general metric spaces. Special cases include functional and high-dimensional data as well as networks, covariance matrices and data on Riemannian manifolds, among other metric space-valued data. A key concept is the profile association, a measure based on distance profiles that intrinsically characterize the distributions of random objects in metric spaces. We rigorously establish a connection between the Hoeffding D statistic and the profile association and derive a permutation test with theoretical guarantees for consistency and power under alternatives to the null hypothesis of independence/no association. We extend this framework to the conditional setting, where the independence between random objects given a Euclidean predictor is of interest. In simulations across various metric spaces, the proposed profile independence test is found to outperform existing approaches. The practical utility of this framework is demonstrated with applications to brain connectivity networks derived from magnetic resonance imaging and age-at-death distributions for males and females obtained from human mortality data.
Classical regression models do not cover non-Euclidean data that reside in a general metric space, while the current literature on non-Euclidean regression by and large has focused on scenarios where either predictors or responses are random objects, i.e., non-Euclidean, but not both. In this paper we propose geodesic optimal transport regression models for the case where both predictors and responses lie in a common geodesic metric space and predictors may include not only one but also several random objects. This provides an extension of classical multiple regression to the case where both predictors and responses reside in non-Euclidean metric spaces, a scenario that has not been considered before. It is based on the concept of optimal geodesic transports, which we define as an extension of the notion of optimal transports in distribution spaces to more general geodesic metric spaces, where we characterize optimal transports as transports along geodesics. The proposed regression models cover the relation between non-Euclidean responses and vectors of non-Euclidean predictors in many spaces of practical statistical interest. These include one-dimensional distributions viewed as elements of the 2-Wasserstein space and multidimensional distributions with the Fisher-Rao metric that are represented as data on the Hilbert sphere. Also included are data on finite-dimensional Riemannian manifolds, with an emphasis on spheres, covering directional and compositional data, as well as data that consist of symmetric positive definite matrices. We illustrate the utility of geodesic optimal transport regression with data on summer temperature distributions and human mortality.
In this paper, we propose a new variable selection method for Nadaraya-Watson-Frechet regression, which is a special case of local Fr & eacute;chet regression. We demonstrate both the practical performance and the theoretical properties of this novel variable selection method. To provide theoretical support, we first derive a refined uniform convergence result for the Nadaraya-Watson-Frechet regression estimator of the Fr & eacute;chet regression function, which then enables us to establish asymptotic selection consistency for the proposed variable selection method. Excellent performance of the proposed variable selection method for Nadaraya-Watson-Frechet regression in finite sample situations is demonstrated both with simulation studies and in real data examples.
Sleep, a fundamental biological process, plays a crucial role in health and longevity across species. This study investigates the intricate relationship between sleep characteristics and lifespan in Mediterranean fruit flies (medflies), transcending the traditional focus on sleep duration. We identify specific sleep characteristics that are associated with remaining lifespan and also characterize how sleep patterns change with aging. Longer, uninterrupted sleep bouts are linked to extended lifespans, while increased nighttime activity, particularly in the first half of the night, correlates with reduced longevity. To our knowledge, this is the first study with continuous monitoring of night activity throughout the entire lifespan. It underscores the important role of sleep patterns beyond duration in shaping lifespan.
Difference-in-differences (DID) is a widely used quasi-experimental design for causal inference, traditionally applied to scalar or Euclidean outcomes, while extensions to outcomes residing in non-Euclidean spaces remain limited. Existing methods for such outcomes have primarily focused on univariate distributions, leveraging linear operations in the space of quantile functions, but these approaches cannot be directly extended to outcomes in general metric spaces. In this paper, we propose geodesic DID, a novel DID framework for outcomes in uniquely geodesic metric spaces that admit a geodesic transport structure, including distributions, networks, and manifold-valued data. To address the absence of algebraic operations in these spaces, we use geodesics as proxies for differences and introduce the geodesic average treatment effect on the treated (ATT) as the causal estimand. We establish the identification of the geodesic ATT and derive the convergence rate of its sample versions, employing tools from metric geometry and empirical process theory. This framework is further extended to the case of staggered DID settings, allowing for multiple time periods and varying treatment timings. To illustrate the practical utility of geodesic DID, we analyze health impacts of the Soviet Union's collapse using age-at-death distributions and assess effects of U.S. electricity market liberalization on electricity generation compositions.
Increasingly complex data analysis tasks motivate the study of the dependency of distributions of multivariate continuous random variables on scalar or vector predictors. Statistical regression models for distributional responses so far have primarily been investigated for the case of one-dimensional response distributions. We investigate here the case of multivariate response distributions while adopting the 2-Wasserstein metric in the distribution space. The challenge is that unlike the situation in the univariate case, the optimal transports that correspond to geodesics in the space of distributions with the 2-Wasserstein metric do not have an explicit representation for multivariate distributions. We show that under some regularity assumptions the conditional Wasserstein barycenters constructed for a geodesic in the Euclidean predictor space form a corresponding geodesic in the Wasserstein distribution space and demonstrate how the notion of conditional barycenters can be harnessed to interpolate as well as extrapolate multivariate distributions. The utility of distributional inter- and extrapolation is explored in simulations and examples. We study both global parametric-like and local smoothing-like models to implement conditional Wasserstein barycenters and establish asymptotic convergence properties for the corresponding estimates. For algorithmic implementation we make use of a Sinkhorn entropy-penalized algorithm. Conditional Wasserstein barycenters and distribution extrapolation are illustrated with applications in climate science and studies of aging.
Regression discontinuity designs have been widely used in observational studies to estimate causal effects of an intervention or treatment at a cutoff point. We propose a generalization of regression discontinuity designs to handle complex non-Euclidean outcomes, such as networks, compositional data, functional data, and other random objects residing in geodesic metric spaces. A key challenge in this setting is the absence of algebraic operations, which makes it difficult to define treatment effects using simple differences. To address this, we define the causal effect at the cutoff as a geodesic between the local Fréchet means of untreated and treated outcomes. This reduces to the classical average treatment effect in the scalar case. Estimation is carried out using local Fréchet regression, a nonparametric method for metric space-valued responses that generalizes local linear regression. We introduce a new bandwidth selection procedure tailored to regression discontinuity designs, which performs competitively even in classical scalar scenarios. The proposed geodesic regression discontinuity design method is supported by theory, including convergence rate guarantees, and is demonstrated in applications where causal inference is of interest in complex outcome spaces. These include changes in daily CO concentration curves following the introduction of the Taipei Metro, and shifts in UK voting patterns measured by vote share compositions after Conservative Party wins. We also develop an extension to fuzzy designs with non-Euclidean outcomes, broadening the scope of causal inference to settings that allow for imperfect compliance with the assignment rule.
We develop an inferential toolkit for analyzing object-valued responses, which correspond to data situated in general metric spaces, paired with Euclidean predictors within the conformal framework. To this end we introduce conditional profile average transport costs, where we compare distance profiles that correspond to one-dimensional distributions of probability mass falling into balls of increasing radius through the optimal transport cost when moving from one distance profile to another. The average transport cost to transport a given distance profile to all others is crucial for statistical inference in metric spaces and underpins the proposed conditional profile scores. A key feature of the proposed approach is to utilize the distribution of conditional profile average transport costs as conformity score for general metric space-valued responses, which facilitates the construction of prediction sets by the split conformal algorithm. We derive the uniform convergence rate of the proposed conformity score estimators and establish asymptotic conditional validity for the prediction sets. The finite sample performance for synthetic data in various metric spaces demonstrates that the proposed conditional profile score outperforms existing methods in terms of both coverage level and size of the resulting prediction sets, even in the special case of scalar Euclidean responses. We also demonstrate the practical utility of conditional profile scores for network data from New York taxi trips and for compositional data reflecting energy sourcing of U.S. states.
Exceedance refers to instances where a dynamic process surpasses given thresholds, e.g., the occurrence of a heat wave. We propose a novel exceedance framework for functional data, where each observed random trajectory is transformed into an exceedance function, which quantifies exceedance durations as a function of threshold levels. An inherent relationship between exceedance functions and probability distributions makes it possible to draw on distributional data analysis techniques such as Fréchet regression to study the dependence of exceedances on Euclidean predictors, e.g., calendar year when the exceedances are observed. We use local linear estimators to obtain exceedance functions from discretely observed functional data with noise and study the convergence of the proposed estimators. New concepts of interest include the force of centrality that quantifies the propensity of a system to revert to lower levels when a given threshold has been exceeded, conditional exceedance functions when conditioning on Euclidean covariates, and threshold exceedance functions, which characterize the size of exceedance sets in dependence on covariates for any fixed threshold. We establish consistent estimation with rates of convergence for these targets. The practical merits of the proposed methodology are illustrated through simulations and applications for annual temperature curves and medfly activity profiles.
Advancements in modern science have led to the increasing availability of non-Euclidean data in metric spaces. This article addresses the challenge of modeling relationships between non-Euclidean responses and multivariate Euclidean predictors. We propose a flexible regression model capable of handling high-dimensional predictors without imposing parametric assumptions. Two primary challenges are addressed: the curse of dimensionality in nonparametric regression and the absence of linear structure in general metric spaces. The former is tackled using deep neural networks, while for the latter we demonstrate the feasibility of mapping the metric space where responses reside to a low-dimensional Euclidean space using manifold learning. We introduce a reverse mapping approach, employing local Fr & eacute;chet regression, to map the low-dimensional manifold representations back to objects in the original metric space. We develop a theoretical framework, investigating the convergence rate of deep neural networks under dependent sub-Gaussian noise with bias. The convergence rate of the proposed regression model is then obtained by expanding the scope of local Fr & eacute;chet regression to accommodate multivariate predictors in the presence of errors in predictors. Simulations and case studies show that the proposed model outperforms existing methods for non-Euclidean responses, focusing on the special cases of probability distributions and networks. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.
Regression analysis for responses taking values in general metric spaces has received increasing attention, particularly for settings with Euclidean predictors X ∈ℝ^p and non-Euclidean responses Y ∈ ( ℳ, d). While additive regression is a powerful tool for enhancing interpretability and mitigating the curse of dimensionality in the presence of multivariate predictors, its direct extension is hindered by the absence of vector space operations in general metric spaces. We propose a novel framework for additive optimal transport regression, which incorporates additive structure through optimal geodesic transports. A key idea is to extend the notion of optimal transports in Wasserstein spaces to general geodesic metric spaces. This unified approach accommodates a wide range of responses, including probability distributions, symmetric positive definite (SPD) matrices with various metrics and spherical data. The practical utility of the method is illustrated with correlation matrices derived from resting state fMRI brain imaging data.
Shape-constrained functional data encompass a wide array of application fields, such as activity profiling, growth curves, healthcare and mortality. Most existing methods for general functional data analysis often ignore that such data are subject to inherent shape constraints, while some specialized techniques rely on strict distributional assumptions. We propose an approach for modeling such data that harnesses the intrinsic geometry of functional trajectories by decomposing them into size and shape components. We focus on the two most prevalent shape constraints, positivity and monotonicity, and develop individual-level estimators for the size and shape components. Furthermore, we demonstrate the applicability of our approach by conducting subsequent analyses involving Fr & eacute;chet mean and Fr & eacute;chet regression and establish rates of convergence for the empirical estimators. Illustrative examples include simulations and data applications for activity profiles for Mediterranean fruit flies during their entire lifespan and for data from the Z & uuml;rich longitudinal growth study.
The modeling of samples of distributions is a major challenge since distributions do not form a vector space. While various approaches exist for univariate distributions, including transformations to a Hilbert space, far less is known about the multivariate case. We utilize a transformation approach to map multivariate distributions to a Hilbert space via a Wasserstein slicing method that is invertible. This approach combines functional data analysis tools, such as functional principal component analysis and modes of variation, with the facility to map back to interpretable distributions. We also provide convergence guarantees for the Hilbert space representations under a broad class of such transforms. The method is illustrated using joint systolic and diastolic blood pressure data.
Unobserved effect modifiers can induce bias when generalizing causal effect estimates to target populations. In this work, we extend a sensitivity analysis framework assessing the robustness of study results to unobserved effect modification that adapts to various generalizability scenarios, including multiple (conditionally) randomized trials, observational studies, or combinations thereof. This framework is interpretable and does not rely on distributional or functional assumptions about unknown parameters. We demonstrate how to leverage the multi-study setting to detect violation of the generalizability assumption through hypothesis testing, showing with simulations that the proposed test achieves high power under real-world sample sizes. Finally, we apply our sensitivity analysis framework to analyze the generalized effect estimate of secondhand smoke exposure on birth weight using cohort sites from the Environmental influences on Child Health Outcomes (ECHO) study.
We propose a generalized notion of Wasserstein-Fr & eacute;chet integral of conditional distributions for the classical situation of a joint distribution between two scalar random variables X and Y by viewing the space of probability distributions as a metric space and defining the WassersteinFr & eacute;chet integral of conditional distributions as a Fr & eacute;chet integral. Within this general framework we illustrate various special cases, focusing on the case where one adopts the 2-Wasserstein metric, which however is only one possible choice of metric to implement the proposed method. We demonstrate that this choice often leads to a useful and interpretable notion of the conditional distribution of Y in situations where Y varies systematically with X when one is interested in the residual distribution of Y after the systematic effect is removed. We provide convergence results for the estimated Wasserstein-Fr & eacute;chet integral of conditional distributions for several commonly encountered data generating mechanisms that are of statistical relevance. These include scatterplot data (Xi, Yi), i = 1, ... , n; data where one has a sample of fully observed (conditional) densities along with predictors Xi; and data where one encounters conditional densities that are not fully observed and instead one has samples of the data that they generate.
Mixed effect modeling for longitudinal data is challenging when the observed data are random objects, which are complex data taking values in a general metric space without linear structure. In such settings the classical additive error model and distributional assumptions are unattainable. Due to the rapid advancement of technology, longitudinal data containing complex random objects, such as covariance matrices, data on Riemannian manifolds, and probability distributions are becoming more common. Addressing this challenge, we develop a mixed-effects regression for data in geodesic spaces, where the underlying mean response trajectories are geodesics in the metric space and the deviations of the observations from the model are quantified by perturbation maps or transports. A key finding is that the geodesic trajectories assumption for the case of random objects is a natural extension of the linearity assumption in the standard Euclidean scenario. Further, geodesics can be recovered from noisy observations by exploiting a connection between the geodesic path and the path obtained by global Fr\'echet regression for random objects. The effect of baseline Euclidean covariates on the geodesic paths is modeled by another Fr\'echet regression step. We study the asymptotic convergence of the proposed estimates and provide illustrations through simulations and real-data applications.
Functional linear and single-index models are core regression methods in functional data analysis and are widely used for performing regression in a wide range of applications when the covariates are random functions coupled with scalar responses. In the existing literature, however, the construction of associated estimators and the study of their theoretical properties is invariably carried out on a case-by-case basis for specific models under consideration. In this work, assuming the predictors are Gaussian processes, we provide a unified methodological and theoretical framework for estimating the index in functional linear, and its direction in single-index models. In the latter case, the proposed approach does not require the specification of the link function. In terms of methodology, we show that the reproducing kernel Hilbert space (RKHS) based functional linear least-squares estimator, when viewed through the lens of an infinite-dimensional Gaussian Stein's identity, also provides an estimator of the index of the single-index model. Theoretically, we characterize the convergence rates of the proposed estimators for both linear and single-index models. Our analysis has several key advantages: (i) it does not require restrictive commutativity assumptions for the covariance operator of the random covariates and the integral operator associated with the reproducing kernel; and (ii) the true index parameter can lie outside of the chosen RKHS, thereby allowing for index misspecification as well as for quantifying the degree of such index misspecification. Several existing results emerge as special cases of our analysis.