Experimental design has emerged as a powerful approach for improving the sample efficiency of A/B testing, yet existing designs rely critically on correctly specified models. We study robust sequential experimental design under model misspecification and develop a unified framework that covers both contextual bandit and dynamic settings. Theoretically, we prove that our design bounds the worst-case mean squared error of the estimated treatment effect. Empirically, we demonstrate the effectiveness of the proposed approach using synthetic and real-world datasets from a leading technology company.
A/B testing has become a gold standard for modern technological companies to conduct policy evaluation. Yet, its application to time series experiments, where treatments are sequentially assigned over time, remains challenging. Existing designs suffer from two limitations: (i) they do not fully leverage the entire history for treatment allocation; (ii) they rely on strong assumptions to approximate the objective function (e.g., the mean squared error of the estimated treatment effect) for optimizing the design. We first establish an impossibility theorem showing that failure to condition on the full history leads to suboptimal designs, due to the dynamic dependencies in time series experiments. To address both limitations simultaneously, we next propose a transformer reinforcement learning (RL) approach which leverages transformers to condition treatment allocation on the entire history and employs RL to directly optimize the MSE without relying on restrictive assumptions. Empirical evaluations on synthetic data, a publicly available dispatch simulator, and a real-world ridesharing dataset demonstrate that our proposal consistently outperforms existing designs.
Recent advancements in large language models demonstrate that injecting perturbations can substantially enhance extrapolation performance. However, current approaches often rely on discrete perturbations with fixed designs, which limits their flexibility. In this work, we propose a framework where token prefixes are perturbed by a learnable transformation of a continuous latent vector within an embedding space. To overcome the challenge of an intractable marginal likelihood, we derive unbiased estimating equations for model parameters and optimize them via stochastic gradient descent. We establish the statistical properties of the resulting estimator in over-parameterized regimes. Empirical evaluations on both synthetic and real-world datasets demonstrate that our proposal yields significant gains in out-of-domain settings over a range of state-of-the-art baseline methods.
Individualized randomized experiments are central to online platforms for optimizing personalized decisions in complex environments. In two-sided markets, however, standard treatment effect estimation is often invalid due to strong temporal and cross-unit interference, a challenge compounded when only aggregated data are available because of privacy or system constraints. To address these issues, we identify the Global Average Treatment Effect (GATE) using only group-level data from treatment and control groups. We first establish identification conditions based on aggregated observations, and then propose the Individualized Randomized Experiment Varying Coefficient Decision Process (IRE-VCDP) model, which accounts for interference through supply-demand dynamics. Building on this framework, we develop a complete procedure for estimation and statistical inference of the GATE, along with theoretical guarantees for the proposed test. Extensive simulations and real-world experiments using data from a leading ridesharing platform demonstrate the effectiveness of our approach.
Understanding the causal effects of organ-specific features from medical imaging on clinical outcomes is essential for biomedical research and patient care. We propose a novel Functional Linear Structural Equation Model (FLSEM) to capture the relationships among clinical outcomes, functional imaging exposures, and scalar covariates like genetics, sex, and age. Traditional methods struggle with the infinite-dimensional nature of exposures and complex covariates. Our FLSEM overcomes these challenges by establishing identifiable conditions using scalar instrumental variables. We develop the Functional Group Support Detection and Root Finding (FGS-DAR) algorithm for efficient variable selection, supported by rigorous theoretical guarantees, including selection consistency and accurate parameter estimation. We further propose a test statistic to test the nullity of the functional coefficient, establishing its null limit distribution. Our approach is validated through extensive simulations and applied to UK Biobank data, demonstrating robust performance in detecting causal relationships from medical imaging.
Alzheimer's Disease Neuroimaging Initiative (ADNI) diagnostic groups present strong heterogeneous associations among demographic, imaging, and cognitive data. We propose a novel PArtially-shared Imaging Regression (PAIR) model to represent imaging coefficients as weighted combinations of smooth spatial components. A Total Variation penalty is applied to enforce spatial smoothness, and a Selective Integration penalty is introduced to adaptively learn partial-sharing structures across groups. Theoretically, we establish minimax-optimal error bounds that dynamically adapt to varying sharing paradigms. Numerically, PAIR achieves predictive accuracy comparable to advanced deep learning models while providing superior interpretability. Applied to ADNI data, PAIR reveals substantial heterogeneity in brain-cognition pathways between cognitively normal (CN) and cognitively impaired (CI) groups, with hippocampal imaging contributing minimally in the CN group but substantially in the CI group, particularly in the CA1, CA3, and presubiculum subfields.
BACKGROUND:Alzheimer's disease (AD), the predominant cause of dementia, is characterized by progressive cognitive deterioration, significantly impacting public health due to the absence of curative treatments. The insidious progression from asymptomatic stages to full-blown dementia underscores the urgency of developing predictive tools for early detection. Recent advancements in proteomics and neuroimaging have identified potential biomarkers and structural brain changes associated with the disease, offering new pathways for understanding and potentially intervening in its progression. METHOD:We designed a novel double machine learning framework to de-bias the analysis of complex block-missing multi-omics data. This approach enabled a comprehensive evaluation of the causal relationships between plasma biomarkers (GFAP, NEFL, GDF15, LTBP2) and hippocampal atrophy in the context of Alzheimer's disease progression. Our methodology combines rigorous statistical analysis with state-of-the-art machine learning techniques to ensure robust and reproducible findings. RESULT:We identified multiple significant protein biomarkers and neuroimaging features associated with Alzheimer's disease (AD). The most prominent effects emerged in three protein biomarkers-GFAP, NEFL, and CST5-implicating astroglial activation, axonal injury, and inflammatory pathways. Functional connectivity disruptions (e.g., Net_Nodes_PC9, Net100_Pair21_38) underscored large-scale network alterations, while additional proteins (BCAN, IGF2R, IGFBP3, FCRL5, PDGFC, ERP44, TNFRSF10A) indicated metabolic and immune-related mechanisms. White matter integrity measures from diffusion MRI (RD_PTR_V3, FA_SFO_V5, AD_SCC_V2, FX.FA, FA_FX_V1, MD_SCC_V2) revealed microstructural damage in tracts tied to cognitive and memory functions. Lastly, hippocampal subfield changes (HippSubfieds_20, HippSubfieds_21, HippSubfieds_71) further highlighted the hippocampus's central role in AD pathology. CONCLUSION:By providing a comprehensive analysis of the key biomarkers and neuroanatomical changes, our study advances the understanding of Alzheimer's disease progression and opens avenues for the development of targeted treatments.
Traditional conformal prediction faces significant challenges with the rise of streaming data and increasing concerns over privacy. In this paper, we introduce a novel online differentially private conformal prediction framework, designed to construct dynamic, model-free private prediction sets. Unlike existing approaches that either disregard privacy or require full access to the entire dataset, our proposed method ensures individual privacy with a one-pass algorithm, ideal for real-time, privacy-preserving decision-making. Theoretically, we establish guarantees for long-run coverage at the nominal confidence level. Moreover, we extend our method to conformal quantile regression, which is fully adaptive to heteroscedasticity. We validate the effectiveness and applicability of the proposed method through comprehensive simulations and real-world studies on the ELEC2 and PAMAP2 datasets.
Alzheimer's disease (AD), the predominant cause of dementia, is characterized by progressive cognitive deterioration, significantly impacting public health due to the absence of curative treatments. The insidious progression from asymptomatic stages to full-blown dementia underscores the urgency of developing predictive tools for early detection. Recent advancements in proteomics and neuroimaging have identified potential biomarkers and structural brain changes associated with the disease, offering new pathways for understanding and potentially intervening in its progression. We designed a novel double machine learning framework to de-bias the analysis of complex block-missing multi-omics data. This approach enabled a comprehensive evaluation of the causal relationships between plasma biomarkers (GFAP, NEFL, GDF15, LTBP2) and hippocampal atrophy in the context of Alzheimer's disease progression. Our methodology combines rigorous statistical analysis with state-of-the-art machine learning techniques to ensure robust and reproducible findings. We identified multiple significant protein biomarkers and neuroimaging features associated with Alzheimer's disease (AD). The most prominent effects emerged in three protein biomarkers—GFAP, NEFL, and CST5—implicating astroglial activation, axonal injury, and inflammatory pathways. Functional connectivity disruptions (e.g., Net_Nodes_PC9, Net100_Pair21_38) underscored large-scale network alterations, while additional proteins (BCAN, IGF2R, IGFBP3, FCRL5, PDGFC, ERP44, TNFRSF10A) indicated metabolic and immune-related mechanisms. White matter integrity measures from diffusion MRI (RD_PTR_V3, FA_SFO_V5, AD_SCC_V2, FX.FA, FA_FX_V1, MD_SCC_V2) revealed microstructural damage in tracts tied to cognitive and memory functions. Lastly, hippocampal subfield changes (HippSubfieds_20, HippSubfieds_21, HippSubfieds_71) further highlighted the hippocampus's central role in AD pathology. By providing a comprehensive analysis of the key biomarkers and neuroanatomical changes, our study advances the understanding of Alzheimer's disease progression and opens avenues for the development of targeted treatments.
This paper studies how to integrate historical control data with experimental data to enhance A/B testing, while addressing the distributional shift between historical and experimental datasets. We propose a pessimistic data integration method that combines two causal effect estimators constructed based on experimental and historical datasets. Our main idea is to conceptualize the weight function for this combination as a policy so that existing pessimistic policy learning algorithms are applicable to learn the optimal weight that minimizes the resulting weighted estimator's mean squared error. Additionally, we conduct comprehensive theoretical and empirical analyses to compare our method against various baseline estimators across five scenarios. Both our theoretical and numerical findings demonstrate that the proposed estimator achieves near-optimal performance across all scenarios.
Human organ structure and function are important endophenotypes for clinical outcomes. Genome-wide association studies (GWAS) have identified numerous common variants associated with phenotypes derived from magnetic resonance imaging (MRI) of the brain and body. However, the role of rare protein-coding variations affecting organ size and function is largely unknown. Here we present an exome-wide association study that evaluates 596 multi-organ MRI traits across over 50,000 individuals from the UK Biobank. We identified 107 variant-level associations and 224 gene-based burden associations (67 unique gene-trait pairs) across all MRI modalities, including PTEN with total brain volume, TTN with regional peak circumferential strain in the heart left ventricle, and TNFRSF13B with spleen volume. The singleton burden model and AlphaMissense annotations contributed 8 unique gene-trait pairs including the association between an approved drug target gene of KCNA5 and brain functional activity. The identified rare coding signals elucidate some shared genetic effects across organs, prioritize previously identified GWAS loci, and are enriched for drug targets. Overall, we demonstrate how rare variants enhance our understanding of genetic effects on human organ morphology and function and their connections to complex diseases.
This paper is motivated by the joint analysis of genetic, imaging, and clinical (GIC) data collected in the Alzheimer's Disease Neuroimaging Initiative (ADNI) study. We propose a regression framework based on partially functional linear regression models to map high-dimensional GIC-related pathways for Alzheimer's Disease (AD). We develop a joint model selection and estimation procedure by embedding imaging data in the reproducing kernel Hilbert space and imposing the L0 penalty for the coefficients of genetic variables. We apply the proposed method to the ADNI dataset to identify important features from tens of thousands of genetic polymorphisms (reduced from millions using a preprocessing step) and study the effects of a certain set of informative genetic variants and the baseline hippocampus surface on thirteen future cognitive scores measuring different aspects of cognitive function. We explore the shared and different heritability patterns of these cognitive scores. Analysis results suggest that both the hippocampal and genetic data have heterogeneous effects on different scores, with the trend that the value of both hippocampi is negatively associated with the severity of cognition deficits. Polygenic effects are observed for all thirteen cognitive scores. The well-known APOE4 genotype only explains a small part of cognitive function. Shared genetic etiology exists, however, greater genetic heterogeneity exists within disease classifications after accounting for the baseline diagnosis status. These analyses are useful in further investigation of functional mechanisms for AD evolution.
Many modern tech companies, such as Google, Uber, and Didi, utilize online experiments (also known as A/B testing) to evaluate new policies against existing ones. While most studies concentrate on average treatment effects, situations with skewed and heavy-tailed outcome distributions may benefit from alternative criteria, such as quantiles. However, assessing dynamic quantile treatment effects (QTE) remains a challenge, particularly when dealing with data from ride-sourcing platforms that involve sequential decision-making across time and space. In this paper, we establish a formal framework to calculate QTE conditional on characteristics independent of the treatment. Under specific model assumptions, we demonstrate that the dynamic conditional QTE (CQTE) equals the sum of individual CQTEs across time, even though the conditional quantile of cumulative rewards may not necessarily equate to the sum of conditional quantiles of individual rewards. This crucial insight significantly streamlines the estimation and inference processes for our target causal estimand. We then introduce two varying coefficient decision process (VCDP) models and devise an innovative method to test the dynamic CQTE. Moreover, we expand our approach to accommodate data from spatiotemporal dependent experiments and examine both conditional quantile direct and indirect effects. To showcase the practical utility of our method, we apply it to three real-world datasets from a ride-sourcing platform. Theoretical findings and comprehensive simulation studies further substantiate our proposal.
This paper studies policy evaluation with multiple data sources, especially in scenarios that involve one experimental dataset with two arms, complemented by a historical dataset generated under a single control arm. We propose novel data integration methods that linearly integrate base policy value estimators constructed based on the experimental and historical data, with weights optimized to minimize the mean square error (MSE) of the resulting combined estimator. We further apply the pessimistic principle to obtain more robust estimators, and extend these developments to sequential decision making. Theoretically, we establish non-asymptotic error bounds for the MSEs of our proposed estimators, and derive their oracle, efficiency and robustness properties across a broad spectrum of reward shift scenarios. Numerical experiments and real-data-based analyses from a ridesharing company demonstrate the superior performance of the proposed estimators.
In this article, we study the transfer learning problem in functional classification, aiming to improve the classification accuracy of the target data by leveraging information from related source datasets. To facilitate transfer learning, we propose a novel transferability function tailored for classification problems, enabling a more accurate evaluation of the similarity between source and target dataset distributions. Interestingly, we find that a source dataset can offer more substantial benefits under certain conditions than another dataset with an identical distribution to the target dataset. This observation renders the commonly-used debiasing step in the parameter-based transfer learning algorithm unnecessary under some circumstances to the classification problem. In particular, we propose two adaptive transfer learning algorithms based on the functional Distance Weighted Discrimination (DWD) classifier for scenarios with and without prior knowledge regarding informative sources. Furthermore, we establish the upper bound on the excess risk of the proposed classifiers, providing the statistical gain via transfer learning mathematically provable. Simulation studies are conducted to thoroughly examine the finite-sample performance of the proposed algorithms. Finally, we implement the proposed method to Beijing air-quality data, and significantly improve the prediction of the PM 2.5 level of a target station by effectively incorporating information from source datasets. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.
A/B testing is critical for modern technological companies to evaluate the effectiveness of newly developed products against standard baselines. This paper studies optimal designs that aim to maximize the amount of information obtained from online experiments to estimate treatment effects accurately. We propose three optimal allocation strategies in a dynamic setting where treatments are sequentially assigned over time. These strategies are designed to minimize the variance of the treatment effect estimator when data follow a non-Markov decision process or a (time-varying) Markov decision process. We further develop estimation procedures based on existing off-policy evaluation (OPE) methods and conduct extensive experiments in various environments to demonstrate the effectiveness of the proposed methodologies. In theory, we prove the optimality of the proposed treatment allocation design and establish upper bounds for the mean squared errors of the resulting treatment effect estimators.
Many modern large-scale longitudinal neuroimaging studies, such as the Alzheimer’s Disease Neuroimaging Initiative (ADNI) study, have collected/are collecting asynchronous scalar and functional variables that are measured at distinct time points. The analyses of temporally asynchronous functional and scalar variables pose major technical challenges to many existing statistical approaches. We propose a class of generalized functional partial-linear varying-coefficient models to appropriately deal with these challenges through introducing both scalar and functional coefficients of interest and using kernel weighting methods. We design penalized kernel-weighted estimating equations to estimate scalar and functional coefficients, in which we represent functional coefficients by using a rich truncated tensor product penalized B-spline basis. We establish the theoretical properties of scalar and functional coefficient estimators including consistency, convergence rate, prediction accuracy, and limiting distributions. We also propose a bootstrap method to test the nullity of both parametric and functional coefficients, while establishing the bootstrap consistency. Simulation studies and the analysis of the ADNI study are used to assess the finite sample performance of our proposed approach. Our real data analysis reveals significant relationship between fractional anisotropy density curves and cognitive function with education, baseline disease status and APOE4 gene as major contributing factors. Supplementary materials for this article are available online.
Motivated by the analysis of longitudinal neuroimaging studies, we study the longitudinal functional linear regression model under asynchronous data setting for modeling the association between clinical outcomes and functional (or imaging) covariates. In the asynchronous data setting, both covariates and responses may be measured at irregular and mismatched time points, posing methodological challenges to existing statistical methods. We develop a kernel weighted loss function with roughness penalty to obtain the functional estimator and derive its representer theorem. The rate of convergence, a Bahadur representation, and the asymptotic pointwise distribution of the functional estimator are obtained under the reproducing kernel Hilbert space framework. We propose a penalized likelihood ratio test to test the nullity of the functional coefficient, derive its asymptotic distribution under the null hypothesis, and investigate the separation rate under the alternative hypotheses. Simulation studies are conducted to examine the finite-sample performance of the proposed procedure. We apply the proposed methods to the analysis of multitype data obtained from the Alzheimer's Disease Neuroimaging Initiative (ADNI) study, which reveals significant association between 21 regional brain volume density curves and the cognitive function. Data used in preparation of this paper were obtained from the ADNI database (adni.loni.usc.edu).
Classical clusterwise linear regression is a useful method for investigating the relationship between scalar predictors and scalar responses with heterogeneous variation of regression patterns for different subgroups of subjects. This paper extends the classical clusterwise linear regression to incorporate multiple functional predictors by representing the functional coefficients in terms of a functional principal component basis. We estimate the functional principal component coefficients based on M-estimation and K-means clustering algorithm, which can classify the data into clusters and estimate clusterwise coefficients simultaneously. One advantage of the proposed method is that it is robust and flexible by adopting a general loss function, which can be broadly applied to mean regression, median regression, quantile regression and robust mean regression. A Bayesian information criterion is proposed to select the unknown number of groups and shown to be consistent in model selection. We also obtain the convergence rate of the set of estimators to the set of true coefficients for all clusters. Simulation studies and real data analysis show that the proposed method is easily implemented, and it consequently improves previous works and also requires much less computing burden than existing methods.
In this paper, we consider composite quantile regression for partial functional linear regression model with polynomial spline approximation. Under some mild conditions, the convergence rates of the estimators and mean squared prediction error, and asymptotic normality of parameter vector are obtained. Simulation studies demonstrate that the proposed new estimation method is robust and works much better than the least-squares based method when there are outliers in the dataset or the random error follows heavy-tailed distributions. Finally, we apply the proposed methodology to a spectroscopic data sets to illustrate its usefulness in practice.