Many biomedical studies collect high-dimensional medical imaging data to identify biomarkers for the detection, diagnosis, and treatment of human diseases. Consequently, it is crucial to develop accurate models that can predict a wide range of clinical outcomes (both discrete and continuous) based on imaging data. By treating imaging predictors as functional data, we propose a residual-based alternative partial least squares (RAPLS) model for a broad class of generalized functional linear models that incorporate both functional and scalar covariates. Our RAPLS method extends the alternative partial least squares (APLS) algorithm iteratively to accommodate additional scalar covariates and non-continuous outcomes. We establish the convergence rate of the RAPLS estimator for the unknown slope function and, with an additional calibration step, we prove the asymptotic normality and efficiency of the calibrated RAPLS estimator for the scalar parameters. The effectiveness of the RAPLS algorithm is demonstrated through multiple simulation studies and an application predicting Alzheimer's disease progression using neuroimaging data from the Alzheimer's Disease Neuroimaging Initiative (ADNI).
Imaging genetics links genetic variations to brain structures and functions, but the computational challenges posed by high-dimensional imaging and genetic data are significant. In voxel-level genome-wide association studies, we introduce a Representation learning-based Voxel-level Genetic Analysis (RVGA) framework that reduces computational time and storage burden by over 200 times. RVGA enhances statistical power by denoising images and shares minimal datasets of summary statistics for associations across the whole genome of the entire image for secondary analyses. Additionally, it introduces a unified estimator for voxel heritability, genetic correlations between voxels, and cross-trait genetic correlations between voxels and non-imaging phenotypes. Applying RVGA to hippocampus shape and white matter microstructure in the UK Biobank (n = 53,454) reveals 39 and 275 novel loci, respectively. We identify heterogeneity in heritability within images and subregions that share genetic bases with 14 brain-related phenotypes, such as the genetic correlation between the hippocampus and educational attainment, and between the anterior corona radiata and schizophrenia. RVGA replicates known genetic associations and uncovers new discoveries.
Experimental design has emerged as a powerful approach for improving the sample efficiency of A/B testing, yet existing designs rely critically on correctly specified models. We study robust sequential experimental design under model misspecification and develop a unified framework that covers both contextual bandit and dynamic settings. Theoretically, we prove that our design bounds the worst-case mean squared error of the estimated treatment effect. Empirically, we demonstrate the effectiveness of the proposed approach using synthetic and real-world datasets from a leading technology company.
Electronic health records (EHRs) are a primary target for foundation models because they capture longitudinal, multimodal, and large-scale clinical trajectories. However, routinely collected clinical data remain difficult to model due to extreme sparsity, irregular temporal sampling, heterogeneous representations, and informative missingness. In this Perspective, we provide an EHR-centered synthesis of foundation models across the translational pipeline, spanning data resources, model architectures, and the requirements for reliable clinical integration. We first examine the data characteristics and real-world barriers that shape EHR modeling, including coding variation and distribution shift. We then review the three dominant model families, including structured-sequence, clinical language, and multimodal models, with a focus on architectural adaptations for longitudinal clinical reasoning. Next, we summarize emerging clinical use cases, from outcome prediction and phenotyping to assistive documentation, digital-twin simulation, and causal inference. Finally, we delineate the engineering and governance conditions necessary for clinical translation, including evaluation under distribution shift, workflow-integrated monitoring, and regulatory oversight. By linking technical advances to the practical constraints of healthcare delivery, this Perspective outlines a roadmap for transitioning EHR foundation models from research benchmarks to safe, impactful clinical technologies.
BACKGROUND:Sleep is crucial for overall physical and mental health, concerning organs such as the brain, heart, eye, liver, kidney, and lung. Nonetheless, a thorough understanding of how sleep relates to anatomical features of these organs, as well as their genetic bases, remains elusive. METHODS:We analyzed ten sleep traits in relation to 623 imaging-derived biomarkers capturing the structure and function of multiple organs from UK Biobank (UKB). We examined phenotypic and genetic sleep-imaging associations, identified shared genetic loci, assessed genetic correlations between sleep traits and a wide range of diseases, and performed mediation analyses to evaluate the role of organ-related diseases in sleep-imaging connections. RESULTS:Here we show that sleep traits are robustly associated with the structure and function of multiple organs at both the phenotypic and genetic levels, including brain functions measured by functional magnetic resonance imaging (fMRI) and body composition traits in abdominal MRI. Sleep and imaging traits share genetic influences across 51 genomic regions, 23 of which show evidence of colocalized causal genetic effects. We also exhibit genetic similarities between sleep traits and diseases affecting multiple organ systems, with psychiatric disorders consistently showing the strongest genetic correlations and causal links. Furthermore, many sleep-imaging associations are mediated by diseases within or across organ systems. CONCLUSIONS:These findings demonstrate that sleep is broadly linked to brain and body health and influenced in part by shared genetic factors. Integrating sleep traits with multi-organ imaging measures provides a framework for characterizing organ-specific sleep associations and their potential relevance to disease.
Relation extraction (RE) is a core task in natural language processing. Traditional approaches typically frame RE as a supervised learning problem, directly mapping context to labels—an approach that often suffers from poor out-of-domain (OOD) generalization. Inspired by the workflow of human annotators, we reframe RE as a reasoning task guided by annotation guidelines and introduce R1-RE, the first reinforcement learning with verifiable reward (RLVR) framework for RE tasks. Our method elicits the reasoning abilities of small language models for annotation tasks, resulting in significantly improved OOD robustness. We evaluate our approach on the public Sem-2010 dataset and a private MDKG dataset. The R1-RE-7B model attains an average OOD accuracy of approximately 70%, on par with leading proprietary models such as GPT-4o. Additionally, our comprehensive analysis provides novel insights into the training dynamics and emergent reasoning behaviors of the RLVR paradigm for RE.
Pathology foundation models (PFMs) have demonstrated strong potential across clinical and scientific applications, yet their performance is often hindered by batch effects, which are non-biological variations across tissue source institutions (TSIs) that distort learned feature representations and impair generalization. Conventional mitigation strategies, such as stain normalization, offer limited success in addressing these high-dimensional, complex artifacts. We present GLMP (General-purpose LLM-Mediated Pathology model), a novel framework that generates robust numerical embeddings from histology image patches through an intermediate textual representation. By leveraging pretrained general-purpose multimodal large language models (MLLMs) and text encoders, GLMP effectively prioritizes biologically meaningful signals over TSI-specific artifacts, thereby improving cross-institutional generalization. To our knowledge, GLMP is the first pathology model to use text descriptions of histological features as an intermediate representation for generating numerical embeddings from histology images. Our results highlight the untapped potential of broad-domain, non-specialized MLLMs in computational pathology and introduce a new paradigm for building versatile, generalizable, and robust pathology models.
Tumor subclonal architecture shapes cancer evolution, yet subclonal reconstruction from bulk sequencing remains difficult to scale due to computational cost and model complexity. We present CliPP, a penalized-likelihood framework that jointly estimates cellular prevalence with pairwise fusion penalties, automatically identifying subclones without requiring extensive priors. Across simulations and 2,778 whole-genome tumors with external consensus reconstructions, CliPP achieves consistently good performances when compared to state-of-the-art approaches while providing substantial runtime reductions. Applied to 7,000+ tumors across >30 cancer types, CliPP quantifies pervasive subclonality and delineates cohort-level subclone landscapes. CliPP enables fast, reproducible large-scale subclonal analysis and is freely available to the community through GitHub and a shiny app.
Sampled-token on-policy distillation (OPD) efficiently transfers capabilities from teacher to student using student-generated tokens, requiring teacher probabilities only for sampled tokens. Yet it frequently suffers from diversity distillation failure: the student's pass@1 improves while its pass@k plateaus, failing to inherit the teacher's diversity. To explain this, we introduce First-Order Local Entropy Influence, a signed first-order proxy that decouples each update's entropy effect into the teacher–student log-probability gap and the student's local probability structure, and empirically links entropy contraction to negative-influence positions. Motivated by this, we propose Influence-Directed Adaptive On-Policy Distillation (IDA-OPD): rather than relying on costly full-vocabulary Forward-KL objectives, it preserves entropy-expanding updates while replacing entropy-contracting ones with divergence-adaptive advantage shrinkage, using only the teacher's sampled-token log-probability. Experiments on reasoning-oriented distillation show IDA-OPD consistently improves pass@k, inheriting the teacher's diversity through distillation, matches the strongest teacher-informed methods at strictly lower cost, and broadly maintains vanilla OPD's pass@1, all without full-vocabulary teacher information.
A/B testing has become a gold standard for modern technological companies to conduct policy evaluation. Yet, its application to time series experiments, where treatments are sequentially assigned over time, remains challenging. Existing designs suffer from two limitations: (i) they do not fully leverage the entire history for treatment allocation; (ii) they rely on strong assumptions to approximate the objective function (e.g., the mean squared error of the estimated treatment effect) for optimizing the design. We first establish an impossibility theorem showing that failure to condition on the full history leads to suboptimal designs, due to the dynamic dependencies in time series experiments. To address both limitations simultaneously, we next propose a transformer reinforcement learning (RL) approach which leverages transformers to condition treatment allocation on the entire history and employs RL to directly optimize the MSE without relying on restrictive assumptions. Empirical evaluations on synthetic data, a publicly available dispatch simulator, and a real-world ridesharing dataset demonstrate that our proposal consistently outperforms existing designs.
This paper introduces Mixtures of Geodesic Factor Analyzers (MGFA) on Riemannian homogeneous spaces. MGFA uses a geodesic factor model within each mixture component, providing greater expressiveness than mixtures of Riemannian radial distributions and enabling clustering of manifold-valued data with anisotropic subpopulations. We establish root-n consistency for the MGFA maximum likelihood estimator (MLE), thereby filling a theoretical gap for mixtures of Riemannian radial distributions as a special case. We also propose an iterative estimation algorithm and implement it on spheres, shape spaces, and hyperbolic spaces. Numerical experiments show that MGFA substantially outperforms competing methods in well-specified regimes while remaining robust under model misspecification. Finally, case studies on corpus callosum and left hippocampus shape datasets demonstrate MGFA's effectiveness for both 2D contour and 3D shape analysis.
Phthalates and replacement plasticizers (PRPs) are ubiquitous exposures in daily life across all age ranges. Exposure to phthalates has been linked to changes in cognitive and behavioral development and associated with increased risk of some developmental disabilities. We examined the extent to which early life exposure to PRPs was associated with changes in connection strength of resting state functional networks or impacted structural morphologies of cortical regions of interest that underlie basic and higher order cognitions. We utilized the UNC Chapel Hill enrollment in the Baby Connectome Project, a longitudinal study of normative brain development of children between 2 weeks and 5 years of age. Non-sedated structural and resting state functional magnetic resonance imaging of the brain during natural sleep were obtained longitudinally, along with urine samples that were analyzed for 17 PRP metabolites. Using kernal weighted estimating equations and generalized linear models, we identified multiple PRP metabolites were associated with alterations in within-network connection strengths in the executive control and dorsal attention networks, with directionality often differing between boys and girls. PRP exposure among boys tended to be associated with lower functional connectivity, whereas PRP exposure among girls tended to be associated with higher functional connectivity. Among girls, MiBP metabolite concentrations were also significantly associated with cortical thinning in several regions of interest in the temporal lobe. Our results indicate that exposure to PRPs in early life has a measurable impact on the developmental trajectory of brain maturation, with potentially important differences by child sex.
This article presents the full, original record of the 2024 Joint Statistical Meetings (JSM) town hall, "Statistics in the Age of AI," which convened leading statisticians to discuss how the field is evolving in response to advances in artificial intelligence, foundation models, large-scale empirical modeling, and data-intensive infrastructures. The town hall was structured around open panel discussion and extensive audience Q&A, with the aim of eliciting candid, experience-driven perspectives rather than formal presentations or prepared statements. This document preserves the extended exchanges among panelists and audience members, with minimal editorial intervention, and organizes the conversation around five recurring questions concerning disciplinary culture and practices, data curation and "data work," engagement with modern empirical modeling, training for large-scale AI applications, and partnerships with key AI stakeholders. By providing an archival record of this discussion, the preprint aims to support transparency, community reflection, and ongoing dialogue about the evolving role of statistics in the data- and AI-centric future.
Echocardiography plays an important role in the screening and diagnosis of cardiovascular diseases. However, automated intelligent analysis of echocardiographic data remains challenging due to complex cardiac dynamics and strong view heterogeneity. In recent years, visual language models (VLM) have opened a new avenue for building ultrasound understanding systems for clinical decision support. Nevertheless, most existing methods formulate this task as a direct mapping from video and question to answer, making them vulnerable to template shortcuts and spurious explanations. To address these issues, we propose EchoTrust, an evidence-driven Actor-Verifier framework for trustworthy reasoning in echocardiography VLM-based agents. EchoTrust produces a structured intermediate representation that is subsequently analyzed by distinct roles, enabling more reliable and interpretable decision-making for high-stakes clinical applications.
Recent advancements in artificial intelligence (AI) have significantly influenced the field of cardiovascular disease (CVD) analysis, particularly in image-based diagnostics. Our article presents an extensive review of AI applications in image-based CVD analysis, offering insights into its current state and future potential. We systematically categorize the literature based on the primary anatomical structures related to CVD, dividing them into nonvessel structures (such as ventricles and atria) and vessel structures (including the aorta and coronary arteries). This categorization provides a structured approach to explore various imaging modalities like computed tomography and magnetic resonance imaging, which are commonly used in CVD research. Our review encompasses these modalities, giving a broad perspective on the diverse imaging techniques integrated with AI for CVD analysis. We conclude with an examination of the challenges and limitations inherent in current AI-based CVD analysis methods and suggest directions for future research to overcome these hurdles.
Human white matter has been linked to inherited variation, circulating molecular state and brain disease, but these layers have rarely been mapped onto the same tract anatomy. Here we measured genetic effects along 6,090 atlas-aligned fiber pathways sampled at 609,000 locations in 72,185 UK Biobank participants, and integrated proteomic and metabolomic profiles within the same anatomical frame. Genetic effects were not whole-tract properties: each locus formed a spatial footprint along fiber trajectories, ranging from single locations to broad multi-tract patterns and reflecting regional polygenicity rather than tract heritability. This map identified 258, 186 and 298 previously unreported loci for fractional anisotropy, mean diffusivity and axial diffusivity; spatial patterns replicated in adults and 157 of 315 FA loci replicated in adolescence in ABCD. Mendelian randomization linked localized genetic effects to neurodegenerative and psychiatric traits, with Alzheimer's disease showing directional effects across 12 of 17 tracts. Multi-omic analyses identified 97 proteomic and 161 metabolomic associations, with the broadest signals from lipid metabolites including linoleic acid and phosphatidylcholines. The strongest lipid-metabolite and genetic signals converged in the corpus callosum, placing inherited variation, disease risk and systemic lipid metabolism on the same localized tract segments.
Large-scale population analyses of structural connectome organization remain challenging because of cross-subject alignment, pathway interpretability and computational burden. No widely adopted standard exists for systematic evaluation across processing methods. We developed connectome-based spatial statistics (CBSS), a scalable framework for anatomically aligned and functionally informed quantification of white-matter microstructure that yields atlas-defined pathways organized into 13 functional networks. Using data from 56,510 UK Biobank participants together with five independent lifespan cohorts, we evaluated the streamline-, voxel- and network-level measures in the aspects of reliability, heritability, structure-function coupling, cognitive and behavioral prediction, brain aging patterns and lifespan trajectories across cohorts. The systematic evaluation workflow compares population-level white-matter representations across methods, spatial scales, tasks and datasets. The results support CBSS as a common connectome reference for large-scale, cross-cohort diffusion MRI studies.
Test-time scaling improves the reasoning performance of large language models but incurs substantial cost in both total computation and latency. Existing adaptive sampling methods partially mitigate this issue by dynamically deciding when to stop sampling, yet they typically rely on heuristic rules or rely on distribution assumptions. In this work, we formulate adaptive sampling as a Markov decision process (MDP). We train a lightweight sampling controller with reinforcement learning (RL) to jointly balance answer correctness, latency, and computation cost. At each round, the controller decides to stop sampling or to acquire additional samples. Our method is lightweight which only relies on statistics of final answers, and can be trained and deployed on CPU. We further show that the resulting framework admits an interpretation as the Lagrangian relaxation of a constrained optimization problem with explicit budget constraints. Experiments against strong baselines such as ASC and ESC show that our method achieves improved trade-offs among answer correctness, sampling rounds, and total samples required.
Ravi Bansal合作论文数Duke University: The Fuqua School of Business16