Recent evidence from small animal models and human electrophysiology suggests that the OFF-pathway is more vulnerable to glaucomatous insult than the ON-pathway. Thus, OFF-pathway based measurements of visual function may be useful in the diagnosis of Glaucoma. The steady-state visually evoked potential (SSVEP) can be used to non-invasively make such functional measurements. Here, we examine whether OFF- and ON-pathway biasing SSVEP measurements differently predict glaucoma diagnosis using a large cohort of 98 glaucoma patients and 71 controls. Using both a logistic regression with k-fold cross-validation and a random forest classifier, we show that OFF-pathway biasing features produce a small improvement in predictive accuracy over ON-pathway biasing features. However, despite our inclusion of many more response features and the retention of both participants' eyes, our classifier did not perform as well as previous reports that used the isolated-check VEP. This is likely a result of the relatively small amount of data we collected for each participant, but may also be explained by the absence of any train-test splitting in preexisting work. Nevertheless, our results support further exploration of the diagnostic potential of OFF-pathway biasing functional biomarkers for glaucoma.
Quasi-experimental evaluations are central for generating real-world causal evidence and complementing insights from randomized trials. The regression discontinuity design (RDD) is a quasi-experimental design that can be used to estimate the causal effect of treatments that are assigned based on a running variable crossing a threshold. Such threshold-based rules are ubiquitous in healthcare, where predictive and prognostic biomarkers frequently guide treatment decisions. However, standard RD estimators rely on complete outcome data, an assumption often violated in time-to-event analyses where censoring arises from loss to follow-up. To address this issue, we propose a nonparametric approach that leverages doubly robust censoring corrections and can be paired with existing RD estimators. Our approach can handle multiple survival endpoints, long follow-up times, and covariate-dependent variation in survival and censoring. We discuss the relevance of our approach across multiple areas of applications and demonstrate its usefulness through simulations and the prostate component of the Prostate, Lung, Colorectal and Ovarian (PLCO) Cancer Screening Trial where our new approach offers several advantages, including higher efficiency and robustness to misspecification. We have also developed an open-source software package, , for the language.
The electrocardiogram (ECG) encodes the electrical activity of the heart across multiple timescales, yet standard clinical analysis collapses this rich signal into a handful of scalar measurements that discard most of the waveform's structure. Whether the frequency signals lost in this reduction carry heritable biological information relevant to cardiovascular disease risk remains unclear. Here we decompose resting 12-lead ECGs from 47,052 White British UK Biobank participants into 84 frequency-specific energy features using Daubechies-6 wavelet analysis across 12 leads and 7 decomposition levels, and perform independent genome-wide association analyses on each feature. We identify 67 independent loci (p < 5 × 10-8) and refine these to 101 high-confidence causal variants (posterior inclusion probability > 0.80) through Bayesian fine-mapping; associated loci converge on genes governing cardiac conduction and myocardial integrity, including SCN5A, TTN, KCNQ1, and DSP, alongside less-characterized cardiomyopathy candidates. SNP-based heritability estimates range from 0.03 to 0.26, with the strongest signals in mid-frequency bands (D6-D4, ~4-32 Hz) of Lead I and aVR, and strong inter-lead genetic correlations indicate a coordinated genetic architecture underlying the waveform. Integrating these features with FinnGen R12 cardiovascular phenotypes reveals genetic correlations reaching 0.56 with heart failure, driven predominantly by energy in the highest-frequency band (D1, 125-250 Hz) - a spectral range routinely filtered from clinical ECGs and previously regarded as acquisition noise. These results reframe the electrocardiogram as a multi-frequency genetic phenotype, expand the set of cardiac loci discoverable from ECG data, and implicate high-frequency cardiac electrical activity as an underexplored dimension of cardiovascular disease risk.
Biobank-scale genomic analyses remain computationally expensive, CPU-bound workflows, particularly when adjusting for confounding. Here, we present CuGen, a GPU-accelerated framework for large-scale genomics. CuGen uses UltraLasso, a novel hierarchical application of univariate-guided sparse regression (uniLasso), to select a compact, phenotype-informed active set of fewer than 30,000 variants. This achieves robust leave-one-chromosome-out (LOCO) confounding control, enabling both downstream GWAS and in-sample fine-mapping. Additionally, we introduce the .cugen file format, a genotype representation designed for memory-optimized, high-throughput streaming and random access on GPU hardware. Building on this substrate, we provide a general GPU-accelerated genomics toolkit handling polygenic prediction, data manipulation, quality control, analysis, and visualization. We demonstrate CuGen's efficacy in the UK Biobank with up to 408,624 individuals, where the full GWAS pipeline and fine-mapping against 6.8 million imputed variants completes in approximately 10 minutes on a single high-throughput GPU with 80 GB of memory. The pipeline scales efficiently to massive phenome-wide analyses with sublinear resource consumption.
Prognostic models in oncology have a profound impact on personalized cancer care and patient profiling, but tend to be heterogeneously developed and implemented in narrow patient cohorts. Here, we develop and benchmark multiple machine learning models to predict survival in pan-cancer and 16 single-cancer settings using a de-identified clinico-genomic database of 28,079 US patients with cancer. We identify key predictors of cancer prognosis, including 15 shared across seven or more cancer types, revealing strong consistency in cancer prognostic factors. We demonstrate that pan-cancer models generally outperform or match single-cancer models in predicting survival and risk stratifying patients, especially in smaller cancer cohorts, suggesting a unique transfer learning advantage of pan-cancer models. This work demonstrates the potential of pan-cancer approaches in enhancing the accuracy and applicability of prognostic models in oncology, paving the way for more personalized and effective cancer care strategies.
Classifying patients into different clusters based on genomic data can offer valuable insights into their projected disease-specific survival trajectories over time. Here we apply two supervised learning methods—Nearest Shrunken Centroids and LASSO— to the METABRIC Breast Cancer data metabric of 1980 patients and 754 genes to perform this task. The pamr R package implements the Nearest Shrunken Centroids classifier and the glmnet R package is used to fit an ungrouped multinomial model, a grouped multinomial model, and a One-Versus-Rest model. Splitting our data into discovery and validation sets, we evaluate all four models' classification performance and the survival implications of their class predictions using cross validation and Kaplan-Meier curves. We find that the One-Versus-Rest model produces the lowest misclassification error of 0.0572 and the lowest median log-rank test of 0.380 statistic measuring the similarity between its Kaplan-Meier curves and the discovery set's true Kaplan-Meier curves. We show that a multinomial classification task split into several LASSO binomial classifiers offers promising results for patient clustering.
Cell cycle (CC) dynamics are reflected in diverse T cell processes such as TCR activation, expansion, contraction, differentiation, senescence, anergy, and exhaustion; linking CC behaviors to functional and dysfunctional T cell states in development and disease. Progression through CC checkpoints is also tightly linked to cell fate decisions across development. Yet, much remains unknown about the connection between CC sensing and T cell differentiation programs. To disentangle the relationship across T cell state, time-since-activation, receptor signaling, division, and CC, we leverage high-throughput single-cell mass cytometry for parallel measurement of these diverse biological states. By modulating CC progression and receptor signaling with inhibitors as well as tonic signaling Chimeric Antigen Receptor (CAR) models of T cell exhaustion, we reveal that earlier G1/S CC programs crosstalk with receptor signaling to control T cell fate, and that exhaustion programs are downstream to aberrant, S-G2 phase CC arrest signatures in tonic CAR signaling in vitro , in situ, and in vivo across human cancers in association with CD8 T-lymphocyte dysfunction.
Cytokine signaling is critical to the function of natural immune cells and engineered immune cell therapies such as chimeric antigen receptor (CAR) T cells. It remains unclear how the limited set of signal transducers and activators of transcription (STATs) and other proteinsa activated by these cytokine receptors can encode the observed diversity of immune cell phenotypes. To understand how signaling downstream of cytokines control immune cell phenotype, we sought to map the structure of Janus kinase (JAK)/STAT signaling domains to cell signaling and resulting CAR T cell function. We recombined 14 signaling motifs to construct a library of ∼30,000 constitutively active synthetic cytokine receptors (SCRs) with intracellular domains composed of novel signaling motif combinations that activate different signaling cascades. We experimentally tested ∼450 SCRs which generated a range of CAR T cell memory, cytotoxicity, and proliferation. SCRs with pSTAT1 and 3 signaling generated effector memory CAR T cells, while SCRs with strong pSTAT5 generated effector CAR T cells with potent anti-tumor activity, measured by flow cytometry and phosphoproteomics. Subsequent kinase substrate enrichment analysis (KSEA) identified key differences of kinases associated with cell signaling networks, such as CDKs that are correlated with high proliferative profile of effector CAR T cells. To map the structure-signaling-phenotype landscape we trained models to predict signaling and CAR T cell phenotype that result from varied motif combinations. From neural network predictions we identified features, including strong STAT5 and Shc signaling, that promote unsafe autonomous CAR T cell proliferation. Models also revealed a trade-off between memory and cytotoxicity, with a Pareto front encoded by a continuous change in signaling. These results demonstrate that recombination of a limited set of signaling motifs creates a continuous spectrum of signaling that encodes a corresponding spectrum of cell phenotype. This work synthetically expands the combinatorial space of JAK/STAT signaling and provides a foundation for rational design of CAR T cells with improved cytotoxicity, memory, and safety profiles. Wansang Cho, Jenny Y. Liu, Alex N. Beckett, Erin Craig, Dain R. Brademan, Judith C. Lunger, Peng Xu, Katie Ho, Ethan E. Chen, Antonio Salcido-Alcantar, Lucas E. Sant'Anna, Kamal Obbad, Nakoa Po, Sophia Joy Ong, Elena Sotillo, Robert Tibshirani, Ruth Huttenhain, Crystal L. Mackall, Kyle G. Daniels. Programmable JAK/STAT signaling drives CAR T cells to enhanced functional states [abstract]. In: Proceedings of the AACR Immuno-Oncology Conference (AACR IO): Discovery and Innovation in Cancer Immunology: Revolutionizing Treatment through Immunotherapy; 2026 Feb 18-21; Los Angeles, CA. Philadelphia (PA): AACR; Cancer Immunol Res 2026;14(2 Suppl):Abstract nr A013.
Motivation Mass spectrometry imaging (MSI) and single-cell RNA sequencing (scRNA-seq) offer high-resolution views into tissue and cellular heterogeneity. However, conventional statistical analyses often treat sub-measurements (pixels or cells) as independent, ignoring their nested origin from individual samples. This assumption inflates the effective sample size, increases false discoveries, and undermines biological interpretation. Pixel- or cell-level averaging avoids this but sacrifices spatial or cellular resolution.Results We evaluated block-SAM on both simulated and real-world datasets. In DESI-MSI data from kidney, lung, and ovarian tumors, block-SAM consistently identified fewer-but more reliable-differential features compared to traditional-SAM. For example, in the kidney dataset, traditional-SAM identified 569 features between tumor and normal tissue, while block-SAM identified 186-all overlapping but excluding 383 likely false positives. Applied to a metastatic RCC scRNA-seq dataset comparing immune checkpoint blockade (ICB)-treated versus untreated patients, traditional-SAM identified over 19,000 differentially expressed genes in malignant cells; block-SAM reduced this to 19. These results demonstrate block-SAM's ability to reduce false discoveries while retaining biologically meaningful signals across diverse high-dimensional datasets.Availability and implementation All code and data associated with this study are deposited on Zenodo (https://doi.org/10.5281/zenodo.18273497). The samr package is freely available on the Comprehensive R Archive Network (CRAN).
Despite being ubiquitous in science, clustering remains a technique whose results are not quantitatively scrutinized via a framework. We present an analysis called evaluating replicability via iterative clustering assignments (ERICA) that is applied to a dataset to determine whether clusters are identified in a replicable manner. The pipeline computes a statistic that describes whether structure is found in a dataset. Quantitative visualization methods are presented to answer important questions such as the similarity between clusters, and the identity of points that may be outliers. When tested on synthetic data, the findings show clusters being discovered in a replicable manner. However, we note a possibility for non-replicable results when the pipeline is applied to three gene expression datasets for breast cancer subtype validation. The study underscores the need for rigorous inspection and offers a practical tool for doing so.
Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling cellular features in microscopy. Realizing this potential, however, requires robust and fine-grained visual perception. Models need to correctly interpret subtle features in images, and they must do so across diverse biomedical modalities, scales, and contexts. Nevertheless, current benchmarks remain limited. To address these gaps, we introduce the Massive Multimodal Biomedical Understanding (MMBU) benchmark. It is the largest biomedical vision and language benchmark to date, covering 35 submodalities with rich structured metadata. It includes both open and closed versions of ungrounded classification, grounded classification, and object detection, enabling systematic evaluation of model performance across biological scales, clinical settings, and imaging modalities. Evaluating 15 open-weight and 2 frontier VLMs, we find that while medical adaptation provides measurable gains for some models, the high accuracy often reported on established benchmarks can mask deficiencies in visual perception and domain generalization.
Oral Cavity Squamous Cell Carcinoma (OCSCC) is a common and deadly form of head and neck cancer which disproportionately afflicts low- and middle-income communities. Since the majority of OCSCC lesions are visible within the oral cavity, visual exam-guided biopsy remains the standard method for detection. However, visual examination can be subjective, biopsy can be painful, and pathological evaluation can be time consuming. With the goal of developing an accessible and noninvasive platform for OCSCC detection, we adapted the MSPen technology as a modular system for direct molecular analysis of OCSCC in vivo. The modular MSPen (Mod-MSPen) system consists of a controller unit and a disposable handheld device that directly contacts tissues, performs gentle liquid extraction of molecules, and collects the sample in a vial for subsequent mass spectrometric analysis. We have used the Mod-MSPen to analyze OCSCC and normal oral tissues spanning various anatomical areas: tongue, buccal mucosa, gingiva, and floor of mouth. The molecular data obtained presented distinct trends in the abundances of metabolites such as glutamate, which was consistently observed in higher relative abundances in OCSCC tissues when compared to normal oral tissues. These trends were corroborated using mass spectrometry imaging, which revealed clear spatio-molecular differences in OCSCC and adjacent normal oral mucosa tissues. Statistical classifiers built from the data provided high performance for disease detection, with >95% agreement with standard histopathology. To evaluate clinical suitability, we also deployed the Mod-MSPen for in vivo analysis of OCSCC and normal oral tissues alongside clinical procedures.
We introduce LLM-Lasso, a novel framework that leverages large language models (LLMs) to guide feature selection in Lasso ℓ_1 regression. Unlike traditional methods that rely solely on numerical data, LLM-Lasso incorporates domain-specific knowledge extracted from natural language, enhanced through a retrieval-augmented generation (RAG) pipeline, to seamlessly integrate data-driven modeling with contextual insights. Specifically, the LLM generates penalty factors for each feature, which are converted into weights for the Lasso penalty using a simple, tunable model. Features identified as more relevant by the LLM receive lower penalties, increasing their likelihood of being retained in the final model, while less relevant features are assigned higher penalties, reducing their influence. Importantly, LLM-Lasso has an internal validation step that determines how much to trust the contextual knowledge in our prediction pipeline. Hence it addresses key challenges in robustness, making it suitable for mitigating potential inaccuracies or hallucinations from the LLM. In various biomedical case studies, LLM-Lasso outperforms standard Lasso and existing feature selection baselines, all while ensuring the LLM operates without prior access to the datasets. To our knowledge, this is the first approach to effectively integrate conventional feature selection techniques directly with LLM-based domain-specific reasoning.
In-context learning with attention enables large neural networks to make context-specific predictions by selectively focusing on relevant examples. Here, we adapt this idea to supervised learning procedures such as lasso regression and gradient boosting, for tabular data. Our goals are to (1) flexibly fit personalized models for each prediction point and (2) retain model simplicity and interpretability. Our method fits a local model for each test observation by weighting the training data according to attention, a supervised similarity measure that emphasizes features and interactions that are predictive of the outcome. Attention weighting allows the method to adapt to heterogeneous data in a data-driven way, without requiring cluster or similarity pre-specification. Further, our approach is uniquely interpretable: for each test observation, we identify which features are most predictive and which training observations are most relevant. We then show how to use attention weighting for time series and spatial data, and we present a method for adapting pretrained tree-based models to distributional shift using attention-weighted residual corrections. Across real and simulated datasets, attention weighting improves predictive performance while preserving interpretability, and theory shows that attention-weighting linear models attain lower mean squared error than the standard linear model under mixture-of-models data-generating processes with known subgroup structure.
We present a scalable framework for computing polygenic risk scores (PRS) in high-dimensional genomic settings using the recently introduced Univariate-Guided Sparse Regression (uniLasso). UniLasso is a two-stage penalized regression procedure that leverages univariate coefficients and magnitudes to stabilize feature selection and enhance interpretability. Building on its theoretical and empirical advantages, we adapt uniLasso for application to the UK Biobank, a population-based repository comprising over one million genetic variants measured on hundreds of thousands of individuals from the United Kingdom. We further extend the framework to incorporate external summary statistics to increase predictive accuracy. Our results demonstrate that uniLasso attains predictive performance comparable to standard Lasso while selecting substantially fewer variants, yielding sparser and more interpretable models. Moreover, it exhibits superior performance in estimating PRS relative to its competitors, such as PRS-CS. Integrating external scores further improves prediction while maintaining sparsity.
In this article, we introduce "uniLasso," a novel statistical method for regression. This two-stage approach preserves the signs of the univariate coefficients and leverages their magnitude. Both of these properties are attractive for stability and interpretation of the model. Through comprehensive simulations and applications to real-world data sets, we demonstrate that uniLasso outperforms lasso in various settings, particularly in terms of sparsity and model interpretability. We prove asymptotic support recovery and mean-squared error consistency under a set of conditions different from the well-known irrepresentability conditions for the lasso. Extensions to generalized linear models (GLMs) and Cox regression are also discussed.
Immune checkpoint inhibition (ICI) has fundamentally changed cancer treatment. However, only a minority of patients with metastatic triple negative breast cancer (TNBC) benefit from ICI, and the determinants of response remain largely unknown. To better understand the factors influencing patient outcome, we assembled a longitudinal cohort with tissue from multiple timepoints, including primary tumor, pre-treatment metastatic tumor, and on-treatment metastatic tumor from 117 patients treated with ICI (nivolumab) in the phase II TONIC trial. We used highly multiplexed imaging to quantify the subcellular localization of 37 proteins in each tumor. To extract meaningful information from the imaging data, we developed SpaceCat, a computational pipeline that quantifies features from imaging data such as cell density, cell diversity, spatial structure, and functional marker expression. We applied SpaceCat to 678 images from 294 tumors, generating more than 800 distinct features per tumor. Spatial features were more predictive of patient outcome, including features like the degree of mixing between cancer and immune cells, the diversity of the neighboring immune cells surrounding cancer cells, and the degree of T cell infiltration at the tumor border. Non-spatial features, including the ratio between T cell subsets and cancer cells and PD-L1 levels on myeloid cells, were also associated with patient outcome. Surprisingly, we did not identify robust predictors of response in the primary tumors. In contrast, the metastatic tumors had numerous features which predicted response. Some of these features, such as the cellular diversity at the tumor border, were shared across timepoints, but many of the features, such as T cell infiltration at the tumor border, were predictive of response at only a single timepoint. We trained multivariate models on all of the features in the dataset, finding that we could accurately predict patient outcome from the pre-treatment metastatic tumors, with improved performance using the on-treatment tumors. We validated our findings in matched bulk RNA-seq data, finding the most informative features from the on-treatment samples. Our study highlights the importance of profiling sequential tumor biopsies to understand the evolution of the tumor microenvironment, elucidating the temporal and spatial dynamics underlying patient responses and underscoring the need for further research on the prognostic role of metastatic tissue and its utility in stratifying patients for ICI.
Rigorous external validation is crucial for assessing the generalizability of prediction models, particularly by evaluating their discrimination (AUROC) on new data. This often involves comparing a new model's AUROC to that of an established reference model. However, many studies rely on arbitrary rules of thumb for sample size calculations, often resulting in underpowered analyses and unreliable conclusions. This paper reviews crucial concepts for accurate sample size determination in AUROC-based external validation studies, making the theory and practice more accessible to researchers and clinicians. We introduce powerROC, an open-source web tool designed to simplify these calculations, enabling both the evaluation of a single model and the comparison of two models. The tool offers guidance on selecting target precision levels and employs flexible approaches, leveraging either pilot data or user-defined probability distributions. We illustrate powerROC's utility through a case study on hospital mortality prediction using the MIMIC database.