Recent advances in spatial transcriptomics (ST) have revolutionized the understanding of cellular functions by providing gene expression profiles with rich spatial context. Effectively learning spatial representations is essential for downstream analyses and requires robust integration of spatial and transcriptomic information. Although existing methods show promise, they often fail to capture local (neighbor-level) and global (tissue-wide) contexts, and methods based on contrastive learning often rely on augmentation strategies that introduce noise and instability. GatorST, a novel and versatile framework, explicitly integrates graph-based modeling with advanced meta-learning strategies to generate spatially informed representations of ST data. Locally, a spot-spot graph connects each node to its nearest neighbors, while two-hop subgraphs capture fine-grained spatial context. Globally, gene expression profiles are clustered to produce pseudo-labels, providing weak supervision for representation learning. An episodic training strategy inspired by meta-learning further enhances GatorST's ability to generalize to new spatial contexts, ensuring robust integration of local and global spatial information. Comprehensive comparisons with fifteen state-of-the-art methods demonstrate that GatorST consistently outperforms existing approaches in identifying spatial domains, imputing gene expression, removing batch effects, and inferring spatial trajectories. By integrating local spatial topology with global gene expression patterns, GatorST provides biologically meaningful representations that advance key downstream analyses.
Most RNA molecules adopt multiple alternative structures, forming dynamic ensembles that cannot be captured by single-structure prediction. Recent advances in chemical probing methods (e.g. DMS-MaPseq and SHAPE-MaP sequencing) now provide single-molecule signals that reflect this structural heterogeneity, enabling computational reconstruction of RNA conformational states. However, existing ensemble-inference approaches based on expectation-maximization (EM) often suffer from instability, convergence to suboptimal local optima, and poor scalability on high-dimensional, sparse mutation matrices, particularly for complex or modification-dependent RNA ensembles. To address these limitations, we developed VIRSE, a variational Bayesian framework that uses coordinate ascent variational inference to achieve efficient, scalable, and noise-robust reconstruction of RNA conformational mixtures from chemical probing data. We evaluated VIRSE using extensive simulations, including mechanism-informed mutation simulations that mimic realistic DMS-MaP-seq behavior (A/C mutation bias, context-dependent dropouts, position-specific mutation rates) and idealized Bernoulli-mixture datasets without experimental artifacts. Across all conditions, especially in high-dimensional and long RNA regimes, VIRSE achieved superior ensemble separation and improved cluster identifiability compared with EM, while maintaining stable posteriors, resolving low-abundance states, and scaling to thousands of nucleotide positions. Applied to experimental datasets, including the human immunodeficiency virus-1 Rev response element, SARS-CoV-2 SHAPE-MaP measurements, and the Escherichia coli mgtL Mg2+-responsive riboswitch, VIRSE successfully recovered biologically meaningful and physically plausible RNA conformational ensembles. VIRSE is freely available at https://github.com/QSong-github/VIRSE.
Systematic characterization of drug-disease relationships is essential for drug discovery and repurposing, yet is hindered by the heterogeneity and rapid growth of biomedical literature. Existing datasets rely on labor-intensive curation and are often incomplete, while LLM-only approaches suffer from hallucination and weak evidence grounding. We introduce UniD^3, a unified framework that integrates Large Language Models with Knowledge Graph-enhanced Retrieval-Augmented Generation (KG-RAG) to extract, organize, and validate drug-disease knowledge across Drug-Disease Matching (DDM), Drug Effectiveness Assessment (DEA), and Drug-Target Analysis (DTA). UniD^3 processes 157,849 PubMed articles with Llama 3.3-70B and constructs knowledge graphs via a dual-stage strategy combining paper-level extraction with KG-level consolidation centered on drug and disease entities. These graphs support KG-RAG-based generation of structured datasets, evaluated through external benchmarks, fuzzy matching with curated resources, and clinician review. UniD^3 produces six knowledge graphs and large-scale datasets, including 28,915 DDM, 15,042 DEA, and over 4,000 DTA QA pairs. External validation shows strong performance (F1: 0.85-0.87 for DDM/DEA; 0.82 for DTA), with clinician review confirming high reliability (AUROC = 0.90). KG-RAG-augmented models outperform standalone LLMs, and the UniD^3 chatbot enables interpretable, citation-supported exploration of drug-disease relationships. UniD^3 provides a scalable, extensible framework for transforming unstructured biomedical literature into high-quality, structured drug-disease knowledge, supporting AI-driven discovery, repurposing, and precision medicine.
Mitochondrial genetic heterogeneity arises from the accumulation of somatic mitochondrial DNA (mtDNA) mutations within individual cells, generating intracellular clonal populations whose selective dynamics in disease remain poorly characterized. Here, we present MitoBayes, a hierarchical Bayesian framework that jointly models mitochondrial clonal lineage structure, allele frequency variation, and single-cell disease-relevant phenotypic burdens to infer clone-specific selection pressures. Extensive benchmarking demonstrates that MitoBayes accurately recovers ground-truth selection coefficients across a wide range of genetic heterogeneity, data sparsity, and lineage complexity scenarios. Application of MitoBayes to single-cell atlases of Alzheimer's disease (AD) cortex, treatment-naïve non-small-cell lung cancer (NSCLC), and chemotherapy-resistant small-cell lung cancer (SCLC) revealed distinct, disease-specific patterns of mitochondrial clonal selection. These include selective expansion of high-risk mitochondrial clones associated with disruption of PVALB interneuron homeostasis in AD; disease-driven clonal remodeling in cycling T/NK cells from NSCLC tumors characterized by increased mitochondrial biogenesis and impaired immune regulatory programs; and preferential enrichment of a tumor-associated MT-ATP6 (m.8859A>G) clone linked to metabolic adaptation and platinum resistance in SCLC. Pan-cancer survival analyses further confirmed the clinical relevance of elevated MT-ATP6 activity, which was associated with inferior chemotherapy outcomes. Additionally, in hepatocellular carcinoma (HCC), a dominant m.2356C>G clone correlated with POLR2A activation and widespread transcriptional amplification, consistent with a mitochondria-nucleus signaling axis contributing to adverse prognosis in this cancer type. Collectively, these findings establish MitoBayes as a robust statistical framework linking mitochondrial genetic diversity to disease phenotypes and highlight mitochondrial clonal selection as a mechanistically and clinically actionable target for therapeutic and diagnostic development.
Metastatic brain disease occurs in up to 30% of patients with lung, melanoma and breast cancers, and the median survival time remains less than a year. Treating these patients is a challenge because surgical approaches are limited and most chemotherapeutic drugs and immunotherapies are ineffective at crossing the blood-brain barrier (BBB). Given the unique abilities of macrophages to cross the BBB and exert their phagocytic function on tumour cells, we genetically engineer macrophages that express a chimaeric antigen receptor (CAR) targeting mesothelin (MSLN). To specifically target metastatic brain tumours, we fused the cells with the immune signalling molecule MyD88. This chimaeric antigen receptor macrophage (CARMA) penetrates the BBB and decreases brain metastasis growth in a humanized mouse model. MSLN-CARMA shows antigen-specific phagocytosis activity against tumour cells and exhibits a bystander effect by releasing TNF to act on surrounding tumour cells lacking the tumour antigen. These features of CARMA represent advantages over other immune therapies and CARMA may serve as a promising therapeutic tool for the treatment of brain metastasis.
Abstract The Florida Cancer Research (FL CARES) Network (https://floridacancernetwork.org/) is a statewide collaboration advancing cancer research and improving patient outcomes through coordinated research initiatives and a focus on understanding population-health variation across demographic categories. Funded by the State of Florida Department of Health’s Bankhead-Coley Cancer Research Program, FL CARES unites six research organizations, including the state’s three NCI-designated Cancer Centers. The Network has established a collaborative research consortium with unified policies, coordinated data and metadata standards, and diverse datasets and analysis tools.To accelerate computational research, FL CARES is developing the Platform for Accelerating Collaborative Computational Cancer Research (PAC3R) (https://pac3r.floridacancernetwork.org). This advanced informatics system enables FAIR (findable, accessible, interoperable, reusable) data management, including standardization, harmonization, and integration of multimodal cancer-related datasets. PAC3R is designed to support secure data sharing, deployment of scalable bioinformatics tools, and collaborative analyses across the FL CARES Network.PAC3R builds upon the Sylvester Data Portal (SDP) (https://sdp.miami.edu/), a cloud-based multi-omics platform that manages clinicogenomic and research data for the Sylvester Comprehensive Cancer Center. PAC3R integrates diverse multimodal cancer datasets to support analyses such as transcriptional perturbation signatures, cell sensitivity data, and small-molecule interactions, alongside local cancer panels from Moffitt Cancer Center and Nova Southeastern University. The platform also incorporates large-scale public genomics resources including The Cancer Genome Atlas (TCGA) and the Clinical Proteomic Tumor Analysis Consortium (CPTAC).Overall, the FL CARES Network and its advanced informatics system, the PAC3R platform, demonstrate the power of a unified semantic data model, harmonized datasets, and computational strategies to detect and understand outcome variations across Florida’s populations, supporting the long-term goal of advancing effective cancer prevention, diagnosis, and treatment for all. Citation Format: Jeronimo Pissinis, Michael S. Sinclair, Marcin Pilarczyk, Caty Chung, Dusica Vidovic, Lukas Rupprecht, Franklin Sotolongo, Oliver Mazariegos, Kathleen M. Jagodnik, Carlos Obregon, Tingyi Li, Ling Cen, Jiang Bian, Ji-Hyun Lee, Qianqian Song, Bikhyat Adhikari, Till Krenz, Ritik Bhandari, Umamaheswari Natarajan, Gogce C. Crynen, Appu Rathinavelu, Stuart Chalk, Mondal M. Ananda, Xuefeng Wang, Vasileios Stathias, Stephan Schürer. Florida Cancer Research (FL CARES) network and the Platform for Accelerating Collaborative Computational Cancer Research (PAC3R) [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 2732.
ABSTRACT Drugs induce coordinated phenotypic changes across multiple modalities, including transcriptional reprogramming and cellular morphological remodeling. Predicting these drug-induced modality changes is central to drug discovery, mechanism-of-action studies and precision therapeutics, however, prediction performance depends critically on how both drug compounds and cellular states are represented. Despite rapid advances in drug molecular and gene representation methods, a systematic evaluation of these methods remains lacking. Herein, we introduce MVCBench, a comprehensive benchmarking framework for evaluating drug molecular and gene representation methods in predicting drug-induced multimodal virtual cell (MVC) phenotypes. MVCBench leverages large-scale transcriptomic and high-content imaging data and systematically evaluates 24 representation methods (12 drug molecular and 12 gene representation methods) across nearly 1.1 million drug-induced profiles, under both in-distribution and out-of-distribution settings spanning unseen compounds, cell lines, assay plates and datasets. Our benchmarking reveals a pronounced modality-dependent asymmetry: advanced drug molecular representations substantially improve the prediction of drug-induced morphological phenotypes but provide only limited gains for gene expression prediction relative to classical fingerprints, whereas task-specific gene representations outperform general-purpose foundation models in predicting drug-induced transcriptomic responses. Predictive performance also deteriorates sharply under distribution shift, highlighting persistent challenges in cross-dataset and cross-platform generalization. We further show that integrating transcriptomic and morphological modalities consistently improves prediction accuracy, and derive practical design principles for MVC architectures, including modality-aware loss calibration and fusion strategies. Together, MVCBench provides a systematic foundation for evaluating representation methods and offers guidance for developing robust MVC models of drug-induced cellular responses.
Importance:BRCA genetic testing is critical for cancer risk assessment, treatment and personalization, yet substantial underutilization persists. Socioeconomic and clinical factors may strongly influence testing uptake; therefore, identifying the potential drivers to BRCA testing and treatment is essential for addressing gaps in access, increasing retention into care, and improving cancer outcomes. Objective:To quantify the putatively causal effects of SDoH on BRCA genetic testing among individuals with breast, ovarian, pancreatic, and prostate cancers and to develop a predictive model to identify patients at risk for underuse of testing. Design Setting and Participants:This observational case-control study used data from a large multistate clinical research data network covering southern US (2012-2023). The network contained records of more than 26 million individuals and was linked with ZIP code-level SDoH variables derived from national socioeconomic datasets. Adults diagnosed with breast, ovarian, pancreatic, or prostate cancer were eligible for cases (received BRCA testing) or controls (did not receive BRCA testing, matched by cancer diagnosis). Exposure:SDoH categories, including economic conditions, education, healthcare access, neighborhood conditions, and social connectedness. Main Outcomes and Measures:The primary outcome was receipt of BRCA genetic testing after cancer diagnosis. Results:Among 3,279 people diagnosed with cancer, 748 received BRCA testing and 2,531 served as controls. Study population's mean [SD] age was 66.8 [15.7] years; 1,758 were women [53.6%], 2,238 [69.6%] were White and 616 [18.8%] were Black or African American. Breast (1,420 [42.8%]) and prostate (1,342 [40.9%]) cancers were the most common diagnoses, followed by pancreatic (242 [7.4%]), ovarian (238 [7.2%]), and multiple cancers (55 [1.7%]). Upon adjusting for potential confounding, higher educational attainment (odds ratio [OR], 1.19), public-sector employment (OR, 1.42), neighborhood safety (OR, 1.28), and social participation (OR, 1.72) showed an increased likelihood of undergoing BRCA testing, whereas economic instability, including housing cost burden and reliance on public insurance, had an effect of reduced testing. A random forest classifier demonstrated good discriminative performance (AUROC, 0.776) to predict cancer patients who were likely to take BRCA testing, where nativity, language, and residential stability ranked among the most influential social determinants according to SHapley Additive exPlanations (SHAP) analysis. Conclusions and Relevance:In this observational case-control study, SDoHs were strongly associated with receipt of BRCA genetic testing among people with cancer. These findings suggest that disparities in genetic testing may reflect structural and social barriers rather than differences in clinical eligibility alone. Efforts to improve equitable access to genetic testing may benefit from integrating social-context information into clinical workflows and targeting outreach or navigation strategies toward socially disadvantaged populations. Key Points:Question: Do socioeconomic and clinical circumstances contribute to care access gaps to BRCA genetic testing among people with breast, ovarian, prostate, and pancreatic cancers? Findings: In a case-control study of 3,279 adults with cancer from a large multi-state US clinical research data network, multiple social determinants of health, including economic conditions, education, healthcare access, neighborhood conditions, and social connectedness, were identified as key determinants of BRCA genetic testing uptake, with variation across age and sex groups. Meaning: Improving equal access to BRCA genetic testing may require interventions that address financial barriers and strengthen social, and community supports that influence engagement with genetic services.
IntroductionProtein-DNA interactions are central to gene regulation, genome stability, and disease mechanisms. Identifying DNA-binding residues (DBRs) is critical for structural modeling, protein engineering, and therapeutic design. Although experimental approaches provide valuable insights, they remain low-throughput and resource-intensive. Computational methods offer scalable alternatives by leveraging protein sequential and structural information to predict DBRs.MethodsWe present PRIMED (Protein Residue Inference using Multilayer perceptron for Enhanced DNA-binding predictions), a machine learning framework that integrates protein representations of distinct biochemical and structural properties from three protein language models: ESM-2, ESM-3, and ESM-C. These representations are concatenated and processed by a multilayer perceptron to perform DBR predictions.ResultsPRIMED demonstrated strong performance across three benchmark datasets: Test-46 and Test-129 from a previous study, CLAPE-DB, and Test-10 K, which we curated from UniProtKB/Swiss-Prot. The model achieves an area under the Receiver Operating Characteristic curve (AUC) of 0.92 and a Matthews Correlation Coefficient (MCC) of 0.64 on Test-46, as well as an AUC of 0.93 and MCC of 0.45 on Test-129. On Test-10 K, PRIMED demonstrates generalizability across proteins with varying DBR percentages, maintaining competitive performance relative to the runner-up method, CLAPE-DB.DiscussionThese results highlight the effectiveness of integrating diverse protein language model representations for accurate, transferable DBR predictions.
Objective:Predicting health outcomes from electronic health records (EHRs) is challenging because traditional models rely on structured data and often ignore external medical knowledge. We propose an approach that integrates structured EHR with text‑based clinical evidence to improve prediction and interpretability. Methods:We introduce PHO-Agents, a multi-agent system powered by large language models (LLMs) for health outcome prediction. Structured EHR sequences are encoded to produce attention-based representations and initial logits, which are converted into patient summaries by a data agent. A retrieval agent gathers relevant clinical guidelines. Research and practical doctor agents independently assess the patient, and a leader agent synthesizes their analyses. Outputs from the EHR-based model and the LLM agents are fused to generate final predictions and explanation reports. PHO-Agents was evaluated on three real-world cohorts: acute kidney injury (AKI) patients (in-hospital mortality), chronic kidney disease patients (AKI onset within two years), and cancer patients receiving immune checkpoint inhibitors (immune-related adverse events within one year). Results:PHO-Agents outperformed single-agent and multi-agent LLM baselines across all cohorts. In the AKI mortality task, it achieved a PR-AUC of 90.20 ± 2.07, compared with 56.46 ± 2.98 for the best single-agent baseline. Similar gains were observed in the ICI and CKD cohorts. Ablation studies showed that both multi-agent reasoning and logit-level fusion contributed to performance improvements, and case analyses demonstrated clinically consistent explanations. Conclusion:PHO-Agents integrates longitudinal EHR modeling with collaborative LLM reasoning, improving predictive performance, interpretability, and robustness across diverse clinical tasks. This hybrid approach offers a trustworthy strategy for real-world clinical decision support.
RNA secondary structure plays a critical role in gene regulation, yet existing computational and experimental tools for structure analysis are often fragmented across prediction, ensemble modeling, and functional interpretation workflows. Here, we present ShapeRNA, a user-friendly web server for integrated RNA secondary structure prediction, ensemble inference, and structure-aware regulatory annotation. ShapeRNA supports three complementary analytical workflows, including sequence-based structure prediction, reactivity-guided modeling using SHAPE or DMS data, and sequencing-guided ensemble inference from high-throughput probing experiments. The platform integrates multiple established prediction algorithms and provides standardized data processing, ensemble clustering, and visualization. In addition, ShapeRNA enables mapping of RNA modification sites, microRNA target regions, and RNA-binding protein interaction motifs onto predicted RNA structures and representative ensemble conformations. We demonstrate the utility of ShapeRNA through applications including analysis of mutation-associated structural changes in MAPT exon 10, characterization of conformational heterogeneity in the HIV-1 Rev Response Element, and regulatory annotation of the oncogenic long non-coding RNA HULC. ShapeRNA provides an accessible and extensible platform for investigating RNA structural heterogeneity and regulatory mechanisms. This website is free and open to all users, and there is no login requirement. The server is accessible at https://shaperna.com.
Alzheimer's disease (AD) has a strong genetic predisposition. Genome-wide association studies have identified multiple risk loci, yet many non-coding variants remain uncharacterized. Machine learning-based polygenic risk scores (PRS) enhance prediction by modeling genetic epistasis and sex-specific risks. This review summarizes AD genetic risk factors, PRS methodologies, and ML-based AD risk prediction. It also highlights challenges such as population bias, functional validation, and integrating multi-omics for precision medicine.
Importance:Combinations of VEGFR tyrosine kinase inhibitors (TKIs) and immune checkpoint inhibitors (ICIs), such as antibodies to programmed cell death-1 (PD-1), or to its ligand PD-L1, are now first-line standard of care for renal cell carcinoma (RCC), but the pivotal clinical trials excluded patients with common comorbidities, leaving their real-world effectiveness uncertain. Objective:To determine whether adding PD-1/PD-L1 inhibitors to VEGFR-TKIs therapy is associated with improved overall survival in a real-world RCC cohort. Design Setting and Participants:This retrospective cohort study used a target trial emulation framework and real-world electronic health records data from the University of Florida Health Integrated Data Repository (IDR). Data was analyzed from September 2009 through June 2023. Adult patients (≥18 years) with confirmed RCC and at least one VEGFR-TKIs prescription were eligible. The date of the first VEGFR-TKIs prescription was defined as the index date, and patients were followed for up to 24 months. Variable-ratio propensity score matching (up to 2:1) across 13 baseline covariates was used to emulate randomized treatment assignments. Of 107,783 patients screened, 387 met eligibility criteria, and 319 remained in the matched cohort. Exposures:VEGFR-TKIs monotherapy (control group) versus VEGFR-TKIs combined with PD-1/PD-L1 inhibitors (experimental group). Main Outcomes and Measures:Overall survival (OS), analyzed by weighted Kaplan-Meier estimation, cluster-robust Cox regression, and restricted mean survival time (RMST) at τ = 24 months, prespecified given anticipated non-proportional hazards. Results:Among 319 matched patients (mean [SD] age, 62 [12] years; 76% male), 107 deaths occurred (33.5%). Twelve-month OS was higher in the combination arm (81.8%; 95% CI, 74.7-89.6%) than VEGFR-TKIs monotherapy (68.1%; 95% CI, 61.1-76.0%), converging by 24 months (61.1% vs 56.7%). The Cox hazard ratio was 0.718 (95% CI, 0.484-1.064; P = 0.0986). RMST was 2.79 months greater with combination therapy (95% CI, 0.93-4.65; P = 0.0033). Conclusions:Adding PD-1/PD-L1 inhibitors to VEGFR-TKIs therapy was associated with a statistically significant and clinically meaningful gain in restricted mean survival, supporting the real-world generalizability of combination therapy and the importance of appropriate treatment effect measures under non-proportional hazards.
Alternative polyadenylation (APA) is a widespread post-transcriptional regulatory mechanism that diversifies transcript isoforms and modulates mRNA stability, localization, and translation. Although single-cell RNA sequencing (scRNA-seq) provides an unprecedented opportunity to study cell-type-specific APA dynamics, existing computational tools are largely designed for bulk RNA-seq data or rely heavily on gene annotations, limiting their applicability to single-cell contexts. Here, we present scDeepAPA, a deep learning framework specifically optimized for scRNA-seq data to enable accurate polyadenylation site (PAS) detection, isoform quantification, and functional interpretation of APA events at single-cell resolution. Trained on high-confidence annotations from PolyASite v3.0, scDeepAPA integrates convolutional feature extraction with Mamba-based state-space modeling and bidirectional LSTM layers to capture both long-range and local sequence dependencies. Comprehensive benchmarking against five state-of-the-art PAS prediction models demonstrates that scDeepAPA consistently achieves superior performance across accuracy, F1 score, and area under the receiver operating characteristic metrics in both human and mouse datasets. Applying scDeepAPA to Alzheimer's disease mouse brain data revealed widespread, cell-type-specific APA remodeling across immune and glial populations, including shifts toward proximal PAS usage and 3 ' UTR shortening. In KRAS-mutant small cell lung cancer, scDeepAPA uncovered global proximal PAS activation and tumor-specific intronic polyadenylation events. Notably, several intronic APA events generated truncated transcripts encoding predicted neoantigenic peptides with strong major histocompatibility complex class I binding affinity, supported by structural modeling and tumor-specific expression patterns. By enabling accurate PAS identification and quantitative APA profiling, scDeepAPA facilitates in-depth downstream analyses of regulatory mechanisms and immunogenic consequences in single-cell transcriptomics, advancing the understanding of post-transcriptional regulation in neurodegeneration and cancer.
Integrating transcriptome-wide single-cell gene expression data with spatial context significantly enhances our understanding of tissue biology, cellular interactions, and disease progression. Although single-cell RNA sequencing (scRNA-seq) provides high-resolution gene expression data, it lacks crucial spatial context, whereas spatial transcriptomics techniques offer spatial resolution but are limited in the transcriptomic coverage. To address these limitations, integrating scRNA-seq and spatial transcriptomics data is essential. We introduce SpaGene, a novel deep learning framework designed to integrate scRNA-seq data and spatial transcriptomics data. SpaGene consists of two encoder-decoder pairs combined with two translators and two discriminators to effectively impute missing gene expressions within spatial transcriptomics datasets. We benchmarked SpaGene against existing state-of-the-art methods across diverse datasets. Across the datasets, SpaGene achieved an average 33% higher Pearson correlation coefficient (PCC), 21% higher Structural similarity index (SSIM), and 6.6% lower Root mean squared error (RMSE) compared to the existing approaches, highlighting its capability to reliably impute missing genes and provide comprehensive transcriptomics profiles. Application of our model to lung tumor tissue revealed immune cell enrichment at tumor boundaries, restricted myeloid cell trafficking in adjacent normal regions, and microenvironmental-driven pathways linked to immune neighborhoods. These results provide novel insight into immune exclusion and tumor-immune interactions that drive tumor progression, highlighting potential avenues for therapeutic development. Thus, SpaGene extends the power of spatial transcriptomics by delivering spatially resolved, enhanced transcriptome data that enable deeper biological understanding.
Drug-information question answering is a high-stakes setting where hallucinated facts can mislead clinical decision-making and the provenance of each cited fact matters as much as the fact itself. We present DrugClaw, a multi-agent retrieval-augmented system that queries a registry of drug and pharmacovigilance skills via a reflection-driven state-machine workflow and returns answers grounded in primary regulatory or peer-reviewed records. We also contribute DrugAudit, a 3,772-item authority-aware benchmark with an evaluation panel that scores upstream-of-gold source match, token-level semantic snippet overlap, and citation faithfulness under a dual-judge LLM-as-judge protocol with inter-judge kappa = 0.88 (almost-perfect). Across DrugAudit plus drug-related subsets of MedQA (751) and PubMedQA (512), DrugClaw is top-1 on every column of the headline table: composite Evidence Index under both judges, judge-mediated answer correctness, primary-source rate (0.918, +10.1 pp over next-best), faithfulness (0.887, +5.9 pp), MedQA (0.920), and PubMedQA (0.693).
Background Alzheimer's Disease (AD) is a complex neurodegenerative disorder, with women comprising nearly two-thirds of individuals with AD. However, sex-specific heterogeneity in AD progression remains insufficiently understood. A data-driven approach is needed to characterise such heterogeneity from longitudinal electronic health records (EHRs). Methods We developed a deep learning-based framework to uncover sex-specific AD sub-phenotypes using longitudinal EHRs from OneFlorida+ Clinical Research Consortium. We constructed temporal representations of these EHRs and employed an autoencoder architecture to generate latent embeddings, followed by clustering to derive sex-specific sub-phenotypes with associated progression patterns. We also performed statistical and survival analyses to unravel the characteristics of our identified sub-phenotypes. Findings From 1665 individuals with AD (961 females, 704 males), we identified five major sex-specific sub-phenotypes of AD with distinct progression pathways and comorbidity patterns. Female-dominant sub-phenotypes presented later AD onset, longer disease duration, and enrichment of respiratory and neurological disorders. Male-dominant sub-phenotypes exhibited earlier onset, shorter duration, and higher prevalence of endocrine and metabolic conditions. Survival analysis showed significant differences in time to AD onset across sub-phenotypes. Interpretation Our findings revealed distinct disease trajectories and comorbidity patterns between male- and female-dominant subgroups with AD. This study provides insight into sex-specific AD progression and demonstrates a data-driven framework for characterising disease heterogeneity using longitudinal EHRs. Funding This study was supported by grants from the Florida Department of Health, the Centers for Disease Control and Prevention, the National Institute of Environmental Health Sciences, and the NIH National Center for Advancing Translational Sciences.
Alzheimer's Disease (AD) is a complex neurodegenerative disorder strongly influenced by sex differences, with women comprising nearly two-thirds of cases. However, sex-specific progression patterns remain underexplored due to unclear clinical and molecular mechanisms. To address this gap, we developed a temporal autoencoder framework to identify sex-specific AD sub-phenotypes using longitudinal electronic health record (EHR) data from the OneFlorida+ Clinical Research Consortium. Sequential EHRs were encoded into latent representations and clustered to derive disease states, which were assembled into progression pathways. This approach uncovered five primary sex-stratified sub-phenotypes with distinct trajectories and phenotypic characteristics. Survival and cumulative prevalence analyses further revealed heterogeneous temporal dynamics of AD onset and comorbidity accumulation between female- and male-dominant groups. By integrating deep learning with large-scale real-world data, our framework advances understanding of sex-based heterogeneity in AD progression and provides a scalable tool for early risk stratification, personalized intervention, and improved clinical trial design.
In image-based drug discovery, accurately capturing cellular phenotypic responses to chemical perturbations is crucial for understanding drug mechanisms and predicting efficacy. However, existing approaches often depend on complex, multi-step pipelines that are computationally intensive and prone to error. PhenoProfiler addresses these challenges with an efficient, end-to-end deep learning framework that directly transforms high-content, multi-channel cellular images into low-dimensional quantitative representations. Evaluated on nearly 400,000 high-content images and 8.42 million single-cell images, PhenoProfiler consistently outperforms state-of-the-art methods by up to 20% in both accuracy and robustness. Its tailored phenotype correction strategy further emphasizes treatment-induced variations, improving the detection of biologically meaningful and reproducible signals. PhenoProfiler also effectively clusters treatments with shared molecular pathways and biological annotations, facilitating mechanistic interpretation and target discovery. Collectively, PhenoProfiler establishes a scalable, interpretable, and generalizable framework for high-throughput phenotypic profiling, paving the way for next-generation AI-driven drug screening, precision therapeutics, and systems-level understanding of cellular responses.