Image classification plays a pivotal role in analyzing biomedical images, serving as a cornerstone for both biological research and clinical diagnostics. We demonstrate that large multimodal models (LMMs), like GPT-4, excel in one-shot learning, generalization, interpretability, and text-driven image classification across diverse biomedical tasks. These tasks include the classification of tissues, cell types, cellular states, and disease status. LMMs stand out from traditional single-modal classification approaches, which often require large training datasets and offer limited interpretability.
Large language models (LLMs) show promise in biomedical research, but their effectiveness for genomic inquiry remains unclear. We developed GeneTuring, a benchmark consisting of 16 genomics tasks with 1600 curated questions, and manually evaluated 48 000 answers from 10 LLM configurations, including GPT-4o (via API, ChatGPT with web access, and a custom Generative Pretrained Transformer (GPT) setup), GPT-3.5, Claude 3.5, Gemini Advanced, GeneGPT (both slim and full), BioGPT, and BioMedLM. A custom GPT-4o configuration integrated with National Center for Biotechnology Information (NCBI) Application Programming Interfaces (APIs), developed in this study as SeqSnap, achieved the best overall performance. GPT-4o with web access and GeneGPT demonstrated complementary strengths. Our findings highlight both the promise and current limitations of LLMs in genomics, and emphasize the value of combining LLMs with domain-specific tools for robust genomic intelligence. GeneTuring offers a key resource for benchmarking and improving LLMs in biomedical research.
Single-cell multi-omics is a transformative technology that measures both gene expression and chromatin accessibility in individual cells. However, most studies concentrate on a single tissue and are unable to determine whether a gene is regulated by a cis-regulatory element (CRE) in just one tissue or across multiple tissues. We developed Compass for comparative analysis of gene regulation across a large number of human and mouse tissues. Compass consists of a database, CompassDB, and an open-source R software package, CompassR. CompassDB contains processed single-cell multi-omics data of more than 2.8 million cells from hundreds of cell types. Building upon CompassDB, CompassR enables visualization and comparison of gene regulation across multiple tissues. We demonstrated that CompassR can identify CRE-gene linkages specific to a tissue type and their associated transcription factors in real examples.
The role of glioma-associated myeloid cells in tumor growth and immune evasion remains poorly understood. We performed single-cell RNA sequencing of immune and tumor cells from 33 gliomas, identifying two distinct myeloid-derived suppressor cell (MDSC) populations in isocitrate dehydrogenase-wild-type (IDT-WT) glioblastoma: an early progenitor MDSC (E-MDSC) population with up-regulation of metabolic and hypoxia pathways and a monocytic MDSC (M-MDSC) population. Spatial transcriptomics demonstrated that E-MDSCs geographically colocalize with metabolic stem-like tumor cells in the pseudopalisading region. Ligand-receptor analysis revealed cross-talk between these cells, where glioma stem-like cells produce chemokines attracting E-MDSCs, which in turn produce growth factors for the tumor cells. This interaction is absent in IDH-mutant gliomas, associated with hypermethylation and repressed gene expression of MDSC-attracting chemokines. Our study elucidates specific MDSCs that may facilitate glioblastoma progression and mediate tumor immunosuppression.
A critical area of recent cancer research is the emergence of transition states between normal and cancer that exhibit increased cell plasticity which underlies tumor cell heterogeneity. Pancreatic ductal adenocarcinoma (PDAC) can arise from the combination of a transition state termed acinar-to-ductal metaplasia (ADM) and a gain-of-function mutation in the proto-oncogene KRAS. During ADM, digestive enzyme-producing acinar cells acquire a transient ductal epithelium-like phenotype while maintaining their geographical acinar organization. One route of ADM initiation is the overexpression of the Krüppel-like factor 4 gene (KLF4) in the absence of oncogenic driver mutations. Here, we asked to what extent cells acquire and retain an epigenetic memory of the ADM transition state in the absence of oncogene mutation. We profiled the DNA methylome and transcriptome of KLF4-induced ADM in transgenic mice at various timepoints during and after recovery from ADM. We validated the identified DNA methylation and transcriptomic signatures in the widely used caerulein model of inducible pancreatitis. We identified differential DNA methylation at Kras-downstream PI3K and Rho/Rac/Cdc42 GTPase pathway genes during ADM, as well as a corresponding gene expression increase in these pathways. Importantly, differential methylation persisted after gene expression returned to normal. Caerulein exposure, which induces widespread digestive system changes in addition to ADM, showed similar changes in DNA methylation in ADM cells. Regions of differential methylation were enriched for motifs of KLF and AP-1 family transcription factors, as were those of human pancreatic intraepithelial neoplasia (PanIN) samples, demonstrating the relevance of this epigenetic transition state memory in human carcinogenesis. Finally, single-cell spatial transcriptomics revealed that these ADM transition cells were enriched for PI3K pathway and AP1 family members. Our comprehensive study of DNA methylation in the acinar-ductal metaplasia transition state links epigenetic memory to cancer-related cell plasticity even in the absence of oncogenic mutation.
Foundation models exhibit strong capabilities for downstream tasks by learning generalized representations through self-supervised pre-training on large datasets. While several foundation models have been developed for single-cell RNA-seq (scRNA-seq) data, there is still a lack of models specifically tailored for single-cell ATAC-seq (scATAC-seq), which measures epigenetic information in individual cells. The principal challenge in developing such a model lies in the vast number of scATAC peaks and the significant sparsity of the data, which complicates the formulation of peak-to-peak correlations. To address this challenge, we introduce EpiFoundation, a foundation model for learning cell representations from the high-dimensional and sparse space of peaks. EpiFoundation relies on an innovative cross-modality pre-training procedure with two key technical innovations. First, EpiFoundation exclusively processes the non-zero peak set, thereby enhancing the density of cell-specific information within the input data. Second, EpiFoundation utilizes dense gene expression information to supervise the pre-training process, aligning peak-to-gene correlations. EpiFoundation can handle various types of downstream tasks, including cell-type annotation, batch correction, and gene expression prediction. To train and validate EpiFoundation, we curated MiniAtlas, a dataset of 100,000+ single cells with paired scRNA-seq and scATAC-seq data, along with diverse test sets spanning various tissues and cell types for robust evaluation. EpiFoundation demonstrates state-of-the-art performance across multiple tissues and diverse downstream tasks.
Spatially variable genes (SVGs) reveal the molecular and functional heterogeneity of cells across different spatial regions of a tissue. Sample-wide SVGs identified by existing methods largely overlap with cell-type marker genes derived from single-cell gene expression, leaving the spatial location information largely underutilized. We develop ctSVG, a computational method specifically tailored for Visium HD spatial transcriptomics at single-cell resolution. We show that cell-type-specific SVGs identified by ctSVG include many new genes that do not overlap with sample-wide SVGs or cell-type marker genes and that these genes reveal important biological functions in real spatial datasets.
DNA methylation (DNAm), an epigenetic modification, regulates gene expression, influences phenotypes, and encodes inheritable information, making it critical for disease diagnosis, treatment, and prevention. While human genome contains approximately 28 million CpG sites where DNAm can be measured, only 1-3% of these sites are typically available in most datasets due to complex experimental protocols and high costs, hindering insights from DNAm data. Leveraging the relationship between gene expression and DNAm offers promise for computational inference, but existing statistical, machine learning, and masking-based generative Transformers face critical limitations: they cannot infer DNAm at unmeasured CpGs or in new samples effectively. To overcome these challenges, we introduce MethylProphet, a gene-guided, context-aware Transformer model designed for DNAm inference. MethylProphet employs a Bottleneck MLP for efficient gene profile compression and a specialized DNA sequence tokenizer, integrating global gene expression patterns with local CpG context through a Transformer encoder architecture. Trained on whole-genome bisulfite sequencing data from ENCODE (1.6B training CpG-sample pairs; 322B tokens), MethylProphet demonstrates strong performance in hold-out evaluations, effectively inferring DNAm for unmeasured CpGs and new samples. In addition, its application to 10842 pairs of gene expression and DNAm samples at TCGA chromosome 1 (450M training CpGsample pairs; 91B tokens) highlights its potential to facilitate pan-cancer DNAm landscape inference, offering a powerful tool for advancing epigenetic research and precision medicine. All codes, data, protocols, and models are publicly available via https://github.com/xk-huang/methylprophet/ .
Here we demonstrate that the large language model GPT-4 can accurately annotate cell types using marker gene information in single-cell RNA sequencing analysis. When evaluated across hundreds of tissue and cell types, GPT-4 generates cell type annotations exhibiting strong concordance with manual annotations. This capability can considerably reduce the effort and expertise required for cell type annotation. Additionally, we have developed an R software package GPTCelltype for GPT-4’s automated cell type annotation.
Successful pancreatic ductal adenocarcinoma (PDAC) immunotherapy requires therapeutic combinations that induce quality T cells. Tumor microenvironment (TME) analysis following therapeutic interventions can identify response mechanisms, informing design of effective combinations. We provide a reference single-cell dataset from tumor-infiltrating leukocytes (TILs) from a human neoadjuvant clinical trial comparing the granulocyte-macrophage colony-stimulating factor (GM-CSF)-secreting allogeneic PDAC vaccine GVAX alone, in combination with anti-PD1, or with both anti-PD1 and CD137 agonist. Treatment with GVAX and anti-PD-1 led to increased CD8+ T cell activation and expression of cytoskeletal and extracellular matrix (ECM)-interacting components. Addition of CD137 agonist increased abundance of clonally expanded CD8+ T cells and increased immunosuppressive TREM2 signaling in tumor associated macrophages (TAMs), identified by comparison of ligand-receptor networks, corresponding to changes in metabolism and ECM interactions. These findings associate therapy with GVAX, anti-PD1, and CD137 agonist with enhanced CD8+ T cell function while inducing alternative immunosuppressive pathways in patients with PDAC.
The analysis of single-cell RNA-sequencing (scRNA-seq) data with multiple biological samples remains a pressing challenge. We present MUSTARD, a trajectory-guided dimension reduction method for multi-sample multi-condition scRNA-seq data. This all-in-one decomposition reveals major gene expression variation patterns along the trajectory and across multiple samples simultaneously, providing opportunities to discover sample endotypes along with associated genes and gene modules. In data-driven simulation, MUSTARD achieves high accuracy in distinguishing sample-level group differences that existing methods fail to capture. MUSTARD also demonstrates a robust ability to capture gene markers and pathways associated with phenotypes of interest across multiple real-world case studies. ### Competing Interest Statement The authors have declared no competing interest.
The performance of seven large language models (LLMs) in generating programming code using various prompt strategies, programming languages, and task difficulties is systematically evaluated. GPT-4 substantially outperforms other LLMs, including Gemini Ultra and Claude 2. The coding performance of GPT-4 varies considerably with different prompt strategies. In most LeetCode and GeeksforGeeks coding contests evaluated in this study, GPT-4, employing the optimal prompt strategy, outperforms 85 percent of human participants in a competitive environment, many of whom are students and professionals with moderate programming experience. GPT-4 demonstrates strong capabilities in translating code between different programming languages and in learning from past errors. The computational efficiency of the code generated by GPT-4 is comparable to that of human programmers. GPT-4 is also capable of handling broader programming tasks, including front-end design and database operations. These results suggest that GPT-4 has the potential to serve as a reliable assistant in programming code generation and software development. A programming assistant is designed based on an optimal prompt strategy to facilitate the practical use of LLMs for programming.
Regulatory T cells (T reg ) are conventionally viewed as suppressors of endogenous and therapy-induced antitumor immunity; however, their role in modulating responses to immune checkpoint blockade (ICB) is unclear. In this study, we integrated single-cell RNA-seq/T cell receptor sequencing (TCRseq) of >73,000 tumor-infiltrating T reg (TIL-T reg ) from anti–PD-1–treated and treatment-naive non–small cell lung cancers (NSCLC) with single-cell analysis of tumor-associated antigen (TAA)–specific T reg derived from a murine tumor model. We identified 10 subsets of human TIL-T reg , most of which have high concordance with murine TIL-T reg subsets. Only one subset selectively expresses high levels of TNFRSF4 (OX40) and TNFRSF18 (GITR), whose engangement by cognate ligand mediated proliferative programs and NF-κB activation, as well as multiple genes involved in T reg suppression, including LAG3 . Functionally, the OX40 hi GITR hi subset is the most highly suppressive ex vivo, and its higher representation among total TIL-T reg correlated with resistance to PD-1 blockade. Unexpectedly, in the murine tumor model, we found that virtually all TIL-T reg –expressing T cell receptors that are specific for TAA fully develop a distinct T H 1-like signature over a 2-week period after entry into the tumor, down-regulating FoxP3 and up-regulating expression of TBX21 ( Tbet) , IFNG , and certain proinflammatory granzymes. Transfer learning of a gene score from the murine TAA-specific T H 1-like T reg subset to the human single-cell dataset revealed a highly analogous subcluster that was enriched in anti–PD-1–responding tumors. These findings demonstrate that TIL-T reg partition into multiple distinct transcriptionally defined subsets with potentially opposing effects on ICB-induced antitumor immunity and suggest that TAA-specific TIL-T reg may positively contribute to antitumor responses.
When analyzing data from in situ RNA detection technologies, cell segmentation is an essential step in identifying cell boundaries, assigning RNA reads to cells, and studying the gene expression and morphological features of cells. We developed a deep-learning-based method, GeneSegNet, that integrates both gene expression and imaging information to perform cell segmentation. GeneSegNet also employs a recursive training strategy to deal with noisy training labels. We show that GeneSegNet significantly improves cell segmentation performances over existing methods that either ignore gene expression information or underutilize imaging information.
The diversity of genetic programs and cellular plasticity of glioma-associated myeloid cells, and thus their contribution to tumor growth and immune evasion, is poorly understood. We performed single cell RNA-sequencing of immune and tumor cells from 33 glioma patients of varying tumor grades. We identified two populations characteristic of myeloid derived suppressor cells (MDSC), unique to glioblastoma (GBM) and absent in grades II and III tumors: i) an early progenitor population (E-MDSC) characterized by strong upregulation of multiple catabolic, anabolic, oxidative stress, and hypoxia pathways typically observed within tumor cells themselves, and ii) a monocytic MDSC (M-MDSC) population. The E-MDSCs geographically co-localize with a subset of highly metabolic glioma stem-like tumor cells with a mesenchymal program in the pseudopalisading region, a pathognomonic feature of GBMs associated with poor prognosis. Ligand-receptor interaction analysis revealed symbiotic cross-talk between the stemlike tumor cells and E-MDSCs in GBM, whereby glioma stem cells produce chemokines attracting E-MDSCs, which in turn produce growth and survival factors for the tumor cells. Our large-scale single-cell analysis elucidated unique MDSC populations as key facilitators of GBM progression and mediators of tumor immunosuppression, suggesting that targeting these specific myeloid compartments, including their metabolic programs, may be a promising therapeutic intervention in this deadly cancer. One-Sentence Summary:Aggressive glioblastoma harbors two unique myeloid populations capable of promoting stem-like properties of tumor cells and suppressing T cell function in the tumor microenvironment.
Cell type annotation is an essential step in single-cell RNA-seq analysis. However, it is a time-consuming process that often requires expertise in collecting canonical marker genes and manually annotating cell types. Automated cell type annotation methods typically require the acquisition of high-quality reference datasets and the development of additional pipelines. We demonstrate that GPT-4, a highly potent large language model, can automatically and accurately annotate cell types by utilizing marker gene information generated from standard single-cell RNA-seq analysis pipelines. Evaluated across hundreds of tissue types and cell types, GPT-4 generates cell type annotations exhibiting strong concordance with manual annotations, and has the potential to considerably reduce the effort and expertise needed in cell type annotation.
Background: Maternal prenatal smoking is known to alter offspring DNA methylation (DNAm). However, there are no effective interventions to mitigate smoking-induced DNAm alteration. Objectives: This study investigated whether 1-carbon nutrients (folate, vitamins B6, and B12) can protect against prenatal smoking-induced offspring DNAm alterations in the aryl hydrocarbon receptor repressor (AHRR) (cg05575921), GFI1 (cg09935388), and CYP1A1 (cg05549655) genes. Methods: This study included mother-newborn dyads from a racially diverse US birth cohort. The cord blood DNAm at the above 3 sites were derived from a previous study using the Illumina Infinium MethylationEPIC BeadChip. Maternal smoking was assessed by self-report and plasma biomarkers (hydroxycotinine and cotinine). Maternal plasma folate, and vitamins B6 and B12 concentrations were obtained shortly after delivery. Linear regressions, Bayesian kernel machine regression, and quantile g-computation were applied to test the study hypothesis by adjusting for covariables and multiple testing. Results: The study included 834 mother-newborn dyads (16.7% of newborns exposed to maternal smoking). DNAmat cg05575921 (AHRR) and at cg09935388 (GFI1) was inversely associated with maternal smoking biomarkers in a dose-response fashion (all P < 7.01 X 10(-13)). In contrast, cg05549655 (CYP1A1) was positively associated with maternal smoking biomarkers (P < 2.4 X 10(-6)). Folate concentrations only affected DNAmlevels at cg05575921 (AHRR, P = 0.014). Regression analyses showed that compared with offspring with low hydroxycotinine exposure (< 0.494) and adequate maternal folate concentrations (quartiles 2-4), an offspring with high hydroxycotinine exposure (< 0.494) and lowfolate concentrations (quartile 1) had a significant reduction inDNAmat cg05575921 (M-value, ss +/- SE - -0.801 +/- 0.117, P - 1.44 X 10(-11)), whereas adequate folate concentrations could cut smoking-induced hypomethylation by almost half. Exposure mixture models further supported the protective role of adequate folate concentrations against smoking-induced aryl hydrocarbon receptor repressor (AHRR) hypomethylation. Conclusions: This study found that adequate maternal folate can attenuate maternal smoking-induced offspring AHRR cg05575921 hypomethylation, which has been previously linked to a range of pediatric and adult diseases.
Pseudotime analysis with single-cell RNA-sequencing (scRNA-seq) data has been widely used to study dynamic gene regulatory programs along continuous biological processes. While many computational methods have been developed to infer the pseudo-temporal trajectories of cells within a biological sample, methods that compare pseudo-temporal patterns with multiple samples (or replicates) across different experimental conditions are lacking. Lamian is a comprehensive and statistically-rigorous computational framework for differential multi-sample pseudotime analysis. It can be used to identify changes in a biological process associated with sample covariates, such as different biological conditions, and also to detect changes in gene expression, cell density, and topology of a pseudotemporal trajectory. Unlike existing methods that ignore sample variability, Lamian draws statistical inference after accounting for cross-sample variability and hence substantially reduces sample-specific false discoveries that are not generalizable to new samples. Using both simulations and real scRNA-seq data, including an analysis of differential immune response programs between COVID-19 patients with different disease severity levels, we demonstrate the advantages of Lamian in decoding cellular gene expression programs in continuous biological processes.