Single-cell RNA-seq (scRNA-seq) enables atlas-scale profiling of complex tissues, revealing rare lineages and transient states. Yet, assigning biologically valid cell identities remains a bottleneck because markers are tissue- and state-dependent, and novel states lack references. We present CellMaster, an AI agent that mimics expert practice for zero-shot cell-type annotation. Unlike existing automated tools, CellMaster leverages LLM-encoded knowledge (e.g., GPT-4o) to perform on-the-fly annotation with interpretable rationales, without pre-training or fixed marker databases. Across 9 datasets spanning 8 tissues, CellMaster improved accuracy by 7.1
Neddylation is a post-translational modification suppressed by the endogenous inhibitor NUB1, and its dysregulation in rheumatoid arthritis (RA) promotes inflammation and NF-κB activation. NUB1 expression in RA and osteoarthritis (OA) synovium and mechanisms underlying defective NUB1 induction in RA fibroblast-like synoviocytes (FLS) were evaluated. NEDD8 and IL-6 protein expression was increased and NUB1 was reduced in RA synovium compared with OA tissue, with significant differences in the lining region that correlated with higher cytokine expression and NF-kB translocation. IL-1β–induced NUB1 induction was impaired in RA FLS at both the mRNA and protein levels. To evaluate the mechanism, we assessed mRNA stability using actinomycin D, examined the role of SNHG12 by siRNA knockdown, analyzed MAP kinase signaling, and measured NUB1 promoter activity with a luciferase reporter assay. None could explain the reduced induction observed in RA FLS. Treatment with the DNA methylation inhibitor 5-azacytidine and the histone methylation inhibitor EPZ6438 partially reversed the difference in NUB1 induction, whereas the histone deacetylase inhibitors ITF2375 and MS275 eliminated it. Therefore, defective NUB1 induction in RA FLS is related to the aberrant epigenetic landscape in RA FLS and is associated with increased neddylation and increased IL-6 expression in rheumatoid synovium. Overcoming increased neddylation in RA represents a novel therapeutic approach.
AMPA receptors (AMPARs) mediate the majority of fast excitatory synaptic transmission throughout the central nervous system. Calcium-permeable AMPARs and GluA4-containing receptors are critical for cerebellar functions, such as motor learning, associative memory, auditory processing, and synaptic plasticity. In contrast to the well-characterized, predominantly GluA2-containing AMPARs of the hippocampus and cortex, cerebellar AMPARs contain a higher proportion of GluA4 and remain poorly understood. Here, we generated a highly GluA4-specific antibody. Using this antibody in combination with antibodies specifically recognizing GluA1 and GluA2, we purified native AMPARs and determined the subunit compositions of both calcium-impermeable and calcium-permeable native AMPARs in the cerebellum. The isolated cerebellar AMPARs that contained both GluA1 and GluA4 were calcium-permeable, with GluA4 occupying mainly the B/D positions, GluA1 occupying the A/C positions, and the complex associated primarily with cornichon 3 (CNIH3). We determined the structures of the complex in distinct functional states, including the resting, active, and desensitized states, and characterized the conformational transitions that underlie its activity. During desensitization, the receptor adopts a pseudo-4-fold configuration of the ligand-binding domain layer, which may be important for its functional properties. This study provides a blueprint for the subunit compositions of AMPARs in the cerebellum and clarifies the gating mechanism of the calcium-permeable native AMPARA1A4-CNIH3 complex, providing significant insight into AMPAR-mediated synaptic transmission in the cerebellum.
Supplementary Figure 1 Identified major cell types, cell proportion analysis and CD112 and CD155 expression in NB tumor scRNA-seq analysis. (A) Violin plots showing the number of features, RNA counts and percent mitochondrial transcripts following quality control. (B) Violin plots showing the average expression of the genes for annotating the major cell types identified in NB tumors. (C) UMAP plots showing the distribution of cell types for each sample. (D) Bar plots showing the proportions of cell types for each sample. (E) UMAP plots showing the expression levels of CD112 and CD155.
Supplementary Figure 3 The expression analysis of CD112 and CD155 in NB. (A) and (B) The expression of CD112 in NB cell lines and primary NB tumor tissues are identified using qPCR and Western blot. (C) and (D) The expression of CD112 in NB cell lines and primary NB tumor tissues are identified using FACS. (E) Soluble CD112 in serum is detected by ELISA (HC, n = 17, NB, n = 22) and is further analyzed between different groups divided according to INSS stage, risk and MYCN status. (F) and (G) The expression of CD155 in NB cell lines and primary NB tumor tissues are identified using qPCR and Western blot. (H) and (I) The expression of CD155 in NB cell lines and primary NB tumor tissues are identified using FACS. (J) Soluble CD155 in serum is detected by ELISA (HC, n = 21, NB, n = 35) and is further analyzed between different groups divided according to INSS stage, risk and MYCN status. The results are expressed as the means ± SEMs from at least three independent experiments. Significant differences between groups are represented by ns no significance, *** p < 0.001.
Supplementary Figure 8 Enriched pathways and functional characteristics of CD8 T and NK cells. (A) The enriched KEGG pathways in CD8 T cells from high-ratio and low-ratio groups are shown in dot plots. The dot size in KEGG enrichment analysis represents enriched gene numbers. (B) and (C) Cell-cycle score and GSVA score analysis of CD8 T cells. (D) Violin plots showing the expression levels of DNAM-1, TIGIT, CD96, PD-1, TIM-3, IFN-γ, granzyme B and perforin in CD8 T cells. (E) The enriched KEGG pathways in NK cells from high-ratio and low-ratio groups are shown in dot plots. The dot size in KEGG enrichment analysis represents enriched gene numbers. (F) and (G) Cell-cycle score and GSVA score analysis of NK cells. (D) and (H) Violin plots showing the expression levels of DNAM-1, TIGIT, CD96, PD-1, TIM-3, IFN-γ, granzyme B and perforin in NK cells.
Supplementary Figure 11 Low-dose doxorubicin affects the expression of CD112 and CD155 in NB cells. (A) The apoptosis of SH-SY5Y and SK-N-BE2 cells treated with low-dose doxorubicin (0.1 μM) is detected by FACS (left) and AV+ proportion is labeled in the representative FACS histograms. (B-F) CD112 and CD155 expression in SH-SY5Y and SK-N-BE2 cells or primary NB tumor cells treated with low-dose doxorubicin (0.1 μM) are detected by FACS and Western blot. MFI is labeled in the representative FACS histograms. (G) and (H) The cytotoxicity assays are performed to detected γδT-cell cytotoxicity against SH-SY5Y and SK-N-BE2 cells pretreated with 0.1 μM doxorubicin. The results are expressed as the means ± SEMs from at least three independent experiments. Significant differences between groups are represented by * p < 0.05 and *** p < 0.001.
Supplementary Figure 10 Identification of Tet-on system, ROC curves and survival curves of different groups for NB tumor. (A) and (B) The variation of CD112 and CD155 expression in SK-N-BE2 cells treated with different concentration of doxycycline (0, 2 nM, 20 nM, 200 nM, 2 μM, 20 μM) is detected by qPCR and Western blot. (C) ROC curves are generated for CD112, CD155, and CD112/CD155 ratio to predict NB patient death, tumor risk or MYCN amplification. The results are expressed as the means ± SEMs from at least three independent experiments. Significant differences between groups are represented by ns no significance, * p < 0.05, ** p < 0.01 and *** p < 0.001.
Supplementary Figure 2 Gating strategies of FACS analysis and the expression pattern of TIGIT in NB. (A) Gating strategies for the FACS analysis of circulating γδT cells. (B) Gating strategies for the FACS analysis of tumor infiltrating immune cells. (C) and (D) The expression analysis of CD96 and CD112R in peripheral blood (PB) γδT cells from healthy controls (HC, n = 10) and neuroblastoma patients (NB, n = 18). The proportions of γδT cells are shown in the left panel statistically and the representative FACS histograms are shown in the right panel with mean fluorescence intensity (MFI) labeled in corresponding histograms. (E) Using GSE49711 from GEO database and R2: Genomics Analysis and Visualization Platform, the expression pattern of TIGIT in NB tumors (n = 498) with different INSS stages, risks, MYCN status and death events is analyzed and illustrated by box plots. (F) The overall and eventfree survival curves are generated by grouping samples with the median of TIGIT expression. (G) The expression correlation between TIGIT and DNAM-1 are analyzed using GSE49711 dataset. The results are expressed as the means ± SEMs. Significant differences between groups are represented by ns no significance, ** p < 0.01 and **** p < 0.0001.
IntroductionInfants exposed to opioids in utero are at risk of developing Neonatal Opioid Withdrawal Syndrome (NOWS). Rodent models of perinatal opioid exposure can reliably recapitulate the acute withdrawal and developmental deficits exhibited in clinical NOWS, but few persisting phenotypes are consistently observed between studies, limiting mechanistic insight into the long-lasting effects of early-life opioid exposure.MethodsTo investigate the enduring impact of perinatal opioid exposure, we employed a multi-region, multi-omic approach. We combined RNA sequencing (RNA-seq) and H3K27ac chromatin immunoprecipitation sequencing (ChIP-seq) from NeuN+ neuronal nuclei in a mouse model of NOWS. Additionally, cytokine levels were measured in the brain and spleen, and physiological responses were assessed under both basal and immune-challenged conditions.ResultsAnalysis revealed differentially expressed genes and H3K27ac modifications were enriched for immune and metabolic pathways in hypothalamic neurons. Transcription factor network inference identified state-dependent rewiring of immune and metabolic regulatory circuits, with Transcription factor 4 (TCF4) emerging as a convergent epigenomic and transcriptional hub under immune-challenged conditions. Consistent with molecular signatures, cytokine levels were suppressed in morphine-exposed mice both in adulthood, both at baseline and under immune-challenged conditions. Several metabolic properties, including changes in weight and basal body temperature, were also altered.DiscussionOur findings suggest that perinatal opioid exposure creates a lasting enhancer imprint in hypothalamic neurons, leading to suppressed baseline immune gene expression and altered regulatory responses to subsequent inflammatory challenges. These epigenomic and transcriptional changes may underlie the long-term physiological impacts of early-life opioid exposure, offering new insights into the enduring consequences of NOWS.
Transcription factors (TFs) govern cell fate through coordinated gene-regulatory networks, yet the full potential of these networks to generate non-native, therapeutically advantageous cell states in vivo remains largely unexplored. We hypothesized that systematic gain-of-function (GOF) overexpression of TFs in CD8+ T cells, central mediators of immune protection, could reveal latent, or "hidden," regulatory programs capable of generating synthetic T cell states with therapeutic utility. To test this, we developed single-cell GOF sequencing (scGOF-seq), a multiplexed platform for unbiased, in vivo mapping of GOF effects on T cell fate in immunocompetent mouse models of infection and cancer. scGOF-seq uncovered unexpected regulators of T cell differentiation and accumulation, including SOX2, OCT4, and GATA2, which are normally silenced during T cell differentiation. Notably, outside its native regulatory context, supraphysiologic cMyc GOF reprogrammed CD8+ T cells into a synthetic stem-effector hybrid state, enabling >5,000-fold antigen-dependent expansion and antitumor activity, contrasting sharply with its native function in driving terminal differentiation. scGOF-seq further identified TF modules that cooperate with cMyc GOF to promote robust CD8+ T cell responses in solid tumors. Together, these findings establish GOF perturbation as a powerful strategy for revealing latent immune regulatory programs and engineering synthetic immune states with therapeutic potential.
CD8+ T cells differentiate into diverse states that shape immune outcomes in cancer and chronic infection1-4. To define systematically the transcription factors (TFs) driving these states, we built a comprehensive atlas integrating transcriptional and epigenetic data across nine CD8+ T cell states and inferred TF activity profiles. Our analysis catalogued TF activity fingerprints, uncovering regulatory mechanisms governing selective cell state differentiation. Leveraging this platform, we focused on two transcriptionally similar but functionally opposing states that are critical in tumour and viral contexts: terminally exhausted T (TEXterm) cells, which are dysfunctional5-8, and tissue-resident memory T (TRM) cells, which are protective9-13. Global TF community analysis revealed distinct biological pathways and TF-driven networks underlying protective versus dysfunctional states. Through in vivo CRISPR screening integrated with single-cell RNA sequencing (in vivo Perturb-seq) we delineated several TFs that selectively govern TEXterm cell differentiation. We also identified HIC1 and GFI1 as shared regulators of TEXterm and TRM cell differentiation and KLF6 as a unique regulator of TRM cells. We discovered new TEXterm-selective TFs, including ZSCAN20 and JDP2, with no previous known function in T cells. Targeted deletion of these TFs enhanced tumour control and synergized with immune checkpoint blockade but did not interfere with TRM cell formation. Consistently, their depletion in human T cells reduces the expression of inhibitory receptors and improves effector function. By decoupling exhaustion TEX-selective from protective TRM cell programmes, our platform enables more precise engineering of T cell states, accelerating the rational design of more effective cellular immunotherapies.
Serine protease inhibitor clade E member 1 (SERPINE1) is involved in various biological processes, but its role in promoting or suppressing tumorigenesis remains controversial. The present study focused on the effects of SERPINE1 downregulation on cell proliferation and invasion across three types of tumors to elucidate the underlying mechanisms. Based on data from an analysis of The Cancer Genome Atlas dataset, high SERPINE1 levels in patients with breast cancer and low‑grade glioma were associated with a poor prognosis, whereas elevated SERPINE1 expression in patients with skin cutaneous melanoma associated with improved outcomes. With respect to cell proliferation phenotypes, SERPINE1 knockdown increased xenograft growth and the proliferation of melanoma C918 cells by promoting cell cycle progression through the modulation of minichromosome maintenance complex component 3 and the activity of p53/SMAD3 regulators; conversely, SERPINE1 knockdown reduced the xenograft growth and proliferation of MDA‑MB‑231 breast cancer cells by decreasing the urokinase‑type plasminogen activator receptor‑mediated ERK/p38 activity ratio and similarly decreased proliferation in H4 glioma cells through an heat shock protein 90‑alpha (HSP90α)‑mediated reduction in the ERK/p38 activity ratio. Regarding invasion and metastasis, SERPINE1 knockdown consistently reduced invasion, matrix metalloproteinase (MMP) activity, and lung metastasis in both C918 and MDA‑MB‑231 cells but paradoxically increased invasion and MMP‑1 activity in H4 cells through the HSP90α‑p38‑MMP‑1 axis. Collectively, these findings suggested that SERPINE1 exerts diverse effects on cell proliferation and invasion through multiple regulatory mechanisms. These findings indicated that therapy targeting SERPINE1, which involves a comprehensive understanding of its diverse mechanisms of function, can increase treatment precision and reduce adverse reactions.
Accurate circulating tumor DNA (ctDNA) detection is limited by high sequencing errors. To address this, we developed Methyltransferase-Assisted Single Duplex sequencing (MASD-seq), which physically links Watson and Crick strands to enable error correction using double-stranded information within a single read pair. MASD-seq employs enzymatic methyl-sequencing to disrupt strand pairing and uses CpG methyltransferase pre-methylation to enable detection of CpG>TpG and transversion mutations in DNA duplexes. In parallel, MspI digestion enriches tumor CpG>TpG mutations and reduces genome complexity. MASD-seq achieves error rates of 5.4 × 10−7 for CpG>TpG and 4.1 × 10−8 for transversions in cell lines, allowing detection of age-associated mutational burdens in blood. Tissue analyses show a 4-fold enrichment of CpG>TpG mutations and a 1.6-fold increase in total mutations in high microsatellite instability (MSI-H) colorectal cancer (CRC), and 3-fold and 1.2-fold enrichments in microsatellite stable (MSS) tumors, respectively, yielding a correctable mutation burden comparable to the whole genome. Clinical testing demonstrates an area under the curve (AUC) of 0.97 for ctDNA detection in CRC patients. MASD-seq is a highly accurate sequencing method, suggesting its potential for future clinical applications. This study presents MASD-seq, a duplex sequencing method for highly accurate circulating tumor DNA (ctDNA) detection. The approach is combined with cost-efficient tumor mutation profiling in tissue samples for minimal residual disease monitoring. This study presents MASD-seq, a duplex sequencing method for highly accurate circulating tumor DNA (ctDNA) detection. The approach is combined with cost-efficient tumor mutation profiling in tissue samples for minimal residual disease monitoring.
Supplementary Figure 7 Metabolic states of high-ratio and low-ratio tumor cells. (A) Enriched metabolic reactions in high-ratio and low-ratio tumor cells revealed by COMPASS analysis. (B) Changes in metabolic KEGG pathways revealed by spatial metabolome analysis. (C) The heatmaps of representing metabolites.
Single-cell RNA sequencing (scRNA-seq) has revolutionized the study of cellular heterogeneity by providing gene expression data at single-cell resolution, uncovering insights into rare cell populations, cell-cell interactions, and gene regulation. Foundation models pretrained on large-scale scRNA-seq datasets have shown great promise in analyzing such data, but existing approaches are often limited to modeling a small subset of highly expressed genes and lack the integration of external gene-specific knowledge. To address these limitations, we present scLong, a billion-parameter foundation model pretrained on 48 million cells. scLong performs self-attention across the entire set of 28,000 genes in the human genome. This enables the model to capture long-range dependencies between all genes, including lowly expressed ones (containing unexpressed genes with zero expressions), which often play critical roles in cellular processes but are typically excluded by existing foundation models. Additionally, scLong integrates gene knowledge from the Gene Ontology using a graph convolutional network, enriching its contextual understanding of gene functions and relationships. In extensive evaluations, scLong surpasses both state-of-the-art scRNA-seq foundation models and task-specific models across diverse tasks, including predicting transcriptional responses to genetic and chemical perturbations, forecasting cancer drug responses, and inferring gene regulatory networks.
Supplementary Figure 13 The effects of CD155 on γδT-cell-cytotoxicity against neuroblastoma, and on proliferation and migration of NB cells. (A) and (B) α-TIGIT is used in the rhCD155 treatment of γδT cells. DNAM-1 is detected by FACS and Western blot. (C-F) rhCD155 is used to treat γδT cells for 24 h. TIGIT, CD96 and CD112R are detected by FACS (left) or Western blot and MFI is labeled in the representative FACS histograms (right). (G) γδT-cell numbers are counted and statistically analyzed during in vitro culture in presence of rhCD155. (H) The CFSE proliferation assays are performed during in vitro culture in presence of rhCD155. γδT cells are detected by FACS (right) and statistically analyzed (left). MFI is labeled in the representative FACS histograms. (I) γδT cells are pre-treated with rhCD155 for 24 h and subjected to calcium flux detection by Fluo-4 AM labeling and low-concentration PMA and ionomycin stimulation. The calcium flux results are used to benchmark the levels of activation. The results are recorded and illustrated by FACS (right) and statistically analyzed (left). (J) The phosphorylation of Akt and ERK1/2 in γδT cells is detected by Western blot after treatment using rhCD155 for 24 h. (K) and (L) The cytotoxicity assays are performed with γδT cells pre-treated with rhCD155 or α-TIGIT for 24 h and NB cell lines. (M) and (N) The cytotoxicity assays are performed to detected γδT-cell cytotoxicity against SH-SY5Y and CHLA-255 cells in presence of α-CD155. (O-Q) IFN-γ, perforin and granzyme B expression are detected by FACS upon PMA and ionomycin stimulation after rhCD155 pre-treatment (left) and MFI is labeled in the representative FACS histograms (right). (R-T) CCK-8, Transwell and wound-healing sassays are performed using SK-N-BE2-NC and SK-N-BE2-KoCD155 cells. Representative images of SK-N-BE2 cell migration obtained from the Transwell (magnification × 100) and wound-healing (magnification × 100) assays are shown (right). The cell numbers obtained from the Transwell assays are counted, and the relative migration rate obtained from the wound-healing assays is calculated by dividing the change in the distance between the scratch edges by the initial distance (left). (U) The phosphorylation of Akt and ERK1/2 in γδT cells is detected by Western blot after SK-N-BE2-NC and SK-N-BE2-KoCD155 stimulation. (V) TIGIT is detected by FACS after 24 h co-culture with SK-N-BE2-NC and SK-N-BE2-KoCD155 cells (left) and MFI is labeled in the representative FACS histograms (right). The results are expressed as the means ± SEMs from at least three independent experiments. Significant differences between groups are represented by ns no significance, * p < 0.05, ** p < 0.01 and *** p < 0.001.
Protein language models (pLMs) have become essential tools in computational biology, powering diverse applications from variant effect prediction to protein engineering. Central to their success is the use of pretrained embeddings-contextualized representations of amino acid sequences-which enable effective transfer learning, especially in data-scarce settings. However, recent studies have revealed that standard masked language modeling objectives used to train these models often produce representations that are misaligned with the needs of downstream tasks. While scaling up model size improves performance in some cases, it does not universally yield better representations. In this study, we investigate two complementary strategies for improving pLM representations: (i) integrating text annotations through contrastive learning, and (ii) combining multiple embeddings via embedding fusion. We benchmark six text-integrated pLMs (tpLMs) and three large-scale pLMs across six biologically diverse tasks, showing that no single model dominates across settings. Fusion of multiple tpLMs embeddings improves performance on most tasks but presents a computational bottleneck due to the combinatorial number of possible combinations. To overcome this, we propose greedier forward selection, a linear-time algorithm that efficiently identifies near-optimal embedding subsets. We validate its utility through two case studies, homologous sequence recovery and protein-protein interaction prediction, demonstrating new state-of-the-art results in both. Our work highlights embedding fusion as a practical and scalable strategy for improving protein representations.