RNA modifications represent a dynamic layer of gene expression regulation, RNA stability, and translation with profound implications for cellular function and disease. However, the critical regulation and functions of RNA-modifying proteins (RMPs) remain poorly understood. Here, we present a large-scale characterization of RMPs through 378 multiomics datasets encompassing genomics, bulk and single-cell transcriptomics, epitranscriptomics, proteomics, and posttranslational modifications (PTMs) across 63 human tissues. Our analysis of experimental perturbations of RMPs revealed dynamic differential modification peaks and expressed genes. We applied nonnegative matrix factorization to annotate RMP-mediated cell types in single-cell transcriptomes. Functional annotations in acute myeloid leukemia (AML) revealed RMPs such as ALKBH5 as critical mediators of m6A dynamics, influencing pathways involved in translation initiation, immune regulation, and tumorigenesis. We revealed cell type-specific modification patterns, including those in ALKBH5-enriched AML stem cells with special ligand‒receptor interactions and genetic variations modulated by m6A. We integrated proteogenomic data to uncover PTM-associated regulatory, mutation, and protein‒protein interaction networks linked to RMPs. We developed RMzyme, a platform that consolidates our findings and provides insights into RMPs and their downstream effects. This resource is expected to facilitate biomedical research into the molecular mechanisms of human diseases through the lens of RNA modifications and multiomics data integration.
Extrachromosomal circular DNAs (eccDNAs) are closed circular DNA molecules widespread across eukaryotic cells, with emerging roles in gene regulation and tumor progression. Experimental assays remain costly and incomplete, underscoring the need for computational approaches. To address this, a deep learning framework termed DeepECC has been established to overcome the challenges posed by eccDNA heterogeneity and its complex biogenesis. Through a two-stage training strategy, DeepECC models the local sequence context flanking both the start and end breakpoints, thereby capturing mechanistically informative features that are often overlooked when analyses focus solely on eccDNA body sequences. Applied to multi-species (human, mouse, gallus) datasets, DeepECC robustly captures conserved breakpoint features, with a marked preference for GC-rich and transcriptionally active regions. Genome-wide scanning reveals non-uniform distributions of human cancer eccDNAs enriched in enhancers, expression quantitative trait loci, and CTCF sites, suggesting regulatory functions in tumor progression. Motif analysis further implicates ribosomal activity, translational regulation, and DNA damage response. Furthermore, genome-wide eccDNA predictions are integrated into the UCSC Genome Browser, enabling convenient querying and visualization of cancer-related eccDNAs associated with specific genes or genomic regions, facilitating functional interpretation for experimental research. Collectively, DeepECC provides a generalizable framework for systematic eccDNA discovery and insights into their functional significance in cancer.
RNA-binding proteins (RBPs) regulate diverse post-transcriptional processes, yet their functional impacts across multiple molecular layers remain incompletely characterized. KnockRBP (https://knockrbp.xu-bioinfo.com) is presented as a comprehensive, multi-omics database systematically cataloging regulatory changes resulting from genetic perturbation of RBPs. The resource integrates 1453 datasets from GEO, ENA, and ENCORE, encompassing RNA-seq, Ribo-seq, microRNA (miRNA)-seq, and 3'RACE-seq across 639 RBPs, 182 cell lines, and 61 diseases. Regulatory outcomes covered include alternative splicing, alternative polyadenylation, RNA editing, transcript abundance, translation efficiency, and miRNA expression, with all events processed through standardized analytical pipelines and annotated with open reading frame consequences, predicted loss-of-function impacts, and established disease or drug associations. The web interface supports dataset browsing, RBP- and gene-centric multi-omics exploration, transcript structure visualization, and interactive regulatory network navigation, while integrated analytical modules enable prioritization of perturbation-responsive events, cross-layer regulatory inference, and evaluation of clinical relevance. In contrast to existing RBP resources, KnockRBP integrates multi-omics data with perturbation-aware functional annotation, enabling mechanistic investigation and identification of RBP-affected genes and events in post-transcriptional regulation. The database is freely accessible and intended for long-term maintenance, serving as a community resource to advance functional genomics research on RBPs and their roles in human disease.
Introduction Advances in single-cell and spatial omics have enabled deep insights into cellular heterogeneity, characterized by molecular features that vary significantly across cell types, functional states, or diseases. However, current differential analysis methods often fail to identify molecular features that are both highly specific and applicable for experimental applications such as cell sorting and clinical diagnostics. Objectives This study aims to develop a novel methodology for identifying cell-type-specific molecular features from single-cell and spatial omics data. Methods We present CSFeatures, a computational method designed to identify cell-type-specific differential features. We performed a comprehensive evaluation on a simulated dataset and 14 real datasets, spanning multiple omics data, including single-cell RNA-seq (scRNA-seq), single-cell assay for transposase-accessible chromatin sequencing (scATAC-seq), spatial transcriptomics (ST), and spatial ATAC-seq (spaATAC-seq). Furthermore, we performed protein-level validation using multiplex immunofluorescence. Results CSFeatures integrates expression levels, distribution smoothness, and proportional representation across cell populations to identify cell-type-specific differential features from single-cell and spatial omics data. Benchmarking with 10 existing methods on 10 scRNA-seq datasets demonstrated that CSFeatures reliably identifies cell-type-specific genes highly expressed in target populations and minimally expressed elsewhere. These genes are enriched in pathways and functional categories relevant to the target cells’ functional states. Crucially, the high specificity of these markers was further validated at the protein level through multiplex immunofluorescence analysis, indicating their applicability for experimental applications such as cell sorting and clinical diagnostics. Moreover, we demonstrate that CSFeatures can robustly identify cell-type-specific features from diverse omics data, including scATAC-seq, ST, and spaATAC-seq, which offers insights into gene regulation and spatial heterogeneity. Conclusions Overall, CSFeatures enhances the identification of cell-type-specific features from multi-modal single-cell and spatial omics data. This yields high-fidelity molecular markers poised for the accurate characterization of cell populations, thereby enabling deeper mechanistic investigations within experimental and clinical settings.
Abstract Advances in single-cell and spatial technologies have transformed the dissection of cell composition and tissue architecture in complex biological systems. However, identifying rare cells critical to disease pathogenesis and biological processes remains difficult, as their low abundance often masks them among dominant cell populations. We introduce RareQ, a fast and scalable framework for rare cell detection by evaluating cliquishness of each cell’s k-nearest neighborhood using single-cell omics data. Extensive benchmarking on diverse simulated and real datasets demonstrates that RareQ exceeds existing methods in accuracy, sensitivity, and efficiency. RareQ also excels at identifying both modality-specific and shared rare cells. Its versatility across various biological contexts enables the discovery of functionally distinct rare cells with unique molecular signatures in physiological and pathological context. RareQ’s application to spatial transcriptomics data reveals anatomically distinct and clinically relevant rare cell populations. Together, RareQ offers an efficient approach for rare cell discovery, enhancing insights into tissue organization and disease mechanisms.
Abstract Inferring early cell fate from single-cell RNA-sequencing data is essential for identifying cellular origins and fate plasticity in development and disease. However, existing methods often fail to exploit tree-structured lineage trajectories, limiting the accuracy and interpretability of fate mapping. Here we present DyMoTree, a computational framework that models cell fate decisions as nonlinear mappings between progenitor and terminal cell states under explicit lineage constraints. By integrating lineage graphs with a tree-structured neural architecture, DyMoTree learns lineage-resolved cell-state transition maps from single-cell transcriptomes, enabling robust inference of early fate bias and identification of fate-specific progenitor substates and driver genes. Across simulations, lineage-tracing experiments, and in vivo systems, DyMoTree outperformed existing methods in resolving early fate biases. Applications to mouse embryogenesis, lung adenocarcinoma progression, and CAR-T immunotherapy revealed regulatory programs underlying developmental and disease-associated transitions. DyMoTree provides a general framework for modeling lineage-resolved cell-state dynamics underlying development and disease progression. Abstract Figure
Functional perturbations of genes do not always cause expression changes, but can manifest through network rewiring or context-specific shifts in regulatory activity. However, inferring functional shifts in genes and linking them to the specific cell populations remains challenging, as most current scRNA-seq data analysis focuses either on differential gene expression or on cell abundance/state changes, but rarely associate gene perturbations with particular cell populations. Here we present scDNS, a framework that quantifies gene-specific functional perturbations by measuring information-theoretic divergence between condition-specific gene interaction network configurations. In simulated stress tests and multiple experimental datasets, scDNS prioritizes key regulators and perturbed cell populations, even when expression changes are minimal but network rewiring is pronounced. Applications to immunodeficiency mutations, stimulus responses, and viral infection reveal hidden regulatory programs and heterogeneous responder cell states. In pancreatic cancer, scDNS nominates TIMM44 as a mitochondrial sensitizer enhancing gemcitabine efficacy. Together, scDNS provides a powerful tool for inferring dynamic gene perturbations in single cells.
Recent advancements in single-cell and spatial omics have enabled the analysis of gene expression patterns in cells and tissues with unprecedented precision. Genes that exhibit significant variation across different cell types or states are typically closely linked to cellular functional states and disease processes. However, current differential analysis methods often fail to adequately account for cell type specificity when identifying differentially expressed genes. CSFeatures effectively addresses this issue by comprehensively considering the expression level, the smoothness of expression distribution, and the proportion of gene expression in the target cell population as well as in each of the other cell populations. Through comprehensive experiments on eight single-cell RNA-seq datasets, it was demonstrated that CSFeatures can effectively identify cell type-specific genes that are highly expressed in the target cell population while being lowly expressed in all other populations. Moreover, it was validated that the underlying principle of CSFeatures can be generalized to other omics data, including single-cell ATAC-seq, spatial transcriptomics, and spatial ATAC-seq data. Furthermore, the differential features identified by CSFeatures are enriched in pathways, functional categories, and regulatory regions that are highly relevant to the specific functional states of the target cell population, providing insights into its gene regulation. ### Competing Interest Statement The authors have declared no competing interest. National Natural Science Foundation of China, 62171365, 62471378, 62473212 Young Talent Support Plan of Xi’an Jiaotong University, YX6J021 Young Elite Scientists Sponsorship Program by CAST, 2023QNRC001 Shaanxi Province Key Research and Development Projects, 2024SF-GJHX-40, QCYRCXM-2022-209
BackgroundHepatocellular carcinoma (HCC) is a major global health challenge with high aggressiveness and recurrence rates. Metabolic reprogramming is a cancer hallmark, enabling tumor cells to sustain rapid growth and evade immune surveillance. Several amino acids have been found to undergo metabolic reprogramming in tumors, and thus are potential anti-tumor targets. However, the characterization and implication of lysine metabolic reprogramming in HCC remain largely unexplored.MethodsWe performed multi-omics profiling, including transcriptomics, proteomics, single-cell omics, immunohistochemistry, and multiplex immunofluorescence on tumor and adjacent normal tissues obtained from 30 HCC patients. Integrative analyses and quantitative evaluations were carried out to characterize the lysine metabolism and investigate its implications for tumor progression, immune microenvironment, and immunotherapy responses.ResultsOur analysis observed a significant downregulation of lysine metabolism in HCC, with inter-patient heterogeneity. Patients with low lysine metabolism in tumors exhibited worse prognoses and a predominance of immunosuppressive tumor immune microenvironment (TIME), characterized by increased infiltration of myeloid-derived suppressor cells (MDSCs), regulatory T cells (Tregs), and exhausted CD8+ T cells (TIM3+CD8+ and LAG3+CD8+). These immunosuppressive cells contribute to immunotherapeutic resistance and promote tumor progression. Notably, our conclusions were consistently supported by observations at both the bulk and single-cell resolutions, as well as T cell receptor (TCR) immune repertoire profiling, reinforcing the robustness of our findings.ConclusionsThis study provides comprehensive evidence that lysine metabolism plays a critical role in shaping the immunosuppressive TIME in HCC and is associated with clinical outcomes and resistance to immunotherapy, offering new insights into clinical molecular subtyping and potential therapeutic strategies.
ObjectiveTo analyse the global burden, trends and cross-country inequalities of female breast and gynaecologic cancers (FeBGCs).DesignPopulation-Based Study.SettingData sourced from the Global Burden of Disease Study 2019.PopulationIndividuals diagnosed with FeBGCs.MethodsAge-standardised mortality rates (ASMRs), age-standardised Disability-Adjusted Life Years (DALYs) rates (ASDRs) and their 95% uncertainty interval (UI) described the burden. Estimated annual percentage changes (EAPCs) and their confidence interval (CI) of age-standardised rates (ASRs) illustrated trends. Social inequalities were quantified using the Slope Index of Inequality (SII) and Concentration Index.Main Outcome MeasuresThe main outcome measures were the burden of FeBGCs and the trends in its inequalities over time.ResultsIn 2019, the ASDRs per 100 000 females were as follows: breast cancer: 473.83 (95% UI: 437.30-510.51), cervical cancer: 210.64 (95% UI: 177.67-234.85), ovarian cancer: 124.68 (95% UI: 109.13-138.67) and uterine cancer: 210.64 (95% UI: 177.67-234.85). The trends per year from 1990 to 2019 were expressed as EAPCs of ASDRs and these: for Breast cancer: -0.51 (95% CI: -0.57 to -0.45); Cervical cancer: -0.95 (95% CI: -0.99 to -0.89); Ovarian cancer: -0.08 (95% CI: -0.12 to -0.04); and Uterine cancer: -0.84 (95% CI: -0.93 to -0.75). In the Social Inequalities Analysis (1990-2019) the SII changed from 689.26 to 607.08 for Breast, from -226.66 to -239.92 for cervical, from 222.45 to 228.83 for ovarian and from 74.61 to 103.58 for uterine cancer. The concentration index values ranged from 0.2 to 0.4.ConclusionsThe burden of FeBGCs worldwide showed a downward trend from 1990 to 2019. Countries or regions with higher Socio-demographic Index (SDI) bear a higher DALYs burden of breast, ovarian and uterine cancers, while those with lower SDI bear a heavier burden of cervical cancer. These inequalities increased over time.
BACKGROUND:Alternative splicing (AS) plays a crucial role in regulating gene expression and governing proteomic diversity by generating multiple protein isoforms from a single gene. Increasing evidence has highlighted the regulation for pre-mRNA splicing of the splicing factors (SFs). This review aims to examine featured mechanisms and examples of SF regulation by AS, focusing on paradigmatic feedback loops and their biological implications. MAIN BODY OF THE ABSTRACT:We specifically focus on the autoregulation and inter-regulation of SFs through AS machinery. These interactions give rise to a feedback system, where the negative feedback loops aid in maintaining cellular homeostasis, and the positive feedback loops play roles in triggering cellular state transitions. We examine the growing evidence highlighting the specific mechanisms employed by SFs to autoregulate their own splicing, including AS-coupled nonsense-mediated mRNA decay (AS-NMD), nuclear retention, and alternative 3'UTR regulation. We showcase the influence of AS feedback in amyotrophic lateral sclerosis (ALS), frontotemporal dementia (FTD), and cancer. Furthermore, we discuss how master splicing factors can dominantly orchestrate splicing cascades, leading to widespread impacts in cellular processes. We also discuss how non-coding RNAs, particularly circular RNAs and microRNAs, engage in the splicing regulatory networks. Lastly, we showcase how negative and positive feedback loops can collaboratively achieve remarkable biological functions during the cell fate decision. SHORT CONCLUSION:This review highlights the regulation of SFs by AS, providing enriched information for future investigations that aim at deciphering the intricate interplay within splicing regulatory networks. KEY POINTS:Negative feedback of alternative splicing maintains cellular homeostasis. Positive feedback of alternative splicing triggers cellular state transitions. Alternative splicing forms integrated feedback networks with circRNAs and microRNAs to reciprocally regulate their expression and function. The coordinated interplay of distinct splicing feedback mechanisms orchestrates precise cell fate transitions. Future directions and therapeutic possibilities that could transform alternative splicing research into treatments.
Alternative splicing (AS) plays a critical role in gene expression by generating protein diversity from single genes. This review provides an overview of the role of AS in regulating cell fate, focusing on its involvement in processes such as cell proliferation, differentiation, apoptosis, and tumorigenesis. We explore how AS influences the cell cycle, particularly its impact on key stages like G1, S, and G2/M. The review also examines AS in cell differentiation, highlighting its effects on mesenchymal stem cells and neurogenesis, and how it regulates differentiation into adipocytes, osteoblasts, and chondrocytes. Additionally, we discuss the role of AS in programmed cell death, including apoptosis and pyroptosis, and its contribution to cancer progression. Importantly, targeting aberrant splicing mechanisms presents promising therapeutic opportunities for restoring normal cellular function. By synthesizing recent findings, this review provides insights into how AS governs cellular fate and offers directions for future research into splicing regulatory networks.
Methamphetamine (METH) is a highly addictive psychostimulant that causes physical and psychological damage and immune system disorder, especially in the liver which contains a significant number of immune cells. Dopamine, a key neurotransmitter in METH addiction and immune regulation, plays a crucial role in this process. Here, we developed a chronic METH administration model and conducted single-cell RNA sequencing (scRNA-seq) to investigate the effect of METH on liver immune cells and the involvement of dopamine receptor D1 (DRD1). Our findings reveal that chronic exposure to METH induces immune cell identity shifts from IFITM3+ macrophage (Mac) and CCL5+ Mac to CD14+ Mac, as well as from FYN+CD4+ T effector (Teff), CD8+ T, and natural killer T (NKT) to FOS+CD4+ T and RORα+ group 2 innate lymphoid cell (ILC2), along with the suppression of multiple functional immune pathways. DRD1 is implicated in regulating certain pathways and identity shifts among the hepatic immune cells. Our results provide valuable insights into the development of targeted therapies to mitigate METH-induced immune impairment.
Male infertility has emerged as a global issue, partly attributed to psychological stress. However, the cellular and molecular mechanisms underlying the adverse effects of psychological stress on male reproductive function remain elusive. We created a psychologically stressed model using terrified-sound and profiled the testes from stressed and control rats using single-cell RNA sequencing. Comparative and comprehensive transcriptome analyses of 11,744 testicular cells depicted the cellular landscape of spermatogenesis and revealed significant molecular alterations of spermatogenesis suffering from psychological stress. At the cellular level, stressed rats exhibited delayed spermatogenesis at the spermatogonia and pachytene phases, resulting in reduced sperm production. Additionally, psychological stress rewired cellular interactions among germ cells, negatively impacting reproductive development. Molecularly, we observed the down-regulation of anti-oxidation-related genes and up-regulation of genes promoting reactive oxygen species (ROS) generation in the stress group. These alterations led to elevated ROS levels in testes, affecting the expression of key regulators such as ATF2 and STAR, which caused reproductive damage through apoptosis or inhibition of testosterone synthesis. Overall, our study aimed to uncover the cellular and molecular mechanisms by which psychological stress disrupts spermatogenesis, offering insights into the mechanisms of psychological stress-induced male infertility in other species and promises in potential therapeutic targets.
Background Different individuals with renal cell carcinoma (RCC) exhibit substantial heterogeneity in histomorphology, genetic alterations in the proteome, immune cell infiltration patterns, and clinical behavior. Objectives This study aims to use single-nucleus sequencing on ten samples (four normal, three clear cell renal cell carcinoma (ccRCC), and three chromophobe renal cell carcinoma (chRCC)) to uncover pathogenic origins and prognostic characteristics in patients with RCC. Methods By using two algorithms, inferCNV and k-means, the study explores malignant cells and compares them with the normal group to reveal their origins. Furthermore, we explore the pathogenic factors at the gene level through Summary-data-based Mendelian Randomization and co-localization methods. Based on the relevant malignant markers, a total of 212 machine-learning combinations were compared to develop a prognostic signature with high precision and stability. Finally, the study correlates with clinical data to investigate which cell subtypes may impact patients’ prognosis. Results & conclusion Two main origin tumor cells were identified: Proximal tubule cell B and Intercalated cell type A, which were highly differentiated in epithelial cells, and three gene loci were determined as potential pathogenic genes. The best malignant signature among the 212 prognostic models demonstrated high predictive power in ccRCC: (AUC: 0.920 (1-year), 0.920 (3-year) and 0.930 (5-year) in the training dataset; 0.756 (1-year), 0.828 (3-year), and 0.832 (5-year) in the testing dataset. In addition, we confirmed that LYVE1+ tissue-resident macrophage and TOX+ CD8 significantly impact the prognosis of ccRCC patients, while monocytes play a crucial role in the prognosis of chRCC patients.
Cellular senescence (CS) is characterized by the irreversible cell cycle arrest and plays a key role in aging and diseases, such as cancer. Recent years have witnessed the burgeoning exploration of the intricate relationship between CS and cancer, with CS recognized as either a suppressing or promoting factor and officially acknowledged as one of the 14 cancer hallmarks. However, a comprehensive characterization remains absent from elucidating the divergences of this relationship across different cancer types and its involvement in the multi-facets of tumor development. Here we systematically assessed the cellular senescence of over 10,000 tumor samples from 33 cancer types, starting by defining a set of cancer-associated CS signatures and deriving a quantitative metric representing the CS status, called CS score. We then investigated the CS heterogeneity and its intricate relationship with the prognosis, immune infiltration, and therapeutic responses across different cancers. As a result, cellular senescence demonstrated two distinct prognostic groups: the protective group with eleven cancers, such as LIHC, and the risky group with four cancers, including STAD. Subsequent in-depth investigations between these two groups unveiled the potential molecular and cellular mechanisms underlying the distinct effects of cellular senescence, involving the divergent activation of specific pathways and variances in immune cell infiltrations. These results were further supported by the disparate associations of CS status with the responses to immuno- and chemo-therapies observed between the two groups. Overall, our study offers a deeper understanding of inter-tumor heterogeneity of cellular senescence associated with the tumor microenvironment and cancer prognosis.
Lung cancer stands as the leading cause of global cancer-related mortality, resulting in an approximate annual toll of 1.6 million lives.1Siegel R.L. Miller K.D. Wagle N.S. Jemal A. Cancer statistics, 2023.CA A Cancer J. Clin. 2023; 73: 17-48Crossref PubMed Scopus (1119) Google Scholar Within this context, non-small cell lung cancer (NSCLC) comprises about 85% of cases, encompassing a range of histological subtypes.2Petrella F. Rizzo S. Attili I. Passaro A. Zilli T. Martucci F. Bonomo L. Del Grande F. Casiraghi M. De Marinis F. Spaggiari L. Stage III Non-Small-Cell Lung Cancer: An Overview of Treatment Options.Curr. Oncol. 2023; 30: 3160-3175Crossref PubMed Scopus (2) Google Scholar Given the intricate genetic modifications underlying the emergence and progression of NSCLC, extensive gene expression profiling has pinpointed a multitude of dysregulated genes.3Rotow J. Bivona T.G. Understanding and targeting resistance mechanisms in NSCLC.Nat. Rev. Cancer. 2017; 17: 637-658Crossref PubMed Scopus (571) Google Scholar These discoveries have significantly informed clinical prognostications and predictions of therapeutic responsiveness. An array of mechanisms, including exon deletion, copy-number variation, and epigenetic modification, have been linked to the disruption of gene expression profiles, as demonstrated by numerous studies.4Xue Y. Hou S. Ji H. Han X. Evolution from genetics to phenotype: reinterpretation of NSCLC plasticity, heterogeneity, and drug resistance.Protein Cell. 2017; 8: 178-190Crossref PubMed Scopus (16) Google Scholar Nonetheless, the comprehensive understanding of the driving forces behind gene expression dysregulation in NSCLC remains an ongoing pursuit. Alternative polyadenylation (APA) has been identified as one of the drivers of aberrant gene expression. Extensive APA events were found in multiple cancer types, which could regulate the expression of oncogenes and tumor suppressors. Accumulated evidence indicates that many APA events are cancer-type and cell-state specific. In a recent issue of Molecular Therapy – Nucleic Acids, Huang et al. describe the identification of heterogeneity of alternative polyadenylation events that occurs in multiple cell types by using bioinformatics pipeline and cell line experiments.5Huang K. Zhang Y. Shi X. Yin Z. Zhao W. Huang L. Wang F. Zhou X. Cell-type specific alternative polyadenylation promotes oncogenic gene expression in non-small cell lung cancer progression.Mol. Ther. Nucleic Acids. 2023; 33: 816-831Abstract Full Text Full Text PDF PubMed Scopus (0) Google Scholar In the current study, Huang et al. collected single-cell RNA sequencing data from previously published studies. Then, the authors systematically analyzed APA events in over 40,000 cells, identifying broadly dysregulated APA events across seven distinct cell types. To ascertain the potential implications of these APA events, the authors pinpointed the loss of microRNA (miRNA)-binding sites resulting from shortened 3′ UTRs, as well as the influence of APA-mediated miRNA regulation. Furthermore, the authors discerned the functional significance of APA-associated genes and validated their roles in cancer cell migration and metastasis. In a bid to explore potential interactions with APA-related genes, the authors conducted drug sensitivity tests on lung cancer cell lines using the Genomics of Drug Sensitivity in Cancer (GDSC) database (Figure 1). According to the analyses, Huang et al. identified widespread APA events in NSCLC across various cell types. Among these events, many are specific to certain cell types and regulate the expression of oncogenes. For instance, the authors noted significant upregulation of eight genes in four distinct cell types in NSCLC samples: SPARC, RGS5, and CD59 in cancer-associated fibroblasts (CAFs), IL1RN in myeloid cells, TMBIM6 in alveolar cells, and RPL22, DERL1, and HM13 in B cells. Notably, the expression of SPARC displayed a robust correlation with patient prognosis. Subsequently, the authors explored the patterns of miRNA-binding site loss induced by APA events in these seven cell types. Prior research had already highlighted the loss of multiple miRNA-binding sites due to APA events. In line with these findings, the authors revealed the widespread occurrence of miRNA-binding site loss in genes affected by APA events. As an example, both bioinformatic analysis and experiments demonstrated the loss of the miR-203a-3p.1-binding site due to APA events in SPARC. It was also shown that miR-203a-3p.1 significantly suppressed the expression of SPARC. Furthermore, the evasion of miR-203a-3p.1 led to the inhibition of tumor-suppressor genes SOCS3 and ETS2. Additionally, the authors uncovered interactions between SPARC and genes regulated by epithelial-mesenchymal transition (EMT), such as COL1A1, TGFBR2, and VCAM1. These interactions are strongly correlated with cancer progression and metastasis. The study then confirmed the role of SPARC in lung cancer proliferation and migration through experiments involving SPARC knockdown cell lines. Drug response analysis also revealed that patients with high SPARC expression levels displayed greater sensitivity to cisplatin treatment compared with the low-SPARC-expression group. Therefore, the analysis suggested that patients with high SPARC expression might derive more benefits from cisplatin treatment. The innovation of cell-type-specific APA represents a significant advancement in the field of NSCLC. Currently, our understanding of dynamic APA regulation in different cell types remains rudimentary. The heterogeneity of APA in distinct cell types may yield valuable insights into the mechanisms of carcinogenesis, progression, and metastasis. While SPARC has been extensively studied in various cancer cell lines, its role in tumorigenesis remains controversial due to its highly cell-type-specific functions.6Arnold S.A. Brekken R.A. SPARC: a matricellular regulator of tumorigenesis.J. Cell Commun. Signal. 2009; 3: 255-273Crossref PubMed Scopus (131) Google Scholar A previous study demonstrated that the low expression of SPARC in lung cancer cells results from its aberrant methylation.7Suzuki M. Hao C. Takahashi T. Shigematsu H. Shivapurkar N. Sathyanarayana U.G. Iizasa T. Fujisawa T. Hiroshima K. Gazdar A.F. Aberrant methylation of SPARC in human lung cancers.Br. J. Cancer. 2005; 92: 942-948Crossref PubMed Scopus (62) Google Scholar However, the cell-type-specific function of SPARC in NSCLC remains unknown. This study not only broadens our understanding of the underlying biological mechanisms of NSCLC but also identifies potential targets for predicting the response to cisplatin treatment in patients with NSCLC. Future research should aim to optimize the therapeutic potential of SPARC in CAFs, uncover its response to other therapies, and comprehensively assess its efficacy as a drug target in mouse models and preclinical trials. In conclusion, this study provides novel insights into the mechanisms underpinning drug resistance in patients with NSCLC. All authors actively contributed to the writing of this commentary. The authors declare no competing interests.
DrugGPT presents a ligand design strategy based on the autoregressive model, GPT, focusing on chemical space exploration and the discovery of ligands for specific proteins. Deep learning language models have shown significant potential in various domains including protein design and biomedical text analysis, providing strong support for the proposition of DrugGPT. In this study, we employ the DrugGPT model to learn a substantial amount of protein-ligand binding data, aiming to discover novel molecules that can bind with specific proteins. This strategy not only significantly improves the efficiency of ligand design but also offers a swift and effective avenue for the drug development process, bringing new possibilities to the pharmaceutical domain. In our research, we particularly optimized and trained the GPT-2 model to better adapt to the requirements of drug design. Given the characteristics of proteins and ligands, we redesigned the tokenizer using the BPE algorithm, abandoned the original tokenizer, and trained the GPT-2 model from scratch. This improvement enables DrugGPT to more accurately capture and understand the structural information and chemical rules of drug molecules. It also enhances its comprehension of binding information between proteins and ligands, thereby generating potentially active drug candidate molecules. Theoretically, DrugGPT has significant advantages. During the model training process, DrugGPT aims to maximize the conditional probability and employs the back-propagation algorithm for training, making the training process more stable and avoiding the Mode Collapse problem that may occur in Generative Adversarial Networks in drug design. Furthermore, the design philosophy of DrugGPT endows it with strong generalization capabilities, giving it the potential to adapt to different tasks. In conclusion, DrugGPT provides a forward-thinking and practical new approach to ligand design. By optimizing the tokenizer and retraining the GPT-2 model, the ligand design process becomes more direct and efficient. This not only reflects the theoretical advantages of DrugGPT but also reveals its potential applications in the drug development process, thereby opening new perspectives and possibilities in the pharmaceutical field.
EDITORIAL article Front. Genet., 08 September 2023Sec. Human and Medical Genomics Volume 14 - 2023 | https://doi.org/10.3389/fgene.2023.1286185