Among the Poly(ADP-ribose) Polymerase (PARP) family in mammals, PARP1 is the first identified and well-studied member that plays a critical role in DNA damage repair and has been proven to be an effective target for cancer therapy. Here, we have reviewed not only the role of PARP1 in different DNA damage repair pathways, but also the working mechanisms of several PARP inhibitors (PARPi), inhibiting Poly-ADP-ribosylation (PARylation) processing and PAR chains production to trap PARP1 on impaired DNA and inducing Transcription- replication Conflicts (TRCs) by inhibiting the PARP1 activity. This review has systematically summarized the latest clinical application of six authorized PARPi, including olaparib, rucaparib, niraparib, talazoparib, fuzuloparib and pamiparib, in monotherapy and combination therapies with chemotherapy, radiotherapy, and immunotherapy, in different kinds of cancer. Furthermore, probable challenges in PARPi application and drug resistance mechanisms have also been discussed. Despite these challenges, further development of new PARP1 inhibitors appears promising as a valuable approach to cancer treatment.
Esophageal squamous cell carcinoma (ESCC) is characterized by substantial intratumoral heterogeneity and poor clinical prognosis. Although metalloproteins are well-documented to drive ESCC malignant progression, incomplete functional annotation of this protein family significantly impedes the clinical translation of related research outcomes. This study reports the development and validation of a reliable prognostic model via integrating AlphaFold2-predicted iron-Sulfur (Fe-S) Cluster/Zinc (Zn)-binding proteins with ESCC multi-omics data. Nine differentially expressed AlphaFold2-predicted Fe-S/Zn-binding proteins significantly associated with ESCC prognosis were identified through integrated analysis of multi-omics and clinical data from public datasets and independent ESCC cohorts. After systematic evaluation of 117 machine learning combinations, a three-Fe-S/Zn-binding protein Prognostic Signature (FZPS) comprising YPEL5, MIB1 and ELAC2 was constructed, and validated as an independent predictor of poor overall survival across cohorts. High FZPS risk correlates with an immune-excluded, stress-adaptive phenotype with p21-driven inflammation and intrinsic immunotherapy resistance, while low-FZPS tumors harbor more actionable mutations and exhibit enhanced sensitivity to targeted therapy and immunotherapy. In vitro assays confirmed YPEL5 knockdown markedly suppresses ESCC cell viability, proliferation and migration. In conclusion, FZPS is a reliable independent prognostic biomarker guiding precision oncology practice for ESCC.
The advent of single-cell lineage-tracing technologies has enabled the simultaneous profiling of gene expression and lineage barcodes. However, accurate, high-resolution reconstruction of cell lineage trees remains challenging because most existing approaches treat these modalities separately and therefore fail to fully exploit their complementary information. Here we present BiLinT, a Bayesian framework that jointly models multimodal single-cell lineage-tracing data for lineage tree reconstruction. BiLinT integrates barcode evolution (a continuous-time Markov chain) with gene expression dynamics (an Ornstein-Uhlenbeck process) within a unified probabilistic model. Across synthetic and real datasets, BiLinT provides accurate lineage-tree reconstruction and reveals differentiation-associated clonal structure and developmental fate biases.
Gene-environment interaction (G×E) analyses play a crucial role in advancing genetic discovery, addressing missing heritability, and facilitating precision medicine. However, existing G×E methods are mostly designed for cross-sectional data, limiting the utility of longitudinal data. Here we propose SAGELD, a scalable and accurate genome-wide G×E method for longitudinal traits that controls for sample relatedness in large-scale datasets. SAGELD uses matrix projection to construct test statistics and the SPAGRM framework to efficiently control for sample relatedness, achieving 10- to 10,000-fold speedups over existing methods while maintaining greater power than cross-sectional analyses. We evaluated SAGELD through extensive simulations and UK Biobank analyses. Using age and body mass index as environmental exposures, we identified 74 loci with genetic × age interactions and 5 loci with genetic × adiposity interactions in the pooled analysis of longitudinal primary care data and cross-sectional assessment data. These results highlight the advantages of leveraging longitudinal data in G×E analyses.
Understanding the molecular mechanisms of complex diseases requires insight into cellular interactions and protein expression. While large-scale sequencing enables disease subtyping and patient stratification, integrating proteomics and transcriptomics data offers a deeper view of cellular states. Recent methods combine scRNA-seq, which provides broad cellular coverage, with transcriptomics and proteomics co-profiling, which provides more comprehensive molecular measurements. However, many models adopt simplistic strategies for joint analysis. We introduce scProca, a deep generative model that incorporates inter-cellular relationships via cross-attention mechanisms to handle heterogeneous inputs, whether from RNA-seq or co-profiling datasets. scProca achieves state-of-the-art integration and imputation, remains robust under high protein sparsity, generalizes across species and tissues, scales to large datasets, and is compatible with experimental batches, demonstrating strong flexibility for complex experimental settings.
The advent of temporal single-cell RNA sequencing (scRNA-seq) data has enabled in-depth investigation of dynamic processes in heterogeneous multicellular systems. Despite remarkable advancements in computational methods for modeling cellular dynamics, integrating cell-cell interactions (CCIs) into these models remains a major challenge. This is particularly true when dealing with high-dimensional gene expression profiles from large populations of interacting cells, where the intricate interplay between cells can be obscured by data complexity. To address this, we present scIMF, a single-cell deep-generative Interacting Mean Field model that learns collective multicellular dynamics. Leveraging the McKean-Vlasov stochastic differential equation framework, scIMF provides a mathematical foundation for describing interacting multicellular systems, where each cell's evolution depends on the population's empirical distribution. By incorporating a cell-wise attention mechanism, the model efficiently captures nonlocal and asymmetric CCIs, enabling realistic reconstruction of complex intercellular relationships in high-dimensional spaces. Benchmarking across diverse temporal scRNA-seq datasets demonstrates that scIMF outperforms state-of-the-art methods in reconstructing gene expression at unobserved time points and in inferring cellular velocities. Furthermore, scIMF uncovers biologically interpretable, non-reciprocal interaction patterns of cells, providing a principled framework for studying complex, particularly non-equilibrium biological systems.
MOTIVATION:Spatial transcriptomics enables the dissection of tissue heterogeneity within native contexts, yet current platforms are inherently constrained by high sparsity and low signal-to-noise ratios that obscure fine-grained biological signals. Current efforts to recover these signals are limited by image registration dependencies or the inherent context-blindness of implicit neural representations. RESULTS:We introduce the Cell Positioning System (CPS), a context-aware implicit neural representation framework designed to map physical coordinates to high-fidelity spatial transcriptomics via a privileged multi-scale context distillation strategy. CPS treats multi-scale tissue niches as privileged information, employing a teacher network equipped with a multi-scale niche attention mechanism to capture adaptive biological interactions during training. This structural knowledge is explicitly distilled into a student coordinate network, enabling the generation of context-aware expression landscapes solely from spatial coordinates during inference. Benchmarking on the DLPFC dataset demonstrates that CPS achieves state-of-the-art performance in spatial and gene expression imputation and denoising. Furthermore, CPS enables super-resolution to recover high-resolution mouse brain anatomical details and offers interpretability by identifying the scale effective size of biological interactions within human breast cancer tissues. Finally, the framework exhibits superior scalability for large-scale datasets with linear computational complexity. AVAILABILITY:Software is available online at https://github.com/tju-zl/CPS.
Time-series single-cell RNA-sequencing (scRNA-seq) datasets offer unprecedented insights into the dynamics and heterogeneity of cellular systems. These systems exhibit multiscale collective behaviors driven by intricate intracellular gene regulatory networks and intercellular interactions of molecules. However, inferring interacting cell population dynamics from time-series scRNA-seq data remains a significant challenge, as cells are isolated and destroyed during sequencing. To address this, we introduce scIMF, a single-cell deep generative Interacting Mean Field model, designed to learn collective multi-cellular dynamics. Our approach leverages a transformer-enhanced stochastic differential equation network to simultaneously capture cell-intrinsic dynamics and intercellular interactions. Through extensive benchmarking on multiple scRNA-seq datasets, scIMF outperforms existing methods in reconstructing gene expression at held-out time points, demonstrating that modeling cell-cell communication enhances the accuracy of multicellular dynamics characterization.Additionally, our model provides biologically interpretable insights into cell-cell interactions during dynamic processes, offering a powerful tool for understanding complex cellular systems.
Motivation: Single-cell RNA sequencing (scRNA-seq) and cellular indexing of transcriptomes and epitopes by sequencing (CITE-seq) have experienced rapid advancements in recent years, accompanied by the development of numerous methods for analyzing scRNA-seq and CITE-seq data. These innovations have enabled deeper insights into cellular heterogeneity and functional phenotypes. However, analyzing scRNA-seq and CITE-seq data within a unified framework remains a significant challenge in the field of single-cell analysis. Specifically, this challenge centers on two primary objectives: aligning scRNA-seq and CITE-seq cells within an integrated representation space and generating antibody-derived tag (ADT) measurements for scRNA-seq cells. Results: By incorporating interrelationships between cells into a deep generative model with cross-attention, we introduced scProca to integrate and generate single-cell proteomics from transcriptomics. scProca delivers state-of-the-art performance in both integration and generation tasks across benchmark datasets. Furthermore, scProca can accommodate cells across experimental batches, showcasing its flexibility in complex experimental contexts. Availability: The code of scProca is available at https://github.com/xiongbiolab/scProca, and replication for this study is available at https://github.com/ZzzsHuqiaAao/scProca-reproducibility. ### Competing Interest Statement The authors have declared no competing interest.
Single-cell RNA sequencing provides detailed insights into cellular heterogeneity and responses to external stimuli. However, distinguishing inherent cellular variation from extrinsic effects induced by external stimuli remains a major analytical challenge. Here, we present scCausalVI, a causality-aware generative model designed to disentangle these sources of variation. scCausalVI decouples intrinsic cellular states from treatment effects through a deep structural causal network that explicitly models the causal mechanisms governing cell-state-specific responses to external perturbations while accounting for technical variations. Our model integrates structural causal modeling with cross-condition in silico prediction to infer gene expression profiles under hypothetical scenarios. Comprehensive benchmarking demonstrates that scCausalVI outperforms existing methods in disentangling causal relationships, quantifying treatment effects, generalizing to unseen cell types, and separating biological signals from technical variation in multi-source data integration. Applied to COVID-19 datasets, scCausalVI effectively identifies treatment-responsive populations and delineates molecular signatures of cellular susceptibility.
Single-cell RNA sequencing provides detailed insights into cellular heterogeneity and responses to external stimuli. However, distinguishing inherent cellular variation from extrinsic effects induced by external stimuli remains a major analytical challenge. Here, we present scCausalVI, a causality-aware generative model designed to disentangle these sources of variation. scCausalVI decouples intrinsic cellular states from treatment effects through a deep structural causal network that explicitly models the causal mechanisms governing cell-state-specific responses to external perturbations while accounting for technical variations. Our model integrates structural causal modeling with cross-condition in silico prediction to infer gene expression profiles under hypothetical scenarios. Comprehensive benchmarking demonstrates that scCausalVI outperforms existing methods in disentangling causal relationships, quantifying treatment effects, generalizing to unseen cell types, and separating biological signals from technical variation in multi-source data integration. Applied to COVID-19 datasets, scCausalVI effectively identifies treatment-responsive populations and delineates molecular signatures of cellular susceptibility.
Effectively interfering with endoplasmic reticulum (ER) function in tumor cells and simultaneously activating an anti-tumor immune microenvironment to attack the tumor cells are promising strategies for cancer treatment. However, precise ER-stress induction is still a huge challenge. In this study, we synthesized a near-infrared (NIR) probe, NIR-715, which induces tumor cell death and inhibits tumor growth without causing apparent side effects. NIR-715 triggers severe ER stress and immunogenic cell death (ICD) after visible light exposure. NIR-715 induced ICD-associated HMGB1 release in vitro and anti-tumor immune responses, including increased cytotoxic T lymphocyte (GZMB+ CD8+ T cell) infiltration and decreased numbers of exhausted T lymphocytes (PD-L1+ CD8+ T cell). These findings suggest that NIR-715 may be a novel agent for "cold" tumor photodynamic therapy (PDT).Schematic illustration of NIR-715 photodynamic therapy for visible light-triggered, endoplasmic reticulum-targeting antitumor therapy.
LINC00094 as a new supper-enhancer (SE)-related long non-coding RNA is associated with poor overall survival of patients with esophageal squamous cell carcinoma (ESCC). However, the transcriptional regulatory mechanism of LINC00094 and the molecular mechanisms by which LINC00094 affects the phenotype of ESCC remains unclear. Here, we found that LINC00094 promoted the proliferation of ESCC cells both in vitro and in vivo. LINC00094 knockdown significantly reduced the expression profiles of transcription activators including transcription factor 3 (TCF3) and Kruppel like factor 5 (KLF5) and lipid metabolism-related genes. Mechanically, TCF3 and KLF5 formed a core regulatory circuitry (CRC) that bound to the SEs of LINC00094 and to their own SEs to regulate the transcriptional expression in a positive feedback loop. LINC00094 recruited TCF3 and KLF5 to form a ternary complex, which forms a new CRC with TCF3 and KLF5 that regulated its own transcription as well as lipid metabolism-related genes. Knockdown of any or all three genes inhibited the expression of genes related to lipid synthesis and consistently reduced total lipid droplet levels. Treatment with SEs inhibitors (THZ1 and JQ1) effectively inhibited the formation of this CRC and the production of lipid droplets in ESCC cells. The high-risk group of CRC-associated signatures were closely associated with poor prognosis in patients with ESCC. Our findings suggest that LINC00094 is involved in the CRC by forming a complex with TCF3 and KLF5, and this regulation model can affect the phenotype of ESCC cells by controlling the expression of lipid metabolism-related genes. 1. We identified a novel functional lncRNA-LINC00094 for esophageal squamous cell carcinoma. 2. LINC00094 forms a complex with the core transcription factors TCF3 and KLF5, thereby forming a core regulatory circuitry to participate in transcriptional regulation in ESCC. 3. A core regulatory circuitry mediated by LINC00094 regulates lipid metabolism in ESCC. ### Competing Interest Statement The authors have declared no competing interest. * ### Abbreviations AUC : area under the curve ChIP-seq : chromatin immunoprecipitation co-IP : co-immunoprecipitation CRC : core transcriptional regulation circuitry DEG : differentially expressed gene ESCC : esophageal squamous cell carcinoma GEO : Gene Expression Omnibus GO : gene ontology H3K27ac : Acetylation at the 27th lysine residue of the histone H3 protein KLF5 : kruppel like factor 5 IGV : integrative genomics viewer LASSO : least shrinkage and selection operator lncRNA : long non-coding RNA NC : Negative control OS : overall survival RIP : RNA binding protein immunoprecipitation ROC : receiver operating characteristic curve RT-qPCR : reverse transcriptase-quantitative real-time PCR SE : super-enhancer SEM : standard error of the mean TCF3 : transcription factor 3 TF : transcription factor
Motivation Single-cell RNA sequencing (scRNA-seq) has become a valuable tool for studying cellular heterogeneity. However, the analysis of scRNA-seq data is challenging because of inherent noise and technical variability. Existing methods often struggle to simultaneously explore heterogeneity across cells, handle dropout events, and account for batch effects. These drawbacks call for a robust and comprehensive method that can address these challenges and provide accurate insights into heterogeneity at the single-cell level.Results In this study, we introduce scVIC, an algorithm designed to account for variational inference, while simultaneously handling biological heterogeneity and batch effects at the single-cell level. scVIC explicitly models both biological heterogeneity and technical variability to learn cellular heterogeneity in a manner free from dropout events and the bias of batch effects. By leveraging variational inference, we provide a robust framework for inferring the parameters of scVIC. To test the performance of scVIC, we employed both simulated and biological scRNA-seq datasets, either including, or not, batch effects. scVIC was found to outperform other approaches because of its superior clustering ability and circumvention of the batch effects problem.Availability and implementation The code of scVIC and replication for this study are available at https://github.com/HiBearME/scVIC/tree/v1.0.
Aims Resistance to targeted therapy is one of the critical obstacles in cancer management. Resistance to trastuzumab frequently develops in the treatment for HER2+ cancers. The role of protein tyrosine phosphatases (PTPs) in trastuzumab resistance is not well understood. In this study, we aim to identify pivotal PTPs affecting trastuzumab resistance and devise a novel counteracting strategy. Methods Four public datasets were used to screen PTP candidates in relation to trastuzumab responsiveness in HER2+ breast cancer. Tyrosine kinase (TK) arrays were used to identify kinases that linked to protein tyrosine phosphate receptor type O (PTPRO)-enhanced trastuzumab sensitivity. The efficacy of small activating RNA (saRNA) in trastuzumab-conjugated silica nanoparticles was tested for PTPRO upregulation and resistance mitigation in cell models, a transgenic mouse model, and human cancer cell line-derived xenograft models. Results PTPRO was identified as the key PTP which influences trastuzumab responsiveness and patient survival. PTPRO de-phosphorated several TKs, including the previously overlooked substrate ERBB3, thereby inhibiting multiple oncogenic pathways associated with drug resistance. Notably, PTPRO, previously deemed “undruggable,” was effectively upregulated by saRNA-loaded nanoparticles. The upregulated PTPRO simultaneously inhibited ERBB3, ERBB2, and downstream SRC signaling pathways, thereby counteracting trastuzumab resistance. Conclusions Antibody-conjugated saRNA represents an innovative approach for targeting “undruggable” PTPs.
MOTIVATION:Learning cellular dynamics through reconstruction of the underlying cellular potential energy landscape (aka Waddington landscape) from time-series single-cell RNA sequencing (scRNA-seq) data is a current challenge. Prevailing data-driven computational methods can be hampered by the lack of physical principles to guide learning from complex data, resulting in reduced prediction accuracy and interpretability when applied to infer cell population dynamics. RESULTS:Here, we propose PI-SDE, a physics-informed neural stochastic differential equation (SDE) framework that combines the Hamilton-Jacobi (HJ) equation and neural SDE to learn cellular dynamics. Grounded in potential energy theory of biological systems, PI-SDE integrates the principle of least action by enforcing the HJ equation when reconstructing cellular potential energy function. This approach not only facilitates accurate predictions, but also improves interpretability, especially in the reconstructed potential energy landscape. Through benchmarking on two real scRNA-seq datasets, we demonstrate the importance of incorporating the HJ regularization term in dynamic inference, especially in predicting gene expression at held-out time points. Meanwhile, the learned potential energy landscape provides biologically interpretable insights into the process of cell differentiation. Our framework enhances model performance, while maintaining robustness and stability. AVAILABILITY:PI-SDE software is available at https://github.com/QiJiang-QJ/PI-SDE.
Spatially resolved transcriptomics data are being used in a revolutionary way to decipher the spatial pattern of gene expression and the spatial architecture of cell types. Much work has been done to exploit the genomic spatial architectures of cells. Such work is based on the common assumption that gene expression profiles of spatially adjacent spots are more similar than those of more distant spots. However, related work might not consider the nonlocal spatial co-expression dependency, which can better characterize the tissue architectures. Therefore, we propose MuCoST, a Multi-view graph Contrastive learning framework for deciphering complex Spatially resolved Transcriptomic architectures with dual scale structural dependency. To achieve this, we employ spot dependency augmentation by fusing gene expression correlation and spatial location proximity, thereby enabling MuCoST to model both nonlocal spatial co-expression dependency and spatially adjacent dependency. We benchmark MuCoST on four datasets, and we compare it with other state-of-the-art spatial domain identification methods. We demonstrate that MuCoST achieves the highest accuracy on spatial domain identification from various datasets. In particular, MuCoST accurately deciphers subtle biological textures and elaborates the variation of spatially functional patterns.
Lysyl oxidase-like 2 (LOXL2) is a member of the lysyl oxidase family and has the ability to catalyze the cross-linking of extracellular matrix collagen and elastin. High expression of LOXL2 is related to tumor cell proliferation, invasion and metastasis. LOXL2 contains 14 exons. Previous studies have found that LOXL2 has abnormal alternative splicing and exon skipping in a variety of tissues and cells, resulting in a new alternatively-spliced isoform denoted LOXL2Δ13. LOXL2Δ13 lacks LOXL2WT exon 13, but its encoded protein has greater ability to induce tumor cell proliferation, invasion and metastasis. However, the molecular events that produce LOXL2Δ13 are still unclear. In this study, we found that overexpression of the splicing factor hnRNPA1 in cells can regulate the alternative splicing of LOXL2 and increase the expression of LOXL2Δ13. The exonic splicing silencer (ESS) exists at the 3′ splice site (3′ SS) and 5′ splice site (5′ SS) of LOXL2 exon 13. HnRNPA1 can bind to the ESS and inhibit the inclusion of exon 13. The RRM domain of hnRNPA1 and phosphorylation of hnRNPA1 at S91 and S95 are important for the regulation of LOXL2 alternative splicing. These results show that hnRNPA1 is a splicing factor that enhances the production of LOXL2Δ13.