Accurate prediction of protein-protein interaction sites (PPISs) plays a crucial role in understanding protein function, elucidating disease mechanisms, and facilitating drug target discovery. Although conventional approaches based on sequence or structural features have shown promising results, they still face several challenges. These challenges include oversmoothing in deep graph neural networks (GNNs) and poor generalization to domain-specific data. To address these issues, we propose RGLLA-PPIS, a novel multimodal prediction model that integrates retrieval-augmented learning and residual GNNs for PPIS identification. In RGLLA-PPIS, protein graphs are constructed by combining AlphaFold3 (AF3)-predicted protein structures with multiple sequence-derived features. To effectively capture both local and global spatial dependencies, the model employs equivariant GNN (EGNN) and GCN modules with residual connections, which help alleviate the oversmoothing problem and preserve node-level variability. Moreover, during prediction, we used the retrieval-augmented knowledge provided by the pretrained protein language model (PLM) Evolla and ChatGPT-4o to construct semantic priors to supplement potential functional site information and enhance the generalization capacity of the prediction model. Extensive experiments on benchmark datasets show that RGLLA-PPIS outperforms several state-of-the-art baselines in both accuracy and robustness. Furthermore, comparison with wet-lab results on a domain-specific protein system reveals a strong correspondence between experimental functional sites and the high-probability regions predicted by RGLLA-PPIS. This demonstrates the model's potential to guide real-world protein engineering tasks. The source code can be found at: https://github.com/MiJia-ID/RGLLA-PPIS.
Ethylene, a cornerstone of the petrochemical industry, constitutes over 75
At least 15% of children with cancer have a pathogenic germline variant in a cancer predisposition gene. Studies of germline cancer predisposition, however, have focused primarily on single nucleotide and copy number variants, leaving other classes such as transposable elements (TEs) largely unexplored. Although rare pathogenic TE insertions have been implicated in inherited cancer, their contribution to pediatric cancer predisposition remains unknown, partly because these repetitive sequences often require whole-genome sequencing for detection. We characterized the germline TE insertion landscape using non-tumor whole genome sequencing data from 2,334 pediatric cancer and 3,447 controls. Across 5,781 genomes, we identified 96,484 TE insertions, most of which were rare and located in intergenic or intronic regions. While global TE burden did not differ between cases and controls, rare TE insertions were significantly enriched in cancer genes in patients with solid tumors, particularly within 3' untranslated regions (p<0.02). Gene-phenotype concordance analysis identified 19 insertions in genes with established dominant cancer predisposition, representing ~0.8% of cases. Integration of RNA-seq data revealed transcriptional impact for a subset of insertions. Notably, a 3' UTR L1 insertion in the tumor suppressor PTEN disrupted alternative polyadenylation, whereas an SVA insertion in a STIM1 intron induced exonization, generating a novel transcript containing SVA sequence. These findings demonstrate that rare germline TE insertions in cancer predisposition genes can have functional consequences at the RNA level. Incorporating TE detection into genomic workflows may improve identification of cancer predisposition syndromes and expand understanding of noncoding contributions to pediatric cancer susceptibility.
Abstract Short tandem repeats (STRs) are among the most mutable regions of the human genome, yet their somatic mosaicism remains poorly characterized due to the technical challenges of distinguishing genuine mutations from high intrinsic polymorphism and sequencing noise. Here, we introduce BulkMonSTR, a computational framework that combines STR-specific error modelling with machine-learning classification to enable accurate detection of mosaic STR mutations from bulk next-generation sequencing data. BulkMonSTR identifies nucleotide-resolution mutations—including insertions, deletions, and single-nucleotide variants (SNVs)—and supports both control-independent and case-control study designs. Leveraging a comprehensive training dataset derived from pedigree-based validation and in silico spike-in simulations, our random forest classifier effectively discriminates true mosaic events from germline variants and technical artifacts. Benchmarking on simulated and real datasets demonstrates that BulkMonSTR achieves substantially improved precision and F1 scores across diverse coverages and variant allele frequencies. In normal samples, cancer samples and controlled in silico mixing experiments, BulkMonSTR consistently outperforms existing methods, capturing a broader spectrum of STR mutations—including those arising on non-reference alleles—while achieving high validation rates. By enabling systematic, genome-wide interrogation of STR mosaicism, BulkMonSTR provides a scalable foundation for investigating the contributions of somatic STR mutations to aging and disease.
Background Recent clinical trials are beginning to show the potential for therapeutic cancer vaccines directed against tumor associated antigens. However, there is a limited set of well-defined antigens that are shared across tumors. Noncanonical coding elements in the genome that are usually silenced in normal cells potentially represent a rich source of “dark” antigens when expressed and presented to the immune system. LINE-1 elements are repetitive, virus-like genomic sequences that can copy themselves to new genomic loci, and which are reactivated in many cancers. In this study, we explore ORF1p, a protein encoded by LINE-1 elements, as a potential shared cancer vaccine antigen. Methods We assessed the tumor specificity of LINE-1 expression in large scale public RNA-seq data sets and in a curated database of public immunopeptidomics studies. We validated these findings using cell line and tissue data and performed an in vitro vaccination assay using healthy donor PBMCs to test the immunogenicity of ORF1p. Results We found widespread RNA expression of LINE-1 across many tumor types with esophageal cancer showing the most highly elevated expression. We also found widespread MHC class I presentation of ORF1p-derived peptides in a range of tumor types, including esophageal cancer. We validated the tumor specific protein expression of ORF1p in a small set of 5 tumor, 5 matched normal, and 5 healthy esophageal samples. Finally, we demonstrated that dendritic cells pulsed with ORF1p peptides were able to induce interferon gamma production in healthy donor human T cells after repeated stimulation in vitro . Conclusions We present data demonstrating that dark antigens derived from LINE-1 elements are 1) expressed in a cancer tissue-specific manner, 2) detected in immunopeptidomics data (both publicly available and newly generated for this study) and 3) immunogenic in vitro . The tumor specificity and immunogenicity of ORF1p peptides point to their potential as a cancer vaccine antigen. ### Competing Interest Statement All authors were employed by ROME Therapeutics during the course of this research. No other competing interests are declared.
Malus Mill., a genus of temperate perennial trees with great agricultural and ecological value, has diversified through hybridization, polyploidy and environmental adaptation. Limited genomic resources for wild Malus species have hindered the understanding of their evolutionary history and genetic diversity. We sequenced and assembled 30 high-quality Malus genomes, representing 20 diploids and 10 polyploids across major evolutionary lineages and geographical regions. Phylogenomic analyses revealed ancient gene duplications and conversions, while six newly defined genome types, including an ancestral type shared by polyploid species, facilitated the detection of strong signals for extensive introgressions. The graph-based pan-genome captured shared and species-specific structural variations, facilitating the development of a molecular marker for apple scab resistance. Our pipeline for analyzing selective sweep identified a mutation in MdMYB5 having reduced cold and disease resistance during domestication. This study advances Malus genomics, uncovering genetic diversity and evolutionary insights while enhancing breeding for desirable traits.
The electric industry is an important factor affecting social progress and economic growth. Compared to conventional thermal power generation, hydroelectric power is applied as a clean and sustainable form of power generation. The relatively short time of hydropower development and the high difficulty in obtaining hydropower data have contributed to collecting a small sample size of data for building an accurate energy production model. Therefore, a novel convolutional neural network (CNN) integrating the synthetic minority over-sampling technique (SMOTE) algorithm (SMOTE-CNN) is proposed to forecast and enhance the energy setting of hydroelectric power plants with precise yield predictions. The SMOTE algorithm is applied to extend the small sample data to increase the diversity of the sample. Then, the CNN is used to process hydropower data and establish a prediction model. Ultimately, the proposed method is applied to predict the actual data of hydroelectric power plants to achieve energy savings and improve energy utilization efficiency. Compared with the back propagation neural network (BP), the radial basis function neural network (RBF), the extreme learning machine (ELM), the gated recurrent unit (GRU), and the long short-term memory (LSTM), the SMOTE-CNN achieves the best performance in terms of the mean relative error (MRE) and the root mean square error (RMSE), with the MRE is 0.0429 and the RMSE is 2373.7366. Additionally, the optimized allocation of resources can improve power generation efficiency.
Somatic mobilization of LINE-1 (L1) has been implicated in cancer etiology. We analyzed a recent TCGA data release comprised of nearly 5000 pan-cancer paired tumor-normal whole-genome sequencing (WGS) samples and ~9000 tumor RNA samples. We developed TotalReCall an improved algorithm and pipeline for detection of L1 retrotransposition (RT), finding high correlation between L1 expression and "RT burden" per sample. Furthermore, we mathematically model the dual regulatory roles of p53, where mutations in TP53 disrupt regulation of both L1 expression and retrotransposition. We found those with Li-Fraumeni Syndrome (LFS) heritable TP53 pathogenic and likely pathogenic variants bear similarly high L1 activity compared to matched cancers from patients without LFS, suggesting this population be considered in attempts to target L1 therapeutically. Due to improved sensitivity, we detect over 10 genes beyond TP53 whose mutations correlate with L1, including ATRX, suggesting other, potentially targetable, mechanisms underlying L1 regulation in cancer remain to be discovered.
Protein-DNA binding directly influences the normal functioning of biological processes by regulating gene expression. Accurate identification of binding sites can reveal the mechanisms of protein-DNA interactions and provide a clear direction for drug target development. However, traditional experimental methods are time-consuming and costly, necessitating the development of efficient computational methods. Although existing computational methods have made significant progress in the field of protein binding site prediction, they have difficulty extracting key residue features and atomic-level features. To address this, we propose a novel method, USPDB, based on a U-shaped Equivariant Graph Neural Network(U-EGNNet) and Subgraph Sampling for Protein-DNA Binding Site Prediction. USPDB reformulates the binding site prediction task by converting the protein into a graph and performing a binary classification for each residue. It leverages protein large language models, such as Protrans, ESM2, and ESM3, to extract sequence and structural features. The General Equivariant Transformer (GET) module is employed to capture geometric features of residues and atoms. Additionally, the U-EGNNet, composed of EGNN and Subgraph Sampling, is utilized to preserve more global information while sampling subgraphs that contain key residues for further computation. Experimental results on DNA_test_181 and DNA_test_129 datasets demonstrate that USPDB achieves prediction accuracies of 0.532 and 0.361, respectively, outperforming all baseline methods. Through interpretability analysis, we observed that USPDB effectively focuses on residues within DNA-binding domains without requiring prior knowledge, thereby enhancing the performance of DNA-binding protein prediction. The code is publicly available at the following link: https://github.com/MiJia-ID/USPDB
BACKGROUND:Long Interspersed Nuclear Elements-1 (LINE-1, L1) are transposable elements that make up roughly 17% of the human genome. These elements can copy and insert themselves into new genomic locations (Kazazian and Moran, N Engl J Med 377:361-370, 2017). Typically, LINE-1 is repressed in healthy tissues but may become activated in various human diseases. LINE-1 expression has been associated with aging (Simon, et al., Cell Metab 29:871-885.e5, 2019; De Cecco et al. Nature 566:73-78, 2019; Della Valle et al. Nat Rev Genet 26:1-12, 2025), neurodegenerative disorders (Roy et al., Acta Neuropathol 148:75, 2024;Frost and Dubnau, Annu Rev Neurosci 47:123-143, 2024; Ravel-Godreuil et al. FEBS Lett 595:2733-2755, 2021), cancer (Rodriguez-Martin et al., Nat Genet 52:306-319, 2020; Taylor et al. Cancer Discov 13:2532-2547, 2023; Solovyov et al. Nat Commun 16:2049, 2025), and autoimmune diseases (Rice et al., N Engl J Med 379:2275-2277, 2018), (Carter et al., Arthritis Rheumatol 72:89-99, 2020). Despite the strong association between LINE-1 expression and disease, the regulatory mechanisms controlling the expression of LINE-1-encoded ORF1p and ORF2p and the link between LINE-1 activity and cancer cell survival remain poorly understood. Gaining insights into these regulatory pathways may help elucidate how LINE-1 contributes to disease pathogenesis. RESULTS:To identify upstream regulators of LINE-1 and genes associated with LINE-1 activity-dependent lethality, we developed a dual-reporter system that simultaneously monitors the protein levels of LINE-1-encoded ORF1p and ORF2p (wild-type or catalytically inactive EN/RT mutant). Using genome-wide CRISPR/Cas9-based screens with this system, we identified candidate genes that may influence LINE-1 regulation at multiple levels, including RNA and protein expression. Alongside known factors such as the HUSH complex, the screens revealed additional genes not previously linked to LINE-1 regulation, suggesting possible new regulatory mechanisms for ORF1p and ORF2p expression. We also identified genes whose loss correlated with reduced viability in a manner dependent on LINE-1 activity. These findings collectively provide a broad resource for exploring cellular factors that may modulate LINE-1 expression and activity. CONCLUSION:This study provides a resource for investigating the cellular regulation of LINE-1, highlighting distinct candidate factors that may modulate ORF1p and ORF2p expression and influence LINE-1 activity-associated cytotoxicity. While functional validation of these candidate regulators remains necessary, the findings offer a foundation for future studies aimed at experimentally confirming their roles and elucidating the molecular mechanisms underlying LINE-1 regulation and its potential contributions to disease contexts.
Mobile element insertions (MEI) shape the human genome in both germline and somatic tissues. While inherited MEIs are well characterized, mapping somatic MEIs (sMEI) in non-cancer tissues remains challenging due to their low allelic fraction and repetitive nature. We established an integrative framework for sMEI analysis leveraging modern sequencing technologies and analytical innovations. We first benchmarked sMEI detection and demonstrated advantages of long-read and MEI-targeted sequencing for ultra-low-frequency events using a mixture of well-established cell lines. We then showed that haplotype phasing and donor-specific assemblies refine sMEI detection, effectively distinguishing from germline and false signals in in-silico tumor-normal mixtures. We further developed a source-tracing strategy based on internal sequence variation, expanding the catalogue of active source elements beyond traditional transduction-based methods. Applying this framework to donor tissues, we identified 18 rare somatic L1 insertions, revealing structural and source diversity. Our work provides a foundational framework and biological insight into sMEIs.
Dwarfing rootstocks have transformed the production of cultivated apples; however, the genetic basis of rootstock-induced dwarfing remains largely unclear. We have assembled chromosome-level, near-gapless and haplotype-resolved genomes for the popular dwarfing rootstock ‘M9’, the semi-vigorous rootstock ‘MM106’ and ‘Fuji’, one of the most commonly grown apple cultivars. The apple orthologue of auxin response factor 3 ( MdARF3 ) is in the Dw1 region of ‘M9’, the major locus for rootstock-induced dwarfing. Comparing ‘M9’ and ‘MM106’ genomes revealed a 9,723-bp allele-specific long terminal repeat retrotransposon/gypsy insertion, DwTE , located upstream of MdARF3 . DwTE is cosegregated with the dwarfing trait in two segregating populations, suggesting its prospective utility in future dwarfing rootstock breeding. In addition, our pipeline discovered mobile mRNAs that may contribute to the development of dwarfed scion architecture. Our research provides valuable genomic resources and applicable methodology, which have the potential to accelerate breeding dwarfing rootstocks for apple and other perennial woody fruit trees.
Industrial process data collected by sensors have characteristics of high dimensionality, non-linearity and dynamics. Consequently, the selection extraction is regarded as a critical part for reducing the dimension of the data and removing irrelevant variables of constructing the production prediction model of industrial processes. Therefore, a novel production prediction model using the Boruta algorithm integrating the convolutional neural network-based Transformer (BCT) is proposed in this paper. Primarily, the Boruta algorithm maps nonlinear high-dimensional data to a low-dimensional space to select features that are meaningful to the yield of the industrial process. Then, the features are extracted adaptively using a convolutional neural network (CNN), which is encoded based on the transformer layer to learn relevant information about the different representation spaces. Furthermore, a linear layer with highway connections is employed to obtain prediction results. Finally, the BCT method is applied to establish a realistic production prediction model of actual liquefied petroleum gas plants for energy saving. Compared with back propagation neural networks, the radial basis function, the extreme learning machine, and the transformer based on the CNN, the BCT method achieves a state-of-the-art level. Furthermore, the BCT method provides the operation guidance on the actual liquefied petroleum gas (LPG) production process with increasing the LPG yield by 17.21%, which can improve production efficiency and reducing energy consumption.
Insertion of active retroelements-L1s, Alus, and SVAs-can disrupt proper genome function and lead to various disorders including cancer. However, the role of de novo retroelements (DNRTs) in birth defects and childhood cancers has not been well characterized due to the lack of adequate data and efficient computational tools. Here, we examine whole-genome sequencing data of 3,244 trios from 12 birth defect and childhood cancer cohorts in the Gabriella Miller Kids First Pediatric Research Program. Using an improved version of our tool xTea (x-Transposable element analyzer) that incorporates a deep-learning module, we identified 162 DNRTs, as well as 2 pseudogene insertions. Several variants are likely to be causal, such as a de novo Alu insertion that led to the ablation of a whole exon in the NF1 gene in a proband with brain tumor. We observe a high de novo SVA insertion burden in both high-intolerance loss-of-function genes and exons as well as more frequent de novo Alu insertions of paternal origin. We also identify potential mosaic DNRTs from embryonic stages. Our study reveals the important roles of DNRTs in causing birth defects and predisposition to childhood cancers.
Abstract Transposable elements are a major source of variation in the human genome. SINE-VNTR-Alus (SVAs), the hominid-specific and youngest active retrotransposon family, are highly polymorphic within the human germline and can contribute to disease risk and progression. In fact, polymorphic SVAs have previously been linked to several human diseases, including rare cases of hereditary cancer syndromes such as neurofibromatosis type I and Lynch syndrome. These examples illustrate that the mobility of SVAs in the genome can significantly impact gene function. Here, we sought to evaluate the potential of SVAs inserting into known tumor suppressors and oncogenes. To this end, we surveyed genomic data from Pan-Cancer Analysis of Whole Genomes (PCAWG) and identified 1956 unique germline SVA insertions in the introns of 1933 genes. Among these, we discovered a highly prevalent polymorphic SVA_E (subfamily) in an intronic region of CASP8. Given caspase-8’s canonical role in apoptosis, we hypothesized that this SVA may impact both caspase-8 expression as well as cell death in cancer. We first created a panel of colon and gastric cancer cell lines genotyped by PCR for the presence or absence of the CASP8-SVA (SVA+/+: RKO, NCI-N87, HT29, NCI-H747, OCUM-1; SVA-/-: HCT116, LoVo, SNU-407, 23132/87, KM12). RT-qPCR analysis showed significantly higher CASP8 expression in SVA+/+ cell lines, which was concordant with caspase-8 protein expression by immunoblot. Treatment of these cell lines with cytotoxic chemotherapy (FOLFOXIRI) in vitro demonstrated increased resistance to cell death in SVA+/+ cell lines. Direct induction of caspase-8-mediated apoptosis via TRAIL (TNF-related apoptosis-inducing ligand) revealed decreased TRAIL-mediated cell death in SVA+/+ cell lines; interestingly, caspase-8 enzymatic activity was also decreased in SVA+/+ cell lines upon TRAIL stimulation despite increased pro-caspase-8 expression. To further isolate the effects of the CASP8-SVA, we generated an isogenic cell line (RKO CASP8-SVA knockout) model via CRISPR-Cas9. Deletion of the CASP8-SVA redemonstrated decreased CASP8 expression at the mRNA and protein levels. Treatment with FOLFOXIRI in the CASP8-SVA KO cell line exhibited increased sensitivity compared to empty vector control. In summary, our findings indicate that an intronic SVA insertion in CASP8 increases gene and protein expression but decreases canonical caspase-8 activity, which was associated with increased resistance to FOLFIRINOX or TRAIL induced cell death. Additional interrogation of SVA-induced effects on transcriptional regulation (e.g., epigenetic effects) as well as non-canonical functions of caspase-8 (e.g., activation of NF-κB) is ongoing. To our knowledge this is the first report describing the mechanistic effects of a germline SVA insertion in the context of cancer and therapeutic resistance. Citation Format: Eric W. Lin, Joshua R. Kocher, Chong Chu, Peter J. Park, David T. Ting. An intronic SINE-VNTR-Alu element insertion in CASP8 alters gene expression and confers resistance to induction of cell death in gastrointestinal cancer [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 5675.
AbstractInsertion of active retroelements—L1s,Alus, and SVAs—can disrupt proper genome function and lead to various disorders including cancer. However, the role ofde novoretroelements (DNRTs) in birth defects and childhood cancers has not been well characterized due to the lack of adequate data and efficient computational tools. Here, we examine whole-genome sequencing data of 3,244 trios from 12 birth defect and childhood cancer cohorts in the Gabriella Miller Kids First Pediatric Research Program. Using an improved version of our tool xTea (x-Transposable element analyzer) that incorporates a deep-learning module, we identified 162 DNRTs, as well as 2 pseudogene insertions. Several variants are likely to be causal, such as ade novo Aluinsertion that led to the ablation of a whole exon in theNF1gene in a proband with brain tumor. We observe a highde novoSVA insertion burden in both high-intolerance loss-of-function genes and exons as well as more frequentde novo Aluinsertions of paternal origin. We also identify potential mosaic DNRTs from embryonic stages. Our study reveals the important roles of DNRTs in causing birth defects and predisposition to childhood cancers.
Achieving rapid, efficient, and cost-effective anaerobic digestion (AD) of food waste is a key means to improve the efficiency of food waste treatment. However, in view of the shortage of historical anaerobic digestion data, the limitation of general neural networks in predicting biogas production, and its sensitivity to abnormal variation points, achieving accurate prediction of biogas production is not easy. This paper proposes a novel biogas production prediction model of food waste AD for energy optimization based on the mixup data augmentation integrating an improved global attention mechanism long short-term memory (LSTM). Taking the AD data of the actual factory as samples, the mixup data augmentation is introduced to generate virtual samples with the similar distribution as original samples. Then original samples and generated virtual samples are used as the input of the global attention mechanism LSTM to establish the food waste AD biogas production prediction model. Finally, the proposed method is applied in the biogas production prediction of actual food waste treatment plants. Compared with other industrial modeling models, the experimental results show that the proposed method has the highest prediction accuracy of 0.988, which performs well in predicting biogas production and can effectively guide and timely adjust feed configuration of AD plants.
When somatic cells acquire complex karyotypes, they often are removed by the immune system. Mutant somatic cells that evade immune surveillance can lead to cancer. Neurons with complex karyotypes arise during neurotypical brain development, but neurons are almost never the origin of brain cancers. Instead, somatic mutations in neurons can bring about neurodevelopmental disorders, and contribute to the polygenic landscape of neuropsychiatric and neurodegenerative disease. A subset of human neurons harbors idiosyncratic copy number variants (CNVs, "CNV neurons"), but previous analyses of CNV neurons are limited by relatively small sample sizes. Here, we develop an allele-based validation approach, SCOVAL, to corroborate or reject read-depth based CNV calls in single human neurons. We apply this approach to 2,125 frontal cortical neurons from a neurotypical human brain. SCOVAL identifies 226 CNV neurons, which include a subclass of 65 CNV neurons with highly aberrant karyotypes containing whole or substantial losses on multiple chromosomes. Moreover, we find that CNV location appears to be nonrandom. Recurrent regions of neuronal genome rearrangement contain fewer, but longer, genes.