As mentioned in the Editorial of the first edition of the Special Issue on “Feature Papers in BioChem”, biochemistry acts as a key cog in the “clock of the knowledge, permitting that wheels from several science areas move each other” [...]
Allosteric effect correlates amino acid residues with entropy transfer with and within proteins in a protein complex. The solvation effect could play roles in shaping the allosteric effects. Here we investigate multiple levels of global allosteric correlations within and among proteins in a quaternary antibody-toxin complex, including perturbations of water molecules within first and second solvation shells during all atom MD simulations. Staphylococcal enterotoxin B (SEB) is a potent exotoxin. While monoclonal antibodies 6D3 and 14G8 bind SEB, neither confers significant protection individually, as their epitopes are distal to the TCR/MHC-II interface. Intriguingly, their combination results in potent synergistic neutralization. Our analysis reveals that the simultaneous binding of 14G8 and 6D3 exerts long-range allosteric effects, altering residue fluctuations within the SEB-TCR-binding region. Transfer entropy analysis further demonstrated that the antibody combination establishes an allosteric network that directly modulates the TCR-binding interface. Finally, conformational and solvent entropy analyses suggest that synergistic antibody-mediated inhibition of SEB-TCR binding is caused by increasing SEB's entropy and saturated entropic dissipation into solvent. This study highlights the importance to incorporate environment factors into the allosteric mechanism and provides a systematic comparison of antibody-induced allostery in SEB and, through transfer entropy modeling, establishes a dominant role for specific antibodies in regulating antigen dynamics, offering novel insights into synergistic neutralization mechanisms.
The A beta peptide contributes to Alzheimer's disease through various mechanisms, including cell membrane disruption. While the fibrillar structure of A beta(1-42) in aqueous medium has been elucidated, its oligomer structure remains elusive. We have combined Fourier transform infrared (FTIR) spectroscopy, transmission electron microscopy (TEM), solid-state NMR (ssNMR), and molecular dynamics (MD) approaches to achieve a structural model for A beta(1-42) octamer in lipid bilayers. FTIR data identify conformational transitions of A beta(1-42) to a stable beta-sheet structure. ssNMR analysis allows assignment of 38 out of 42 A beta(1-42) residues, with three additional inter-residue contacts to define the tertiary fold. Combined, MD simulations produce a structural model of A beta(1-42) octamers in a novel sushi-roll fold of in-register cross-beta motif with a lipid-filled internal cavity. The membrane-embedded structure of A beta(1-42) and the mode of peptide-lipid interactions provide a better understanding of A beta neurotoxicity.
The UGA-independent substitution of methionine (Met) and cysteine (Cys) with their selenium (Se) analogues, selenomethionine (SeMet) and selenocysteine (Sec), represents a non-canonical but widespread pathway for the biosynthesis of selenium-enriched proteins. Although well-documented across prokaryotes and eukaryotes, the associated cellular adaptive strategies and phenotypes remain poorly understood. Here, we investigated these substitution patterns and their functional consequences in Bifidobacterium longum (B. longum), a probiotic bacterium that adapts efficiently to high Se stress. Using high-resolution mass spectrometry, we systematically identified and compared SeMet and Sec incorporation sites within the B. longum proteome under Se-enriched conditions. SeMet incorporation proved extensive, substituting over 90% of Met residues, with limited cellular damage. Ribosomal proteins exhibited the highest SeMet incorporation, which did not significantly alter the translational rate. In contrast, Sec incorporation was markedly restricted, characterized by significantly fewer substitution sites and lower substitution proportions. This restriction, accompanied by severely delayed bacterial growth, indicates a profound state of cellular stress, which was further corroborated by the upregulation of protein quality control machinery and the remodeling of sulfur metabolism pathways. However, a subset of proteins with a high probability of Sec incorporation remained, primarily found in catalytic enzymes, yet not localized within their active sites. Notably, both SeMet and Sec incorporations occurred preferentially in high-abundance proteins, without distinct sequence preferences. This work provides the first systematic comparison of SeMet and Sec incorporation patterns in a bacterial proteome, establishing a framework to analyze noncanonical Se incorporation and the specific adaptation strategies bacteria employ against environmental Se challenges.
Abstract c-Src, the first identified oncogene, and KRas, one of the most frequently mutated proteins in human cancers, can activate each other to promote tumorigenesis. Although normal KRas proteins cycle between the inactive GDP bound and the active GTP bound forms, oncogenic mutants are predominantly GTP bound. Importantly, c-Src preferentially recognizes and phosphorylates the GTP-bound form of KRas but the molecular mechanism underlying this specific recognition remains unknown. Here, we employ molecular simulation tools to identify the mechanism underlying the c-Src recognition of the GTP-loaded state of the G12D mutant of KRas4B (the most prevalent in humans). Combining extensive all-atom Molecular Dynamics simulations and Markov State Models analysis, we found that the most populated states of GTP-bound KRas maintain more open and dynamic switch regions, facilitating easier access of c-Src to the phosphorylation sites of KRas (Tyr32 and Tyr64). These states are sparsely populated for GDP loaded KRas. Docking calculations refined by molecular dynamics identify two c-Src specific regions (residues 340-359 and 453-473) able to stabilize phosphorylation-competent KRas conformations. Therefore, c-Src engages highly populated macrostates of GTP-loaded KRas, while interactions with the GDP form are limited to rare conformations. These KRas conformations selectively recognized by c-Src constitute privileged targets for the rational design of peptide-based or small-molecule inhibitors that specifically target active KRas4B-G12D while sparing the inactive GDP-bound form.
Alzheimer's disease (AD) is a progressive neurodegenerative disorder characterized by the accumulation of amyloid-β (Aβ) aggregates, which play a central role in disease pathogenesis according to the amyloid cascade hypothesis. While soluble Aβ oligomers and protofibrils have been identified as the most neurotoxic species, their structural heterogeneity has posed significant challenges for therapeutic development. Current antibody therapies targeting Aβ show differential clinical efficacy, but the molecular basis for their selective recognition of various Aβ polymorphs remains unclear. This critical knowledge gap stems from the lack of experimental structures of antibody-oligomer complexes, which hinders rational drug design. In this study, we thoroughly simulated possible interactions between Aβ oligomer and three antibodies recently approved for targeting Aβ as AD therapy. Our results reveal fundamental differences in their recognition mechanisms. Aducanumab shows polymorph-dependent binding, targeting N-terminal epitopes in full-length Aβ but maintaining non-specific contacts to cross-β structures. Lecanemab uniquely engages multiple N-termini simultaneously through an extended flat-binding interface. Donanemab employs a conserved CDRL1-dominated mode to recognize F4-H13 aggregates, with the pE3 modification acting as a structural anchor that reinforces binding stability. These structural insights provide a molecular basis for observed clinical outcomes and establish design principles for improved therapeutics targeting specific pathological aggregates.
CXCL14 is a highly conserved chemokine with potential roles in tumor progression and immune modulation. This study investigates the functional impact of CXCL14 on colon cancer by exploring its effects on tumor cell behavior and the immune microenvironment. We generated stable cell lines overexpressing CXCL14 in mouse MC38 and CT26 cells and human HCT15 colon cancer cells, and used these models to assess tumor growth, invasion, and immune cell infiltration. Our results demonstrate that CXCL14 suppresses colon cancer cell proliferation, migration, and metastasis. In vitro, CXCL14 inhibited the expression of matrix metalloproteinases (MMPs), key regulators of epithelial-mesenchymal transition (EMT), suggesting a role in promoting mesenchymal-epithelial transition (MET). Additionally, in vivo studies using a subcutaneous tumor model showed that CXCL14 not only suppressed tumor growth but also enhanced the infiltration of immune cells, including NK cells, dendritic cells (DCs), and T cells, converting the tumor microenvironment from a "cold" to a "hot" phenotype. RNA sequencing and pathway analyses revealed that CXCL14 regulates the expression of genes associated with angiogenesis, immune response, and cell signaling, particularly through the MAPK pathway. Furthermore, CXCL14's influence on tumor progression was confirmed in a spleen-to-liver metastasis model, where its overexpression reduced metastatic spread. In conclusion, CXCL14 inhibits colon cancer progression by modulating both tumor cell behavior and the immune landscape, making it a promising candidate for targeted immunotherapy. Our findings highlight CXCL14's potential to enhance anti-tumor immunity and provide new insights into its therapeutic applications in colon cancer.
Predicting Antibody-Antigen (Ab-Ag) docking and structure-based design represent significant long-term and therapeutically important challenges in computational biology. We present SAGERank, a general, configurable deep learning framework for antibody design using Graph Sample and Aggregate Networks. SAGERank successfully predicted the majority of epitopes in a cancer target dataset. In nanobody-antigen structure prediction, SAGERank, coupled with a protein dynamics structure prediction algorithm, slightly outperforms Alphafold3. Most importantly, our study demonstrates the real potential of inductive deep learning to overcome the small dataset problem in molecular science. The SAGERank models trained for antibody-antigen docking can be used to examine general protein-protein interaction tasks, such as T Cell Receptor-peptide-Major Histocompatibility Complex (TCR-pMHC) recognition, classification of biological versus crystal interfaces, and prediction of ternary complexes of molecular glues. In the cases of ranking docking decoys and identifying biological interfaces, SAGERank is competitive with or outperforms state-of-the-art methods.
Biochemistry, or the chemistry of life, is an interdisciplinary science that uses strategies and methods from all exact and natural sciences [...]
Aptamers has drawn significant attention in light of the emerging prominence of nucleic acid-based therapeutics and diagnosis. Aptamers are single-stranded oligonucleotides or short peptides characterized by a distinctive three-dimensional architecture comprising of 20 to 100 nucleotides (nt). They exhibit high affinity and specificity towards target molecules. They have great potential in the detection and medical fields. The SELEX technique is an empirical experimental method. Aptamers obtained by this method are often time-consuming to produce and may have low affinity. With the development of computational technology, artificial intelligence algorithms have demonstrated excellent performance in the field of nucleic acids. Several machine learning approaches have published to predict protein-aptamer interaction. Here, we present SelfTrans-Ensemble, a deep learning model that integrates sequence information models and structural information models to extract multi-scale features for predicting aptamer-protein interactions (APIs). The model employs two pre-trained models, ProtBert and RNA-FM, to encode protein and aptamer sequences, along with features generated from primary sequence and secondary structural information. To address the data imbalance in the aptamer dataset imbalance, we incorporated short RNA-protein interaction data in the training set. We have compiled a dataset consists of 1422 aptamer/RNA sequences and 848 protein sequences, for a total of 1934 aptamer/RNA-protein interaction entries. Our model resulted in a training accuracy of 98.9% and a test accuracy of 88.0%, demonstrating the model*s effectiveness in accurately predicting APIs. We investigated the attention learned for aptamer and protein sequences to explore the enabling residue/nucleotides for APIs, and evaluate if the applied transformer-based network is capable to capture the short-range and long-range dependencies efficiently for aptamer and protein sequences. For a DNA aptamer binding to Von Willebrand Factor (VWF, PDB 3HXO), we found that the attention layer strongly associate with binding correlation, which is consistent with previous structural analyses. Additionally, analysis using molecular simulation indicated that SelfTrans-Ensemble is sensitive to aptamer sequence mutations. SelfTrans-Ensemble exhibits an F1 score of 0.896 and an AUC of 0.9232, indicating that the model is capable of effectively predicting APIs. We further explored the sensitivity of the model by assessing its response to double mutations in RNA sequences and found that the transformer-based model is capable of capturing small mutations in sequences,providing insights of the model*s applicability to facilitate RNA design approach aimed at targeting specific proteins. Our approach holds potential to serve as a rapid and reliable screening approach for binding aptamer sequences towards target proteins, improving the cost-effectiveness and efficiency of SELEX in aptamer screening.
Numerous c-mesenchymal-epithelial transition (c-MET) inhibitors have been reported as potential anticancer agents. However, most fail to enter clinical trials owing to poor efficacy or drug resistance. To date, the scaffold-based chemical space of small-molecule c-MET inhibitors has not been analyzed. In this study, we constructed the largest c-MET dataset, which included 2,278 molecules with different structures, by inhibiting the half maximal inhibitory concentration (IC50) of kinase activity. No significant differences in drug-like properties were observed between active molecules (1,228) and inactive molecules (1,050), including chemical space coverage, physicochemical properties, and absorption, distribution, metabolism, excretion, and toxicity (ADMET) profiles. The higher chemical diversity of the active molecules was downscaled using t-distributed stochastic neighbor embedding (t-SNE) high-dimensional data. Further clustering and chemical space networks (CSNs) analyses revealed commonly used scaffolds for c-MET inhibitors, such as M5, M7, and M8. Activity cliffs and structural alerts were used to reveal "dead ends" and "safe bets" for c-MET, as well as dominant structural fragments consisting of pyridazinones, triazoles, and pyrazines. Finally, the decision tree model precisely indicated the key structural features required to constitute active c-MET inhibitor molecules, including at least three aromatic heterocycles, five aromatic nitrogen atoms, and eight nitrogen-oxygen atoms. Overall, our analyses revealed potential structure-activity relationship (SAR) patterns for c-MET inhibitors, which can inform the screening of new compounds and guide future optimization efforts.
T-cell receptors (TCRs) recognize peptide-MHC (pMHC) complexes through intricate structural interactions, which is a core component of adaptive immunity. However, the diverse and cross-reactive nature of TCRs poses great challenges for accurate prediction of TCR-epitope interactions, hampering the advancement and broad application of TCR-related therapies. Here, we present SageTCR, a bi-level graph neural network (GNN) framework that leverages structural data to predict TCR-pMHC binding possibilities. Harnessing the pretrained language models, SageTCR encodes detailed structural arrangement at both residue-level and atomic-level and effectively integrates the bimodal representations via attention mechanisms. To tackle the deficiency of experimental structures, we explore comprehensive data augmentation strategies to enrich the training and increase the generalizability while concurrently preserving the characteristic TCR-pMHC diagonal binding mode. SageTCR demonstrates superior performance compared to six methods with different deep learning architectures. Furthermore, SageTCR offers the interpretability by identifying and focusing on the conformational features of pivotal contact residues on the interface, which can provide valuable insights for TCR engineering and immunotherapy design.
Protein structure prediction has reached revolutionary levels of accuracy on single structures, implying biophysical energy function can be learned from known protein structures. However apart from single static structure, conformational distributions and dynamics often control protein biological functions. Alphafold2/3 currently predict static protein structure only. Several machine learning approaches have been developed to train model using conformations generated from molecular dynamics (MD) simulations. In this work, we tested a hypothesis that protein energy landscape and conformational dynamics can be learned from experimental structures in PDB and coevolution data. Towards this goal, we develop DeepConformer, a diffusion generative model for sampling protein conformation distributions from a given amino acid sequence. We combined three approaches to allow deep learning techniques to extract hidden dynamics energy landscape information: expanded sequence-structure mapping, large scale 50% structure masking, and MSA clustering.Results: Despite the lack of MD simulation data in training process, DeepConformer captured conformational flexibility and dynamics (RMSF and covariance matrix correlation) similar to MD simulation and reproduced experimentally observed conformational variations.DeepConformer can generate conformation close to different native structures and locate intermediate pathway conformations, as illustrated in the Fold switch of KaiB protein (Figure 1).In the case of large conformation change of an interferon-inducible DNA-sensor protein IFI16, Deepconformer predicted close and open conformer transition, which is hard to obtain using MD simulation.For intrinsically disordered protein, Deepconformer generated conformation ensembles agree with experimental conformation distributions. Our study demonstrated that DeepConformer learned energy landscape can be used to efficiently explore protein conformational distribution and dynamics. DeepConformer-generated structures has similar dynamic properties to that of MD simulation sampled structures and can cover distinct native structures of a single sequence. As DeepConformer achieved this without using protein-specific models or training with data like molecular dynamics trajectories, we expect it to be widely applicable for exploring the conformational dynamics of both natural and designed proteins to understand and optimize their function and regulation. In IVD applications, many biomarkers have flexible and even disordered conformations. Deepconformer can help to understand the biomarkers’ functions and engineering detection methods.
KRAS remains a challenging therapeutic target with limited effective inhibitors currently available. Here, we report the discovery of MCB-294, a potent dual-state pan-KRAS inhibitor capable of binding both the active (GTP-bound) and inactive (GDP-bound) forms of KRAS. MCB-294 engages the switch-II pocket through a water-mediated hydrogen-bond network and selectively inhibits KRAS over NRAS and HRAS. It effectively suppresses oncogenic KRAS signaling, inhibits the growth of KRAS-dependent cancer cells and patient-derived organoids, and reduces tumor progression in multiple preclinical models. MCB-294 also demonstrates superior activity compared to the inactive-state selective pan-KRAS inhibitor Bl-2865 and the KRASG12D inhibitor MRTX1133. Building upon MCB-294 as a pan-KRAS-targeting warhead, we further develop MCB-36, a von Hippel-Lindau (VHL)-recruiting pan-KRAS degrader that induces sustained KRAS degradation. Notably, both MCB-294 and MCB-36 effectively suppress KRASG12C inhibitor-resistant cancer cells and remodel the tumor immune microenvironment. These findings highlight a promising therapeutic strategy for broadly targeting KRAS-driven tumors and overcoming drug resistance.
Identifying functional sites of RNA, particularly those where small molecules bind, is crucial for understanding related biological processes and advancing drug design. Small molecule therapies, compared to traditional protein-targeted therapies, have the potential to pioneer novel RNA-specific therapeutic strategies. However, the challenge lies in developing accurate and efficient computational methods, requiring novel computational models that can better characterize RNA and precisely predict RNA-small molecule binding sites. In this study, we introduced GATRsite, an efficient deep learning framework leveraging graph attention networks (GATs) and Pretrained RNA Language Models to predict RNA-ligand binding sites. GATRsite regards RNA nucleotides as nodes, and its main component is an RNA graph with nodes that comprehensively incorporates both sequential and structural features. Furthermore, it integrates embeddings derived from advanced Pretrained RNA Language Models, which precisely capture the intricate structural and functional complexities of RNA molecules. GATRsite outperforms other state-of-the-art methods, particularly in terms of recall rates, Matthew's correlation coefficient, and F1 score on benchmark test sets. Moreover, GATRsite exhibits significant robustness regarding the predicted RNA structures. A user-friendly online server for GATRsite is freely available at https://malab.sjtu.edu.cn/GATRsite/.
This study aimed to create a new recombinant virus by modifying the EV-A71 capsid protein, serving as a useful tool and model for studying human Enteroviruses. We developed a new screening method using EV-A71 pseudovirus particles to systematically identify suitable insertion sites and tag types in the VP1 capsid protein. The pseudovirus’s infectivity and replication can be assessed by measuring postinfection luciferase signals. We reported that the site after the 100th amino acid within the VP1 BC loop of EV-A71 is particularly permissive for the insertion of various tags. Notably, the introduction of S and V5 tags at this position had minimal effect on the fitness of the tagged pseudovirus. Furthermore, recombinant infectious EV-A71 strains tagged with S and V5 epitopes were successfully rescued, and the stability of these tags was verified. Computational analysis suggested that viable insertions should be compatible with capsid assembly and receptor binding, whereas non-viable insertions could potentially disrupt the capsid’s binding with heparan sulfate. We expect the tagged recombinant EV-A71 to be a useful tool for studying the various stages of the enterovirus life cycle and for virus purification, immunoprecipitation, and research in immunology and vaccine development. Furthermore, this study serves as a proof of principle and may help develop similar tags in enteroviruses, for which there are fewer available tools.
Aptamers are single-stranded DNA/RNAs or short peptides with unique tertiary structures that selectively bind to specific targets. They have great potential in the detection and medical fields. Here, we present SelfTrans-Ensemble, a deep learning model that integrates sequence information models and structural information models to extract multi-scale features for predicting aptamer-protein interactions (APIs). The model employs two pre-trained models, ProtBert and RNA-FM, to encode protein and aptamer sequences, along with features generated from primary sequence and secondary structural information. To address the data imbalance in the aptamer dataset imbalance, we incorporated short RNA-protein interaction data in the training set. This resulted in a training accuracy of 98.9
Protein structure prediction has reached revolutionary levels of accuracy on single structures, implying biophysical energy function can be learned from known protein structures. However apart from single static structure, conformational distributions and dynamics often control protein biological functions. In this work, we tested a hypothesis that protein energy landscape and conformational dynamics can be learned from experimental structures in PDB and coevolution data. Towards this goal, we develop DeepConformer, a diffusion generative model for sampling protein conformation distributions from a given amino acid sequence. Despite the lack of molecular dynamics (MD) simulation data in training process, DeepConformer captured conformational flexibility and dynamics (RMSF and covariance matrix correlation) similar to MD simulation and reproduced experimentally observed conformational variations. Our study demonstrated that DeepConformer learned energy landscape can be used to efficiently explore protein conformational distribution and dynamics. ### Competing Interest Statement The authors have declared no competing interest.
Amyloidosis involves the deposition of misfolded proteins. Even though it is caused by different pathogenic mechanisms, in aggregate, it shares similar features. Here, we tested and confirmed a hypothesis that an amyloid antibody can be engineered by a few mutations to target a different species. Amyloid light chain (AL) and β-amyloid peptide (Aβ) are two therapeutic targets that are implicated in amyloid light chain amyloidosis and Alzheimer’s disease, respectively. Though crenezumab, an anti-Aβ antibody, is currently unsuccessful, we chose it as a model to computationally design and prepare crenezumab variants, aiming to discover a novel antibody with high affinity to AL fibrils and to establish a technology platform for repurposing amyloid monoclonal antibodies. We successfully re-engineered crenezumab to bind both Aβ42 oligomers and AL fibrils with high binding affinities. It is capable of reversing Aβ42-oligomers-induced cytotoxicity, decreasing the formation of AL fibrils, and alleviating AL-fibrils-induced cytotoxicity in vitro. Our research demonstrated that an amyloid antibody could be engineered by a few mutations to bind new amyloid sequences, providing an efficient way to reposition a therapeutic antibody to target different amyloid diseases.