Protein-protein and protein-peptide interactions are fundamental to biological processes, making the accurate prediction of their binding affinity crucial for drug design and mutational analysis. Here, we develop HyBind-NN, a multimodal graph neural network that integrates protein language models (PLMs) with 3D structural and dynamic datasets to predict protein-protein and protein-peptide affinity. First, we demonstrate that combining ESM-2 sequence embeddings with precise 3D Voronoi spatial geometry enables accurate affinity predictions across diverse structural datasets. Next, we show that the inherent limitations of static rigid-body structures can be mitigated through a multi-task learning framework. By utilizing residue-level root mean square fluctuations (RMSF) derived from molecular dynamics (MD) as an auxiliary training target, the model implicitly learns to capture the conformational entropy of flexible peptides without requiring computationally expensive MD simulations during inference. In our benchmarking study, we observe that this multimodal architecture outperforms both purely sequence-based and strictly structural state-of-the-art algorithms, achieving a mean absolute error of 0.89 for pKD (1.12 kcal/mol for ∆G) on the independent benchmark. Finally, we confirmed through ablation analysis that while the PLM provides the dominant predictive signal, geometric representations and dynamic regularization are crucial for resolving subtle conformational rearrangements. This study highlights the synergistic potential of combining PLMs with physics-aware architectures and demonstrates their application towards the robust prediction of intermolecular binding affinity.
The full complement of chromatin-associated proteins-collectively referred to as the chromatome-enables genome functioning in eukaryotes by participating in a wide range of physico-chemical processes. These include mediating diverse specific and nonspecific intermolecular interactions, catalyzing in situ synthesis and modification of macromolecules, facilitating ATP-dependent chromatin remodeling, etc. Despite considerable progress in epigenomics and the structural characterization of many nuclear proteins and their complexes, our understanding of chromatin organization at the proteome scale remains incomplete. This gap hinders the development of a holistic view of genome regulation. In this study, we present a state-of-the-art characterization of the human chromatome based on an integrative meta-analysis of diverse data sources describing the composition, abundance, and sub-nuclear localization of chromatin proteins. This effort is complemented by original analyses of their physico-chemical properties, domain architectures, and interaction patterns. To support and streamline these analyses, we developed a reference dataset of chromatin proteins, integrated with an empirical, function-based classification ontology and an associated interactive web resource-SimChrom-accessible at https://simchrom.intbio.org/. The reference dataset was carefully curated by reconciling data among protein databases, localization, and mass spectrometry-based experimental studies. Sequence-based and AI-assisted structural analyses revealed previously unannotated domains within chromatin proteins that warrant experimental validation, as well as the widespread use of multivalent interaction strategies that underpin chromatin organization. Together, our findings establish a robust framework for future studies aimed at elucidating genome function through detailed analysis of protein-protein and protein-nucleic acid interactions within chromatin.
MOTIVATION:Single-cell chemical perturbation profiling offers a powerful opportunity to organize drugs by shared mechanism-associated transcriptional responses, but observed transcriptional responses are entangled with contextual variation from cell identity, dose and treatment time. As a result, models that perform well in perturbation-response prediction may still learn latent spaces dominated by context-associated structure rather than transferable drug-associated signal. We developed DECANT to learn mechanism-aligned perturbation representations that remain stable across context shifts while preserving response fidelity. RESULTS:DECANT represents each perturbation as a matched treated-control cell set and separates a context-suppressed, mechanism-aligned perturbation representation from context-dependent response information. The resulting mechanism-aligned perturbation space is shaped to support drug-level retrieval and biological interpretation. Under a fixed drug-level unseen-compound benchmark, DECANT achieved the strongest overall response-difference profile among adapted published perturbation models and strong pseudo-bulk baselines across gene- and program-level metrics. Beyond prediction, DECANT produced embeddings that remained stable across changes in dose, cell line and treatment time, recovered drug neighborhoods enriched for shared mechanism-family annotations, and linked these neighborhoods to interpretable downstream consequence programs. Ablation analyses showed that mechanism-context decoupling provided the main signal-separation backbone, whereas retrieval-oriented shaping was critical for organizing local representation-space geometry. These results support DECANT as a framework for learning context-robust, mechanism-aligned perturbation representations from single-cell transcriptional responses, providing a basis for mechanism-aligned perturbation analysis and representation-based compound prioritization. AVAILABILITY AND IMPLEMENTATION:The DECANT web server is publicly available at http://bliulab.net/DECANT. All source code and analysis scripts are available at https://github.com/bliulab/DECANT and archived on Zenodo at https://doi.org/10.5281/zenodo.21216567. SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
During various DNA-centered processes in the cell nucleus, the minimal structural units of chromatin organization, nucleosomes, are often transiently converted to hexasomes and tetrasomes missing one or both H2A/H2B histone dimers, respectively. However, the structural and functional properties of the subnucleosomes and their impact on biological processes in the nuclei are poorly understood. Here, using biochemical approaches, molecular dynamics simulations, single-particle Förster resonance energy transfer microscopy, and nuclear magnetic resonance spectroscopy, we have shown that, surprisingly, removal of both dimers from a nucleosome results in much higher mobility of both histones and DNA in the tetrasome. Accordingly, DNase I footprinting shows that DNA–histone interactions in tetrasomes are greatly compromised, resulting in formation of a much lower barrier to transcribing RNA polymerase II than nucleosomes. The data suggest that tetrasomes are remarkably dynamic structures and their formation can strongly affect various biological processes.
Anti-angiogenic therapy is a clinically validated method for cancer treatment. It was previously revealed that concurrent targeting of angiogenic and death receptor signaling pathways by a multivalent DR5-specific cytokine TRAIL variant DR5-B genetically fused with the effector peptides, SRH-DR5-B-iRGD, enhances solid tumor suppression and prolongs survival. The SRH peptide is aimed at blocking the tumor neoangiogenesis by preventing activation of the VEGFR2 receptor, while the iRGD peptide interferes with the activation of integrin αvβ3, and enhances the tumor penetration. Here, we investigated how the antiangiogenic activity of the SRH-DR5-B-iRGD fusion protein contributes to its antitumor effects. An integrated approach has been applied involving molecular modeling of SRH-DR5-B-iRGD binding to DR5 receptor, optoacoustic (OA) and optical coherence tomography-based microangiography (OCT-MA) imaging of the vessel networks in xenografts of human glioblastoma and pancreatic adenocarcinoma in nude mice, supported by immunohistochemical (IHC) staining for vascularization marker CD31, and in vitro and in vivo bioactivity studies. Molecular modeling has demonstrated that genetic fusion of DR5-B with the SRH and iRGD peptides not only enables the engagement of additional tumor targets VEGFR2 and integrin αvβ3/NRP-1, but also improves the interaction with DR5 receptor. OA imaging of the vessel network in xenograft tumor nodes of human glioblastoma and pancreatic adenocarcinoma displayed a decrease in the vessel fraction in DR5-B-treated xenograft tumors, with the effect being even more pronounced in SRH-DR5-B-iRGD-treated tumor nodes. This data was consistent with the reduction in the number of perfused vessels in DR5-B and SRH-DR5-B-iRGD-treated tumors as quantified by OCT-MA, and also correlate well with the data obtained by IHC staining and tumor growth inhibition. Ameliorated interaction with the DR5 receptor and imparting antiangiogenic properties to the multivalent fusion protein SRH-DR5-B-iRGD resulted in improved antitumor activity compared to DR5-B. Thereby, SRH-DR5-B-iRGD can be considered as a promising candidate for the treatment of vascularized solid tumors.
Artificial intelligence (AI) is revolutionizing the field of drug development, particularly in addressing key challenges such as drug response prediction, drug combination design, drug repositioning, and drug molecule generation. Traditional drug discovery is hindered by long timelines, high costs, and low success rates, necessitating innovative technologies to accelerate the process. AI technologies, such as deep learning, graph neural networks, and generative models, have demonstrated significant potential in enhancing the accuracy of drug response predictions, optimizing drug combination strategies, identifying opportunities for drug repositioning, and generating drug molecules with specific biological activities. These advancements not only accelerate the drug development process but also open up new possibilities for precision medicine. This review discusses the latest applications and developments of AI in drug discovery, highlighting the breakthroughs and challenges AI addresses in drug development. By summarizing the current research progress, this study provides theoretical support and practical guidance for further applications of AI in drug development.
Nucleosomes are fundamental elements of chromatin organization that participate in compacting genomic DNA and serve as targets for the binding of numerous regulatory proteins. Currently, over 500 different nucleosome structures are known. Despite the large number of nucleosome structures, all of them were formed on only about twenty different DNA sequences. Using cryo-electron microscopy, we determined the structure of the nucleosome formed on a high-affinity Widom 603 DNA sequence at 4 Å resolution; an atomic model was built. We proposed an integrative modeling approach to study the nucleosomal DNA unwrapping based on the cryoelectron microscopy (cryo-EM) data. We also demonstrated the DNA unwrapping of the Widom 603 nucleosome using small angle X-ray scattering and single particle Förster resonance energy transfer measurements. Our results are consistent with the asymmetry of nucleosomal DNA unwrapping. Our data revealed the dependence of nucleosome structure and dynamics on the sequence of nucleosomal DNA.
In the Big Data era, a change of paradigm in the use of molecular dynamics is required. Trajectories should be stored under FAIR (findable, accessible, interoperable and reusable) requirements to favor its reuse by the community under an open science paradigm.
The cytokine TRAIL is distinguished by its remarkable ability to preferentially induce apoptosis in transformed, but not in normal, cells. The recombinant TRAIL extracellular domain and other first-generation agonists of DR4 and DR5 death receptors (DRs) have shown very limited antitumor activity in clinical trials. To enhance the antitumor effect, we developed the multitarget recombinant fusion protein SRH-DR5-B-p48 based on the DR5-selective TRAIL variant DR5-B to simultaneously affect tumor cells (DR5-B-mediated apoptosis) and tumor microenvironment, in particular, to suppress angiogenesis. For this purpose, we modeled and produced the recombinant SRH-DR5-B-p48 fusion protein containing antagonistic synthetic peptides (SRH and p48) to VEGFR2 and FGFR1 receptors, respectively. Analysis of molecular trajectories using molecular dynamics methods showed that the SRH and p48 peptides form non-specific temporary contacts with the DR5-B domain. Using enzyme-linked immunosorbent assay, we showed that SRH-DR5-B-p48 was similar to DR5-B in its affinity for the death receptor DR5 and demonstrated a high affinity for VEGFR2 and FGFR1 with nanomolar dissociation constants. SRH-DR5-B-p48 killed tumor cells of various origin more efficiently than DR5-B and destroyed tumor-like structures in 3D cell models, as well as inhibited FGF2-mediated stimulation of fibroblast proliferation. Therefore, the SRH-DR5-B-p48 fusion protein can be considered as a promising agent for the therapy of solid tumors of various origin.
Epigenome engineering, particularly utilizing CRISPR/dCas-based systems, is a powerful strategy to modulate gene expression and genome functioning without altering the DNA sequence. In this review we summarized current achievements and prospects in dCas-mediated epigenome editing, primarily focusing on its applications in biomedicine, but also providing a wider context for its applications in biotechnology. The diversity of CRISPR/dCas architectures is outlined, recent innovations in the design of epigenetic editors and delivery methods are highlighted, and the therapeutic potential across a wide range of diseases, including hereditary, neurodegenerative, and metabolic disorders, is examined. Opportunities for the application of dCas-based tools in animal, agricultural, and industrial biotechnology are also discussed. Despite substantial progress, challenges, such as delivery efficiency, specificity, stability of induced epigenetic modifications, and clinical translation, are emphasized. Future directions aimed at enhancing the efficacy, safety, and practical applicability of epigenome engineering technologies are proposed.
The COVID-19 pandemic has become a serious challenge for the healthcare system and the economy of many states, and understanding the molecular mechanisms of the pathogenesis of this disease has become a significant challenge for modern science. At the same time, for the first time, a number of high-precision and high-throughput methods for analyzing molecular processes were available to scientists, including technologies for studying changes in chromatin at the genomic level. In this review, we discuss various modern methods that have been used or can be used to study changes in the structure and dynamics of chromatin during infection with SARS-CoV-2 and present the results of currently available studies on the role of these changes in the pathogenesis of COVID-19, and in conclusion, we review the currently known molecular mechanisms of chromatin modulation that occur during infection with SARS-CoV-2.
Histone proteins that play an important role in the chromatin dynamics and regulation of gene activity are a key epigenetic factor. They are divided into two broad classes: canonical histones and their variants. The canonical histones are expressed mainly during the S phase of the cell cycle, since they are involved in DNA packaging in the process of cell division. The histone variants are histone genes that are expressed and regulate the chromatin dynamics during the entire cell cycle. Due to the functional and species diversity, different families of variant histones are distinguished. Some proteins are characterized by minor differences from the canonical histones, while others, on the contrary, can have many important structural and functional peculiarities affecting the nucleosome stability and chromatin dynamics. In order to estimate the variability of histones of the H2A family and their effect on the nucleosome structure, we carried out a bioinformatics analysis of amino acid sequences of the H2A family histones. Clustering conducted using a UPGMA method allowed to distinguish two main subfamilies of H2A proteins: short H2A and other H2A variants that demonstrate higher conservatism of amino acid sequences. We also constructed and analyzed multiple alignments for different H2A histone subfamilies. It is important to note that the proteins of short H2A subfamily are not only the least conservative within their family, but also have the peculiarities that have a significant effect on the nucleosome structural properties. In addition, we conducted a phylogenetic analysis of short H2A histones, as a result of which the subfamilies corresponding to the H2A.B, H2A.P, H2A.Q, H2A.L variants were characterized in more detail.
Bgl2p is a major, conservative, constitutive glucanosyltransglycosylase of the yeast cell wall (CW) with amyloid amino acid sequences, strongly non-covalently anchored in CW, but is able to leave it. In the environment, Bgl2p can form fibrils and/or participate in biofilm formation. Despite a long study, the question of how Bgl2p is anchored in CW remains unclear. Earlier, it was demonstrated that Bgl2p lost the ability to attach in CW and to fibrillate after the deletion of nine amino acids in its C-terminal region (CTR). Here, we demonstrated that a Bgl2p anchoring is weakened by substitution Glu-233/Ala in the active center. Using AlphaFold and molecular modeling approach, we demonstrated the role of CTR on Bgl2p attachment and supposed the conformational possibilities determined by the presence or absence of an intramolecular disulfide bond, forming by Cys-310, leading to accessibility of amyloid sequence and β-turns localized in CTR of Bgl2p for protein interactions. We hypothesized the mode of Bgl2p attachment in CW. Using atomic force microscopy, we investigated fibrillar structures formed by peptide V187MANAFSYWQ196 and suggested that it can serve as a factor leading to the induction of amyloid formation during interaction of Bgl2p with other proteins and is of medical interest being located close to the surface of the molecule.
Viral infections, including SARS-CoV-2, are accompanied by signs of systemic inflammation, which can cause long-term sequela for the patient. Time-stable changes in the organism may be caused by epigenetic shifts inherited in a series of cell divisions, in particular, by changes in the DNA methylation profile in cells of various organs and tissues in response to proinflammatory cytokines. IL1B is a key inflammatory factor, and it was shown that CpG methylation level in its promoter can change upon pro-inflammatory stimuli, and that it was associated with significant increase in IL1B expression. In particular, a specific CpG site in the promoter of the IL1B gene located 299 bp upstream from the transcription start site (CpG3) was previously shown to be an important player in these processes. In this study, we examined methylation/demethylation levels of this CpG3 in publicly available genome-wide methylation studies. A total of 15 dataset were analyzed that comprised data from stromal cells in normal and inflammation-associated states, immune cells of healthy young and aging donors, patients during COVID-19 and after recovery. The level of CpG3 demethylation was found to be higher in osteoarthritis samples of cartilage as compared to healthy donors in one dataset. In blood samples of patients with rheumatoid arthritis CpG3 demethylation was also found to be statistically higher than in healthy donors. In COVID-19 studies, blood samples obtained from patients with severe symptoms had higher CpG3 demethylation levels compared to samples obtained from patients with mild symptoms and controls. The level of CpG3 demethylation increased with age in healthy people as judged by whole blood samples. The same dependency was seen for in vitro cultures of mesenchymal cells obtained from healthy donors. Taken together we showed that demethylation level of a single CpG site in IL1B promoter increases in several cell types due to conditions associated with local and systemic inflammation, including SARS-CoV-2 infection, and in aging. These data suggest a possibility that a history of conditions associated with inflammation within an organism may be recorded, preserved, and encoded in its DNA methylation pattern. While the specificity of these “records of inflammation” is an open question, decoding the history of pathological events associated with inflammation that had been faced by the organism is an intriguing possibility.
Histone proteins form the building blocks of chromatin—nucleosomes. Incorporation of alternative histone variants instead of the major (canonical) histones into nucleosomes is a key mechanism enabling epigenetic regulation of genome functioning. In humans, H2A.J is a constitutively expressed histone variant whose accumulation is associated with cell senescence, inflammatory gene expression, and certain cancers. It is sequence-wise very similar to the canonical H2A histones, and its effects on the nucleosome structure and dynamics remain elusive. This study employed all-atom molecular dynamics simulations to reveal atomistic mechanisms of structural and dynamical effects conferred by the incorporation of H2A.J into nucleosomes. We showed that the H2A.J C-terminal tail and its phosphorylated form have unique dynamics and interaction patterns with the DNA, which should affect DNA unwrapping and the availability of nucleosomes for interactions with other chromatin effectors. The dynamics of the L1-loop and the hydrogen bonding patterns inside the histone octamer were shown to be sensitive to single amino acid substitutions, potentially explaining the higher thermal stability of H2A.J nucleosomes. Taken together, our study demonstrated unique dynamical features of H2A.J-containing nucleosomes, which contribute to further understanding of the molecular mechanisms employed by H2A.J in regulating genome functioning.
Understanding the function of eukaryotic genomes, including the human genome, is undoubtedly one of the major scientific challenges of the 21st century. The cornerstone of eukaryotic genome organization is nucleosomes—elementary building blocks of chromatin about 10 nm in size that wrap DNA around an octamer of histone proteins. Nucleosomes are integral players in all genomic processes, including transcription, DNA replication and repair. They mediate genome regulation at the epigenetic level, bridging the discrete nature of the genetic information encoded in DNA with the analog physical nature of the intermolecular interactions required to access that information. Due to their relatively large size and dynamic nature, nucleosomes are difficult objects for experimental characterization. Molecular dynamics (MD) simulations have emerged over the years as a useful tool to complement experimental studies. Particularly in recent years, advances in computing power, refinement of MD force fields and codes have opened up new frontiers in terms of simulation timescales and quality for nucleosomes and related systems. It has become possible to elucidate in atomistic detail their functional dynamics modes such as DNA unwrapping and sliding, to characterize the effects of epigenetic modifications, DNA and protein sequence variation on nucleosome structure and stability, to describe the mechanisms governing nucleosome interactions with chromatin‐associated proteins and the formation of supranucleosome structures. In this review, we systematically analyzed all‐atom MD simulation studies of nucleosomes and related structures published since 2018 and discussed their relevance in the context of older studies, experimental data, and related coarse‐grained and multiscale studies. This article is categorized under: Software > Molecular Modeling Molecular and Statistical Mechanics > Molecular Dynamics and Monte‐Carlo Methods Structure and Mechanism > Computational Biochemistry and Biophysics
In eukaryotic organisms, genomic DNA associates with histone proteins to form nucleosomes. Nucleosomes provide a basis for genome compaction, epigenetic markup, and mediate interactions of nuclear proteins with their target DNA loci. A negatively charged (acidic) patch located on the H2A-H2B histone dimer is a characteristic feature of the nucleosomal surface. The acidic patch is a common site in the attachment of various chromatin proteins, including viral ones. Acidic patch-binding peptides present perspective compounds that can be used to modulate chromatin functioning by disrupting interactions of nucleosomes with natural proteins or alternatively targeting artificial moieties to the nucleosomes, which may be beneficial for the development of new therapeutics. In this work, we used several computational and experimental techniques to improve our understanding of how peptides may bind to the acidic patch and what are the consequences of their binding. Through extensive analysis of the PDB database, histone sequence analysis, and molecular dynamic simulations, we elucidated common binding patterns and key interactions that stabilize peptide–nucleosome complexes. Through MD simulations and FRET measurements, we characterized changes in nucleosome dynamics conferred by peptide binding. Using fluorescence polarization and gel electrophoresis, we evaluated the affinity and specificity of the LANA1-22 peptide to DNA and nucleosomes. Taken together, our study provides new insights into the different patterns of intermolecular interactions that can be employed by natural and designed peptides to bind to nucleosomes, and the effects of peptide binding on nucleosome dynamics and stability.
Chromatin organization plays an important role in regulating the genetic machinery of the cell. A nucleosome is a basic unit of chromatin packaging and harbors about 145 bp of DNA. The packaging of genetic material and its accessibility to transcription enzymes and other regulatory chromatin proteins depend on the positions of nucleosomes. MNase sequencing is used to examine the nucleosome positions in a genome. MNase sequencing data are sufficient for detecting the presence of nucleosomes on a sequence, but their precise locations can be problematic to establish. Additional data filtering and processing are required for accurate determination of nucleosome positions. A combined method was developed using a geometric analysis of the molecular models of nucleosome chains to select the possible nucleosome positions on the basis of MNase sequencing data. The algorithm efficiently eliminates the inaccessible nucleosome chain combinations and conformationally prohibited nucleosome positions.
Nucleosomes are basic building blocks of chromatin, comprising DNA wrapped around an octamer of eight histone proteins. They play key roles in DNA compaction, epigenetic mark up of the genome and actively participate in chromatin dynamics. X-ray and later cryo-EM studies have contributed greatly to our understanding of nucleosome structure, their interactions and dynamics. With over 470 nucleosome containing structures in the Protein Data Bank there is a wealth of information to be gleaned from these structures, especially through comparative analysis. However, due to the variability in their representation (chain naming, residue numbering, other artifacts), these structures cannot be systematically analyzed “as is”. To address this issue, we developed a framework for analyzing and classifying nucleosome structures and their complexes, resulting in the creation of the NucleosomeDB database and the corresponding web-service. NucleosomeDB allows researchers to search, explore, and compare nucleosomes with each other, despite differences in composition and peculiarities of their representation. By utilizing the information contained within the NucleosomeDB, researchers can gain valuable insights into how nucleosomes interact with DNA and other proteins, assess the implications of mutations and protein binding on nucleosome structure. The detailed information contained within NucleosomeDB can contribute to a better understanding of the structure and function of nucleosomes, and ultimately, the functioning of chromatin and gene regulation. NucleosmeDB is freely available at https://nucldb.intbio.org .
EDITORIAL article Front. Mol. Biosci., 07 March 2023Sec. Molecular Biophysics Volume 10 - 2023 | https://doi.org/10.3389/fmolb.2023.1171714