To address the rapid growth of scientific publications and data in biomedical research, knowledge graphs (KGs) have become a critical tool for integrating large volumes of heterogeneous data to enable efficient information retrieval and automated knowledge discovery. However, transforming unstructured scientific literature into KGs remains a significant challenge, with previous methods unable to achieve human-level accuracy. Here we used an information extraction pipeline that won first place in the LitCoin Natural Language Processing Challenge (2022) to construct a large-scale KG named iKraph using all PubMed abstracts. The extracted information matches human expert annotations and significantly exceeds the content of manually curated public databases. To enhance the KG’s comprehensiveness, we integrated relation data from 40 public databases and relation information inferred from high-throughput genomics data. This KG facilitates rigorous performance evaluation of automated knowledge discovery, which was infeasible in previous studies. We designed an interpretable, probabilistic-based inference method to identify indirect causal relations and applied it to real-time COVID-19 drug repurposing from March 2020 to May 2023. Our method identified around 1,200 candidate drugs in the first 4 months, with one-third of those discovered in the first 2 months later supported by clinical trials or PubMed publications. These outcomes are very challenging to attain through alternative approaches that lack a thorough understanding of the existing literature. A cloud-based platform ( https://biokde.insilicom.com ) was developed for academic users to access this rich structured data and associated tools. This study presents iKraph, a large-scale biomedical knowledge graph built using an award-winning natural language processing pipeline with expert-level accuracy. Using probabilistic semantic reasoning, iKraph enables automated knowledge discovery with excellent performance.
The structures of metalloproteins are essential for comprehending their functions and interactions. The breakthrough of AlphaFold has made it possible to predict protein structures with experimental accuracy. However, the type of metal ion that a metalloprotein binds and the binding structure are still not readily available, even with the predicted protein structure. In this study, we present DisDock, a deep learning method for predicting protein-metal docking. DisDock takes distogram of randomly initialized protein-ligand configuration as input and outputs the distogram of the predicted binding complex. It combines the U-net architecture with self-attention modules to enhance model performance. Taking inspiration from the physical principle that atoms in closer proximity display a stronger mutual attraction, this predictor capitalizes on geometric information to uncover latent characteristics indicative of atom interactions. To train our model, we employ a high-quality metalloprotein dataset sourced from the Mother of All Databases (MOAD). Experimental results demonstrate that our approach outperforms other existing methods in prediction accuracy for various types of metal ions.
Dynamic mutations in some human genes containing trinucleotide repeats are associated with severe neurodegenerative and neuromuscular disorders—known as Trinucleotide (or Triplet) Repeat Expansion Diseases (TREDs)—which arise when the repeat number of triplets expands beyond a critical threshold. While the mechanisms causing the DNA triplet expansion are complex and remain largely unknown, it is now recognized that the expandable repeats lead to the formation of nucleotide configurations with atypical structural characteristics that play a crucial role in TREDs. These nonstandard nucleic acid forms include single-stranded hairpins, Z-DNA, triplex structures, G-quartets and slipped-stranded duplexes. Of these, hairpin structures are the most prolific and are associated with the largest number of TREDs and have therefore been the focus of recent single-molecule FRET experiments and molecular dynamics investigations. Here, we review the structural and dynamical properties of nucleic acid hairpins that have emerged from these studies and the implications for repeat expansion mechanisms. The focus will be on CAG, GAC, CTG and GTC hairpins and their stems, their atomistic structures, their stability, and the important role played by structural interrupts.
Myotonic dystrophy type 1 is the most frequent form of muscular dystrophy in adults caused by an abnormal expansion of the CTG trinucleotide. Both the expanded DNA and the expanded CUG RNA transcript can fold into hairpins. Co-transcriptional formation of stable RNA·DNA hybrids can also enhance the instability of repeat tracts. We performed molecular dynamics simulations of homoduplexes associated with the disease, d(CTG)n and r(CUG)n, and their corresponding r(CAG)n:d(CTG)n and r(CUG)n:d(CAG)n hybrids that can form under bidirectional transcription and of non-pathological d(GTC)n and d(GUC)n homoduplexes. We characterized their conformations, stability, and dynamics and found that the U·U and T·T mismatches are dynamic, favoring anti–anti conformations inside the helical core, followed by anti–syn and syn–syn conformations. For DNA, the secondary minima in the non-expanding d(GTC)n helices are deeper, wider, and longer-lived than those in d(CTG)n, which constitutes another biophysical factor further differentiating the expanding and non-expanding sequences. The hybrid helices are closer to A-RNA, with the A-T and A-U pairs forming two stable Watson–Crick hydrogen bonds. The neutralizing ion distribution around the non-canonical pairs is also described.
Most of the biomedical knowledge the research community has acquired during the past few decades has been deposited in scientific literature as unstructured text. Converting the unstructured text into the structured form will enable novel methodologies and applications for scientific discovery that can fully harness the power of the existing knowledge. To this end, two fundamental questions need to be addressed: named entity recognition (NER) and relation extraction (RE). NER deals with identifying the concepts or entities in texts, such as diseases, genes/proteins, chemical compounds, etc. while RE aims to extract the relations among these entities. Together, the extracted information forms a knowledge graph (KG) where the nodes are entities in the texts and the edges represent their relationships. KGs can link concepts within existing research to allow researchers to find connections that may have been difficult to discover without them. The LitCoin Natural Language Processing (NLP) Challenge was recently organized by NCATS of NIH and NASA to spur innovation by rewarding the most creative and high-impact uses of biomedical text to create KGs. Our team participated in the challenge and ranked first place. We have applied the methods we developed for the LitCoin NLP challenge to all PubMed abstracts and constructed the largest-scale biomedical KG to date. We show that powerful and versatile query functions can be implemented on top of the KG to enable highly specific and accurate knowledge retrieval and inference of causal and indirect relationships. Citation Format: Xin Sui, Yuan Zhang, Feng Pan, Donghu Sun, Menghan Chung, Jinfeng Zhang. Constructing the largest-scale knowledge graph using all PubMed abstracts and its application for highly specific and accurate knowledge retrieval. [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2023; Part 1 (Regular and Invited Abstracts); 2023 Apr 14-19; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2023;83(7_Suppl):Abstract nr 5421.
The number of biomedical publications is growing at an accelerated speed. This ever-increasing amount of scientific literature has made reading all the published articles regularly impossible even for a very specific research area. A solid grasp of existing literature is essential for coming out with novel and plausible scientific ideas. To bridge the gap between the published scientific findings and our incapability of manually processing them, we need to convert the unstructured text into structured form to enable automated methods to use the structured, machine-readable information to generate novel hypotheses, which can then be manually validated. A plausible approach for converting unstructured text into structured form is to use named entity recognition (NER) and relation extraction (RE) methods to identify the biological entities and extract their relations to construct knowledge graphs (KGs). KGs can link concepts within existing research to allow researchers to find connections that may have been difficult to discover without them. The LitCoin Natural Language Processing (NLP) Challenge was recently organized by NCATS of NIH and NASA to spur innovation by rewarding the most creative and high-impact uses of biomedical, publication-free text to create KGs. Our team participated in the challenge and ranked first place. Using the pipelines developed for the LitCoin NLP challenge, we have constructed the largest-scale biomedical KG using all PubMed articles. We further develop advanced deep-learning methods to predict new links from the constructed KG. We demonstrate the power of this new framework using several examples important for drug discovery. Citation Format: Yuan Zhang, Feng Pan, Xin Sui, Donghu Sun, Menghan Chung, Jinfeng Zhang. Constructing the largest-scale biomedical knowledge graph using all PubMed articles and its application in automated knowledge discovery. [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2023; Part 1 (Regular and Invited Abstracts); 2023 Apr 14-19; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2023;83(7_Suppl):Abstract nr 5366.
Millions of new papers are published in biomedical sciences every year. In many disciplines, it has become impossible to read all the published new papers to learn what is happening in the frontier of a particular area. This gap is widening as the publishing speed has been accelerating in recent years. To address this challenge, one can convert unstructured text data into a structured form, which can then support highly accurate information retrieval, information integration, and automated knowledge discovery. A plausible approach for such a task is to use named entity recognition (NER) and relation extraction (RE) methods to identify important biological entities and extract their relations to construct knowledge graphs (KGs). The LitCoin Natural Language Processing (NLP) Challenge was recently organized by NCATS of NIH and NASA to spur innovation by rewarding the most creative and high-impact uses of biomedical text to create KGs. In addition to entities and relations, the manually annotated LitCoin dataset also contains the annotations of relations being new discoveries or background knowledge. Our team participated in the challenge and ranked first place. The novelty prediction model of our pipeline has achieved an F1 score of 0.90. We have applied our model to all the PubMed abstracts published previously and the newly published ones to extract the novel discoveries in each article. A web portal has been created to allow scientists to view the latest discoveries in cancer research. The web portal is updated daily with versatile visualization tools for cancer researchers to quickly grasp the latest discoveries in a particular area. It also offers powerful functions to explore the existing literature and make sophisticated inferences about causal and indirect relationships. Citation Format: Feng Pan, Yuan Zhang, Xin Sui, Donghu Sun, Menghan Chung, Jinfeng Zhang. Extracting novel knowledge from scientific literature to build a web portal for cancer researchers to keep up with the latest scientific discoveries. [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2023; Part 1 (Regular and Invited Abstracts); 2023 Apr 14-19; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2023;83(7_Suppl):Abstract nr 5365.
AmberTools is a free and open-source collection of programs used to set up, run, and analyze molecular simulations. The newer features contained within AmberTools23 are briefly described in this Application note.
DNA trinucleotide repeat (TRs) expansion beyond a threshold often results in human neurodegenerative diseases. The mechanisms causing expansions remain unknown, although the tendency of TR ssDNA to self-associate into hairpins that slip along their length is widely presumed related. Here we apply single molecule FRET (smFRET) experiments and molecular dynamics simulations to determine conformational stabilities and slipping dynamics for CAG, CTG, GAC and GTC hairpins. Tetraloops are favored in CAG (89%), CTG (89%) and GTC (69%) while GAC favors triloops. We also determined that TTG interrupts near the loop in the CTG hairpin stabilize the hairpin against slipping. The different loop stabilities have implications for intermediate structures that may form when TR-containing duplex DNA opens. Opposing hairpins in the (CAG) center dot (CTG) duplex would have matched stability whereas opposing hairpins in a (GAC) center dot (GTC) duplex would have unmatched stability, introducing frustration in the (GAC) center dot (GTC) opposing hairpins that could encourage their resolution to duplex DNA more rapidly than in (CAG) center dot (CTG) structures. Given that the CAG and CTG TR can undergo large, disease-related expansion whereas the GAC and GTC sequences do not, these stability differences can inform and constrain models of expansion mechanisms of TR regions.
Protein ligand docking is an indispensable tool for computational prediction of protein functions and screening drug candidates. Despite significant progress over the past two decades, it is still a challenging problem, characterized by the still limited understanding of the energetics between proteins and ligands, and the vast conformational space that has to be searched to find a satisfactory solution. In this project, we developed a novel reinforcement learning (RL) approach, the asynchronous advantage actor-critic model (A3C), to address the protein ligand docking problem. The overall framework consists of two models. During the search process, the agent takes an action selected by the actor model based on the current location. The critic model then evaluates this action and predict the distance between the current location and true binding site. Experimental results showed that in both single- and multi-atom cases, our model improves binding site prediction substantially compared to a naïve model. For the single-atom ligand, copper ion (Cu 2+ ), the model predicted binding sites have a median root-mean-square-deviation (RMSD) of 2.39 Å to the true binding sites when starting from random starting locations. For the multi-atom ligand, sulfate ion (SO 4 2− ), the predicted binding sites have a median RMSD of 3.82 Å to the true binding sites. The ligand-specific models built in this study can be used in solvent mapping studies and the RL framework can be readily scaled up to larger and more diverse sets of ligands.
Abstract Large volumes of publications are being produced in biomedical sciences nowadays with ever-increasing speed. To deal with the large amount of unstructured text data, effective natural language processing (NLP) methods need to be developed for various tasks such as document classification and information extraction. BioCreative Challenge was established to evaluate the effectiveness of information extraction methods in biomedical domain and facilitate their development as a community-wide effort. In this paper, we summarize our work and what we have learned from the latest round, BioCreative Challenge VII, where we participated in all five tracks. Overall, we found three key components for achieving high performance across a variety of NLP tasks: (1) pre-trained NLP models; (2) data augmentation strategies and (3) ensemble modelling. These three strategies need to be tailored towards the specific tasks at hands to achieve high-performing baseline models, which are usually good enough for practical applications. When further combined with task-specific methods, additional improvements (usually rather small) can be achieved, which might be critical for winning competitions. Database URL: https://doi.org/10.1093/database/baac066
By blocking the DEK protein, DEK-targeted aptamers (DTAs) can reduce the formation of neutrophil extracellular traps (NETs) to reveal a strong anti-inflammatory efficacy in rheumatoid arthritis. However, the poor stability of DTA has greatly limited its clinical application. Thus, in order to design an aptamer with better stability, DTA was modified by methoxy groups (DTA_OMe) and then the exact DEK-DTA interaction mechanisms were explored through theoretical calculations. The corresponding 2 '-OCH3-modified nucleotide force field was established and the molecular dynamics (MD) simulations were performed. It was proved that the 2 '-OCH3-modification could definitely enhance the stability of DTA on the premise of comparative affinity. Furthermore, the electrostatic interaction contributed the most to the binding of DEK-DTA, which was the primary interaction to maintain stability, in addition to the non-specific interactions between positively-charged residues (e.g., Lys and Arg) of DEK and the negatively-charged phosphate backbone of aptamers. The H-bond network analysis reminded that eight bases could be mutated to probably enhance the affinity of DTA_OMe. Therein, replacing the 29th base from cytosine to thymine of DTA_OMe was theoretically confirmed to be with the best affinity and even better stability. These research studies imply to be a promising new aptamer design strategy for the treatment of inflammatory arthritis.
Loops in proteins play essential roles in protein functions and interactions. The structural characterization of loops is challenging because of their conformational flexibility and relatively poor conservation in multiple sequence alignments. Many experimental and computational approaches have been carried out during the last few decades for loop modeling. Although the latest AlphaFold2 achieved remarkable performance in protein structure predictions, the accuracy of loop regions for many proteins still needs to be improved for downstream applications such as protein function prediction and structure based drug design. In this paper, we proposed two novel deep learning architectures for loop modeling: one uses a combined convolutional neural network (CNN)-recursive neural network (RNN) structure (DeepMUSICS) and the other is based on refinement of histograms using a 2D CNN architecture (DeepHisto). In each of the methods, two types of models, conformation sampling model and energy scoring model, were trained and applied in the loop folding process. Both methods achieved promising results and worth further investigations. Since multiple sequence alignments (MSA) were not used in our architecture, the energy scoring models have less bias from MSA. We believe the methods may serve as good complements for refining AlphaFold2 predicted structures.
Atypical DNA and RNA secondary structures play a crucial role in simple sequence repeat (SSR) diseases, which are associated with a class of neurological and neuromuscular disorders known as "anticipation diseases," where the age of disease onset decreases and the severity of the disease is increased as the intergenerational expansion of the SSR increases. While the mechanisms underlying these diseases are complex and remain elusive, there is a consensus that stable, non-B-DNA atypical secondary structures play an important - if not causative - role. These structures include single-stranded DNA loops and hairpins, G-quartets, Z-DNA, triplex nucleic acid structures, and others. While all of these structures are of interest, structures based on nucleic acid triplexes have recently garnered increased attention as they have been implicated in gene regulation, gene repair, and gene engineering. Our work here focuses on the construction of DNA triplexes and RNA/DNA hybrids formed from GAA/TTC trinucleotide repeats, which underlie Friedreich's ataxia. While there is some software, such as the Discovery Studio Visualizer, that can aid in the initial construction of DNA triple helices, the only option for the triple helix is constrained to be that of an antiparallel pyrimidine for the third strand. In this protocol, we illustrate how to build up more generalized DNA triplexes and DNA/RNA mixed hybrids. We make use of both the Discovery Studio Visualizer and the AMBER simulation package to construct the initial triplexes. Using the steps outlined here, one can - in principle - build up any triple nucleic acid helix with a desired sequence for large-scale molecular dynamics simulation studies.
Solving the half-century old protein structure prediction problem by DeepMind’s AlphaFold is certainly one of the greatest breakthroughs in biology in the twenty first century. This breakthrough paved the way for tackling some previously highly challenging or even infeasible problems in structural biology. In this study, we propose strategies to use AlphaFold to address several fundamental problems: (1) protein engineering by predicting the experimentally measured stability changes using the representations extracted from AlphaFold models; (2) estimating the designability of a given protein structure by combining a protein design method (e.g. ProDCoNN), sequential Monte Carlo, and AlphaFold. The designability of a protein structure is defined as the number of sequences that encode that protein structure.; (3) predicting protein stabilities using natural sequences and designed sequences as training data, and representations extracted from AlphaFold models as input features; and (4) understanding the sequence-structure relationship of proteins by computational mutagenesis and testing the foldability of the mutants by AlphaFold. We found the representations extracted from AlphaFold models can be used to predict the experimentally measured stability changes accurately. For the first time, we have estimated the designability for a few real proteins. For example, the designability of chain A of FLT3 ligand (PDB ID: 1ETE) with 134 residues was estimated as 3.12±2.14E85.
Dominant conformations of F19W 3Aβ11–40 immersed in transmembrane DPPC lipid bilayer submerged in aqueous solution.
Pathogenic DNA secondary structures have been identified as a common and causative factor for expansion in trinucleotide, hexanucleotide, and other simple sequence repeats. These expansions underlie about fifty neurological and neuromuscular disorders known as "anticipation diseases". Cell toxicity and death have been linked to the pathogenic conformations and functional changes of the RNA transcripts, of DNA itself and, when trinucleotides are present in exons, of the translated proteins. We review some of our results for the conformations and dynamics of pathogenic structures for both RNA and DNA, which include mismatched homoduplexes formed by trinucleotide repeats CAG and GAC; CCG and CGG; CTG(CUG) and GTC(GUC); the dynamics of DNA CAG hairpins; mismatched homoduplexes formed by hexanucleotide repeats (GGGGCC) and (GGCCCC); and G-quadruplexes formed by (GGGGCC) and (GGGCCT). We also discuss the dynamics of strand slippage in DNA hairpins formed by CAG repeats as observed with single-molecule Fluorescence Resonance Energy Transfer. This review focuses on the rich behavior exhibited by the mismatches associated with these simple sequence repeat noncanonical structures.
The total number of amino acid sequences that can fold to a target protein structure, known as “designability”, is a fundamental property of proteins that contributes to their structure and function robustness. The highly designable structures always have higher thermodynamic stability, mutational stability, fast folding, regular secondary structures, and tertiary symmetries. Although it has been studied on lattice models for very short chains by exhaustive enumeration, it remains a challenge to estimate the designable quantitatively for real proteins. In this study, we designed a new deep neural network model that samples protein sequences given a backbone structure using sequential Monte Carlo method. The sampled sequences with proper weights were used to estimate the designability of several real proteins. The designed sequences were also tested using the latest AlphaFold2 and RoseTTAFold to confirm their foldabilities. We report this as the first study to estimate the designability of real proteins.
DNA trinucleotide repeats (TRs) can exhibit dynamic expansions by integer numbers of trinucleotides that lead to neurodegenerative disorders. Strand slipped hairpins during DNA replication, repair and/or recombination may contribute to TR expansion. Here, we combine single-molecule FRET experiments and molecular dynamics studies to elucidate slipping dynamics and conformations of (CAG)(n) TR hairpins. We directly resolve slipping by predominantly two CAG units. The slipping kinetics depends on the even/odd repeat parity. The populated states suggest greater stability for 5'-AGCA-3' tetraloops, compared with alternative 5'-CAG-3' triloops. To accommodate the tetraloop, even(odd)-numbered repeats have an even(odd) number of hanging bases in the hairpin stem. In particular, a paired-end tetraloop (no hanging TR) is stable in (CAG)(n) (= even), but such situation cannot occur in (CAG)(n) (= odd), where the hairpin is "frustrated" and slips back and forth between states with one TR hanging at the 5' or 3' end. Trinucleotide interrupts in the repeating CAG pattern associated with altered disease phenotypes select for specific conformers with favorable loop sequences. Molecular dynamics provide atomic-level insight into the loop configurations. Reducing strand slipping in TR hairpins by sequence interruptions at the loop suggests disease-associated variations impact expansion mechanisms at the level of slipped hairpins.