The small GTPase Ras, a central switch in signal transduction of cell growth, is a crucial element in the development of many forms of cancer. Recent studies have shown that membrane association of Ras is essential to its downstream effector recruitment and signal transduction. Here we investigate the membrane association and interaction of K-Ras4B, the most frequently found Ras isoform in tumor cells, using all-atom molecular dynamics simulations totaling 8.5 µs. We find that the isoform-specific, farnesylated hypervariable region (HVR) of Ras plays an important role in the organization and oligomerization of its globular domain (G-domain) on the surface of the membrane. We show that the overall pose of the Ras G-domain on the membrane can be determined by its HVR. The HVR thus can control downstream effector binding by positioning the G-domain on the membrane in a specific orientation. Furthermore, we observe that the HVR alone already transiently dimerizes in the membrane and recruits negatively charged phosphatidylserine lipids around the K-Ras isoform specific poly-lysine region. The observed HVR dimerization potentially can initiate the dimerization of the full K-Ras, which is essential for downstream effector signaling via various pathways. These atomistic details behind the molecular mechanism of the HVR provide novel insight into Ras interaction with membrane and its availability for downstream effectors.
The severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2) replication transcription complex (RTC) is a multi-domain protein responsible for replicating and transcribing the viral mRNA inside a human cell. Attacking RTC function with pharmaceutical compounds is a pathway to treating COVID-19. Conventional tools, e.g. cryo-electron microscopy and all-atom molecular dynamics (AAMD), do not provide sufficiently high resolution or timescale to capture important dynamics of this molecular machine. Consequently, we develop an innovative workflow that bridges the gap between these resolutions, using mesoscale fluctuating finite element analysis (FFEA) continuum simulations and a hierarchy of AI-methods that continually learn and infer features for maintaining consistency between AAMD and FFEA simulations. We leverage a multi-site distributed workflow manager to orchestrate AI, FFEA, and AAMD jobs, providing optimal resource utilization across HPC centers. Our study provides unprecedented access to study the SARS-CoV-2 RTC machinery, while providing general capability for AI-enabled multi-resolution simulations at scale.
Machine learning (ML)-based steering can improve the performance of ensemble-based simulations by allowing for online selection of more scientifically meaningful computations. We present DeepDriveMD, a framework for ML-driven steering of scientific simulations that we have used to achieve orders-ofmagnitude improvements in molecular dynamics (MD) performance via effective coupling of ML and HPC on large parallel computers. We discuss the design of DeepDriveMD and characterize its performance. We demonstrate that DeepDriveMD can achieve between 100-1000x acceleration for protein folding simulations relative to other methods, as measured by the amount of simulated time performed, while covering the same conformational landscape as quantified by the states sampled during a simulation. Experiments are performed on leadershipclass platforms on up to 1020 nodes. The results establish DeepDriveMD as a high-performance framework for ML-driven HPC simulation scenarios, that supports diverse MD simulation and ML back-ends, and which enables new scientific insights by improving the length and time scales accessible with current computing capacity.
For successful signaling, Ras associates with its downstream effector, Raf. Ras dimerization likely plays an important role in this process, as the active form of Raf is known to be a dimer. However, it is still unclear whether Ras dimerizes first, facilitating the Raf dimerization and further downstream signaling, or a different sequence of events is in play. Given the conformational heterogeneity in both Ras and Raf, probing their dimerization mechanism experimentally is difficult. Hence, we used simulations and deep learning methods to characterize the Ras dimerization interface. Given that the signaling mechanism is different for wild-type (WT) and mutants (G12D, D154Q), we probed if the dimerization of WT and mutants are different. We studied the dimerization of a total of 4 systems (WT-WT, WT-G12D, G12D-G12D and D154Q-D154Q), each simulated for 10 μs, for a total of 40 μs. The systems consisted of the globular domain (G-domain), the hypervariable region (HVR) in a POPC:POPS (70:30) all atom membrane. Using latent-space representations learned from these simulations, we determined stable binding interfaces in three of the four systems. Our analyses suggest these dimers tend to be different in terms of proximity to the membrane, their orientation as well as the binding interface. Contact maps show specific residue interactions at the surface of the dimers of the G-domains, and not the linkers, suggesting the dimerization is driven by the contact between the G-domains, and not the position or interaction of the linkers in the membrane. Furthermore, we also characterized the interactions of POPC and POPS lipids with the proteins, using adversarial autoencoders (AAE). Our analyses revealed a highly nonlinear correlation in the interaction landscape of the Ras and the lipid bilayer.
The race to meet the challenges of the global pandemic has served as a reminder that the existing drug discovery process is expensive, inefficient and slow. There is a major bottleneck screening the vast number of potential small molecules to shortlist lead compounds for antiviral drug development. New opportunities to accelerate drug discovery lie at the interface between machine learning methods, in this case developed for linear accelerators, and physics-based methods. The two in silico methods, each have their own advantages and limitations which, interestingly, complement each other. Here, we present an innovative infrastructural development that combines both approaches to accelerate drug discovery. The scale of the potential resulting workflow is such that it is dependent on supercomputing to achieve extremely high throughput. We have demonstrated the viability of this workflow for the study of inhibitors for four COVID-19 target proteins and our ability to perform the required large-scale calculations to identify lead antiviral compounds through repurposing on a variety of supercomputers.
COVID-19 has claimed more than 2.7 × 106 lives and resulted in over 124 × 106 infections. There is an urgent need to identify drugs that can inhibit SARS-CoV-2. We discuss innovations in computational infrastructure and methods that are accelerating and advancing drug design. Specifically, we describe several methods that integrate artificial intelligence and simulation-based approaches, and the design of computational infrastructure to support these methods at scale. We discuss their implementation, characterize their performance, and highlight science advances that these capabilities have enabled.
We develop a generalizable AI-driven workflow that leverages heterogeneous HPC resources to explore the time-dependent dynamics of molecular systems. We use this workflow to investigate the mechanisms of infectivity of the SARS-CoV-2 spike protein, the main viral infection machinery. Our workflow enables more efficient investigation of spike dynamics in a variety of complex environments, including within a complete SARS-CoV-2 viral envelope simulation, which contains 305 million atoms and shows strong scaling on ORNL Summit using NAMD. We present several novel scientific discoveries, including the elucidation of the spike's full glycan shield, the role of spike glycans in modulating the infectivity of the virus, and the characterization of the flexible interactions between the spike and the human ACE2 receptor. We also demonstrate how AI can accelerate conformational sampling across different systems and pave the way for the future application of such methods to additional studies in SARS-CoV-2 and other molecular systems.
Despite the recent availability of vaccines against the acute respiratory syndrome coronavirus 2 (SARS-CoV-2), the search for inhibitory therapeutic agents has assumed importance especially in the context of emerging new viral variants. In this paper, we describe the discovery of a novel non-covalent small-molecule inhibitor, MCULE-5948770040, that binds to and inhibits the SARS-Cov-2 main protease (M pro ) by employing a scalable high throughput virtual screening (HTVS) framework and a targeted compound library of over 6.5 million molecules that could be readily ordered and purchased. Our HTVS framework leverages the U.S. supercomputing infrastructure achieving nearly 91% resource utilization and nearly 126 million docking calculations per hour. Downstream biochemical assays validate this M pro inhibitor with an inhibition constant ( K i ) of 2.9 µ M [95% CI 2.2, 4.0]. Further, using room-temperature X-ray crystallography, we show that MCULE-5948770040 binds to a cleft in the primary binding site of M pro forming stable hydrogen bond and hydrophobic interactions. We then used multiple µ s-timescale molecular dynamics (MD) simulations, and machine learning (ML) techniques to elucidate how the bound ligand alters the conformational states accessed by M pro , involving motions both proximal and distal to the binding site. Together, our results demonstrate how MCULE-5948770040 inhibits M pro and offers a springboard for further therapeutic design. Significance Statement The ongoing novel coronavirus pandemic (COVID-19) has prompted a global race towards finding effective therapeutics that can target the various viral proteins. Despite many virtual screening campaigns in development, the discovery of validated inhibitors for SARS-CoV-2 protein targets has been limited. We discover a novel inhibitor against the SARS-CoV-2 main protease. Our integrated platform applies downstream biochemical assays, X-ray crystallography, and atomistic simulations to obtain a comprehensive characterization of its inhibitory mechanism. Inhibiting M pro can lead to significant biomedical advances in targeting SARS-CoV-2 treatment, as it plays a crucial role in viral replication.
Emerging hardware tailored for artificial intelligence (AI) and machine learning (ML) methods provide novel means to couple them with traditional high performance computing (HPC) workflows involving molecular dynamics (MD) simulations. We propose StreamAI-MD, a novel instance of applying deep learning methods to drive adaptive MD simulation campaigns in a streaming manner. We leverage the ability to run ensemble MD simulations on GPU clusters, while the data from atomistic MD simulations are streamed continuously to AI/ML approaches to guide the conformational search in a biophysically meaningful manner on a wafer-scale AI accelerator. We demonstrate the efficacy of Stream-AI-MD simulations for two scientific use-cases: (1) folding a small prototypical protein, namely beta beta alpha-fold (BBA) FSD-EY and (2) understanding protein-protein interaction (PPI) within the SARS-CoV-2 proteome between two proteins, nsp16 and nsp10. We show that Stream-AI-MD simulations can improve time-to-solution by similar to 50X for BBA protein folding. Further, we also discuss performance trade-offs involved in implementing AI-coupled HPC workflows on heterogeneous computing architectures.
We seek to completely revise current models of airborne transmission of respiratory viruses by providing never-before-seen atomic-level views of the SARS-CoV-2 virus within a respiratory aerosol. Our work dramatically extends the capabilities of multiscale computational microscopy to address the significant gaps that exist in current experimental methods, which are limited in their ability to interrogate aerosols at the atomic/molecular level and thus obscure our understanding of airborne transmission. We demonstrate how our integrated data-driven platform provides a new way of exploring the composition, structure, and dynamics of aerosols and aerosolized viruses, while driving simulation method development along several important axes. We present a series of initial scientific discoveries for the SARS-CoV-2 Delta variant, noting that the full scientific impact of this work has yet to be realized.
The use of ML methods to dynamically steer ensemble-based simulations promises significant improvements in the performance of scientific applications. We present DeepDriveMD, a tool for a range of prototypical ML-driven HPC simulation scenarios, and use it to quantify improvements in the scientific performance of ML-driven ensemble-based applications. We discuss its design and characterize its performance. Motivated by the potential for further scientific improvements and applicability to more sophisticated physical systems, we extend the design of DeepDriveMD to support stream-based communication between simulations and learning methods. It demonstrates a 100x speedup to fold proteins, and performs 1.6x more simulations per unit time, improving resource utilization compared to the sequential framework. Experiments are performed on leadership-class platforms, at scales of up to O(1000) nodes, and for production workloads. We establish DeepDriveMD as a high-performance framework for ML-driven HPC simulation scenarios, that supports diverse simulation and ML back-ends, and Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than ACM must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. SC’21, November 14–19, 2021, St Louis, MO © 2018 Association for Computing Machinery. ACM ISBN 978-1-4503-9999-9/18/06. . . $15.00 https://doi.org/10.1145/1122445.1122456 which enables new scientific insights by improving lengthand time-scale accessed.
Calmodulin (CaM) is a ubiquitous Ca2+ sensing protein that binds to and modulates numerous target proteins and enzymes during cellular signaling processes. A large number of CaM-target complexes have been identified and structurally characterized, revealing a wide diversity of CaM-binding modes. A newly identified target is creatine kinase (CK), a central enzyme in cellular energy homeostasis. This study reports two high-resolution X-ray structures, determined to 1.24 Å and 1.43 Å resolution, of calmodulin in complex with peptides from human brain and muscle CK, respectively. Both complexes adopt a rare extended binding mode with an observed stoichiometry of 1:2 CaM:peptide, confirmed by isothermal titration calorimetry, suggesting that each CaM domain independently binds one CK peptide in a Ca2+-depended manner. While the overall binding mode is similar between the structures with muscle or brain-type CK peptides, the most significant difference is the opposite binding orientation of the peptides in the N-terminal domain. This may extrapolate into distinct binding modes and regulation of the full-length CK isoforms. The structural insights gained in this study strengthen the link between cellular energy homeostasis and Ca2+-mediated cell signaling and may shed light on ways by which cells can ‘fine tune’ their energy levels to match the spatial and temporal demands.
The drug discovery process currently employed in the pharmaceutical industry typically requires about 10 years and $2-3 billion to deliver one new drug. This is both too expensive and too slow, especially in emergencies like the COVID-19 pandemic. In silicomethodologies need to be improved to better select lead compounds that can proceed to later stages of the drug discovery protocol accelerating the entire process. No single methodological approach can achieve the necessary accuracy with required efficiency. Here we describe multiple algorithmic innovations to overcome this fundamental limitation, development and deployment of computational infrastructure at scale integrates multiple artificial intelligence and simulation-based approaches. Three measures of performance are:(i) throughput, the number of ligands per unit time; (ii) scientific performance, the number of effective ligands sampled per unit time and (iii) peak performance, in flop/s. The capabilities outlined here have been used in production for several months as the workhorse of the computational infrastructure to support the capabilities of the US-DOE National Virtual Biotechnology Laboratory in combination with resources from the EU Centre of Excellence in Computational Biomedicine.
Ras proteins are small GTPases involved in key cell signaling pathways regulating cell growth, proliferation and division. Overactivity of Ras proteins has been a hallmark of diverse forms of cancer in humans. Interconversion between the active (GTP-bound) and inactive (GDP-bound) forms is mediated by two key protein partners that bind and interact with Ras on the surface of the membrane: hydrolysis of GTP, and inactivation of Ras, is mediated by GAP, while exchange of GDP by GTP is facilitated by GEF, reactivating Ras. While the G-domain, the globular domain which hydrolyzes GTP, of Ras isoforms is highly conserved sequentially and structurally, the difference lies in the hypervariable region (HVR). The HVR plays a crucial role in anchoring the G-domain into the cytosolic membrane. We used K-Ras as it is the most common isoform whose mutations lead to a number of cancer type. The orientation of the globular domain on the membrane is essential to its ability to interact with downstream effectors. We simulated the GTP- and GDP-bound domains on HMMM membrane composed of 70:30 PS/PC lipids. The G-domain initial orientation in solution was varied by 25-30 degrees with respect to the membrane, to avoid convergence of results due to initial placement bias. While the differences in structure are small in the GTP- and GDP- bound domains, they bind different effectors, therefore their individual orientation on the membrane is important. Our simulations of over 3.5 microseconds show the G-domain, in the absence of the linker samples many conformations on the surface of the membrane, contrasting results from simulations where the protein was bound to the membrane by the HVR, converging to two distinct binding modes. Our findings suggest the linker highly constrains the orientations available for the G-domain to sample.
The cellular membrane constitutes one of the most fundamental compartments of a living cell, where key processes such as selective transport of material and exchange of information between the cell and its environment are mediated by proteins that are closely associated with the membrane. The heterogeneity of lipid composition of biological membranes and the effect of lipid molecules on the structure, dynamics, and function of membrane proteins are now widely recognized. Characterization of these functionally important lipid-protein interactions with experimental techniques is however still prohibitively challenging. Molecular dynamics (MD) simulations offer a powerful complementary approach with sufficient temporal and spatial resolutions to gain atomic-level structural information and energetics on lipid-protein interactions. In this review, we aim to provide a broad survey of MD simulations focusing on exploring lipid-protein interactions and characterizing lipid-modulated protein structure and dynamics that have been successful in providing novel insight into the mechanism of membrane protein function.
John E. Stone合作论文数University of Illinois at Urbana-Champaign4