KRAS4a and KRAS4b are important regulators of signaling, and their interactions with the plasma membrane are dynamic and influenced by lipid composition. KRAS 4a and 4b have nearly identical globular domains but differ in their membrane-associated hyper variable region (HVR). The functional distinctions between these isoforms remain unclear, particularly with regards to their dependence on specific lipids and the membrane environment. Previous work showed that the membrane orientation of KRAS4b affects its ability to bind to RAF kinase RBDCRD and that the KRAS-RBDCRD complex adopts different poses on the membrane as well as influences the size and composition of the lipid environment. To model differences between KRAS 4a and 4b protein-lipid interactions, we extended the Multiscale Machine-Learned Modeling Infrastructure (MuMMI) to incorporate continuum simulations in the grand canonical ensemble, enabling sampling across macroscopic, coarse-grained, and all-atom resolutions. Using this framework, we systematically altered PIP2 concentrations, KRAS 4a versus 4b, and RAF RBDCRD complexation to assess impacts on membrane-protein interactions and dynamics. Our results reveal that reducing PIP2 shifts and broadens the membrane orientational preference of both KRAS 4b and 4a, with stronger effects on 4b HVR localization versus 4a. We demonstrate that with depletion of the strong negatively charged PIP2 lipid, the less charged phosphatidylserine replaces PIP2. Our findings highlight similarities and distinctions in the dynamics and lipid dependency of KRAS isoforms and suggest that ordering of the local lipid composition by HVRs is a shared property and key modulator of RAS-mediated signaling at the plasma membrane.
To gain molecular and mechanistic insights into initiation of the RAS-RAF signaling cascade, we developed and used a combination of multiscale simulation and experimental approaches. The influence and impact of the membrane on RAS and RAF proteins is a factor we are just beginning to understand and appreciate in more detail. Molecular simulation is an ideal methodology to further study this complicated relationship between the membrane and associated proteins. Our previous work using Multiscale Machine-learned Modeling Infrastructure investigated different lipid compositions solely around the KRAS4b protein and the interplay between protein behavior and these membrane environments. Multiscale Machine-learned Modeling Infrastructure uses machine learning to couple adjacent simulation scales and has been efficiently scaled across some of the world’s largest high-performance computers. Recently, we have expanded this multiresolution framework to include the all-atom simulation scale and to incorporate the RAF RBDCRD domains. Here, we present the overall analysis results from this new simulation campaign comprising a mixture of RAS and RAF RBDCRD proteins. Approximately 35,000 coarse-grained and 10,000 all-atom molecular dynamics simulations were completed, sampled from a variety of protein/lipid composition configurations that were generated from a micron-scale continuum simulation containing hundreds of copies of the proteins.Our studies suggest that orientations of the RAS-RBDCRD complex on the membrane occupy distinct configurational states, and the spatial patterns of lipid arrangements around these different protein states are unique to each state. The extent and size of lipid “fingerprints” imposed on the membrane by the RAS-RBDCRD protein complex are significantly larger than observed for just the RAS protein on its own. These protein complexes strongly associate, but we do not observe statistically significant preferred protein-protein orientations. These observations indicate that spatial colocalization of RAS-RBDCRD proteins in the same vicinity may be assisted by specific membrane environments, acting to increase the probability of signaling complex formation.
Large-scale diffusion MRI tractography remains a significant challenge. Users must orchestrate a complex sequence of instructions that requires many software packages with complex dependencies and high computational costs. We developed MaPPeRTrac, an edge-centric tractography pipeline that simplifies and accelerates this process in a wide range of high-performance computing (HPC) environments. It fully automates either probabilistic or deterministic tractography, starting from a subject's magnetic resonance imaging (MRI) data, including structural and diffusion MRI images, to the edge density image (EDI) of their structural connectomes. Dependencies are containerized with Singularity (now called Apptainer) and decoupled from code to enable rapid prototyping and modification. Data derivatives are organized with the Brain Imaging Data Structure (BIDS) to ensure that they are findable, accessible, interoperable, and reusable following FAIR principles. The pipeline takes full advantage of HPC resources using the Parsl parallel programming framework, resulting in the creation of connectome datasets of unprecedented size. MaPPeRTrac is publicly available and tested on commercial and scientific hardware, so it can accelerate brain connectome research for a broader user community. MaPPeRTrac is available at: https://github.com/LLNL/mappertrac .
Interdependence across time and length scales is common in biology, where atomic interactions can impact larger-scale phenomenon. Such dependence is especially true for a well-known cancer signaling pathway, where the membrane-bound RAS protein binds an effector protein called RAF. To capture the driving forces that bring RAS and RAF (represented as two domains, RBD and CRD) together on the plasma membrane, simulations with the ability to calculate atomic detail while having long time and large length- scales are needed. The Multiscale Machine-Learned Modeling Infrastructure (MuMMI) is able to resolve RAS/RAF protein-membrane interactions that identify specific lipid-protein fingerprints that enhance protein orientations viable for effector binding. MuMMI is a fully automated, ensemble-based multiscale approach connecting three resolution scales: (1) the coarsest scale is a continuum model able to simulate milliseconds of time for a 1 μm2 membrane, (2) the middle scale is a coarse-grained (CG) Martini bead model to explore protein-lipid interactions, and (3) the finest scale is an all-atom (AA) model capturing specific interactions between lipids and proteins. MuMMI dynamically couples adjacent scales in a pairwise manner using machine learning (ML). The dynamic coupling allows for better sampling of the refined scale from the adjacent coarse scale (forward) and on-the-fly feedback to improve the fidelity of the coarser scale from the adjacent refined scale (backward). MuMMI operates efficiently at any scale, from a few compute nodes to the largest supercomputers in the world, and is generalizable to simulate different systems. As computing resources continue to increase and multiscale methods continue to advance, fully automated multiscale simulations (like MuMMI) will be commonly used to address complex science questions.
Introduction:A defendant who is deemed incompetent to stand trial may go through competency restoration consisting of mental health treatment and legal education. Antipsychotics are often used in treatment; however, there is little data examining their role.Methods:This retrospective study included subjects opined competent to stand trial from July 2016 to February 2020 and prescribed an antipsychotic. The primary outcome was difference in time to competency between antipsychotics. Secondary outcomes included difference in time to competency between groups of antipsychotics, difference in length of stay after opined competent based on medication availability in jail, individual antipsychotics, and formulations.Results:There were 117 subjects included for analysis. There were no differences in time to competency between individual antipsychotics, first- and second-generation antipsychotics, or formulations. Length of stay after opined competent was significantly longer for subjects who were prescribed a long-acting injectable antipsychotic (103 days vs 56 days), who were not able to receive their antipsychotic in jail (104 days vs 54 days), or who were prescribed any formulation of paliperidone compared with olanzapine (88 days vs 35 days).Discussion:Since there were no differences in time to competency, patient-specific factors should be used to choose an agent for competency restoration. Length of stay differences are likely related to the antipsychotic access differences between jails and state psychiatric facilities. Therefore, policies related to antipsychotic access should better align between state psychiatric facilities and jails to improve the capacity of the system and provide better care.
The anatomic validity of structural connectomes remains a significant uncertainty in neuroimaging. Edge-centric tractography reconstructs streamlines in bundles between each pair of cortical or subcortical regions. Although edge bundles provides a stronger anatomic embedding than traditional connectomes, calculating them for each region-pair requires exponentially greater computation. We observe that major speedup can be achieved by reducing the number of streamlines used by probabilistic tractography algorithms. To ensure this does not degrade connectome quality, we calculate the identifiability of edge-centric connectomes between test and re-test sessions as a proxy for information content. We find that running PROBTRACKX2 with as few as 1 streamline per voxel per region-pair has no significant impact on identifiability. Variation in identifiability caused by streamline count is overshadowed by variation due to subject demographics. This finding even holds true in an entirely different tractography algorithm using MRTrix. Incidentally, we observe that Jaccard similarity is more effective than Pearson correlation in calculating identifiability for our subject population.
Wastewater-based epidemiology (WBE) is a popular tool for the early indication of community spread of infectious diseases. WBE emerged as an effective tool during the COVID-19 pandemic and has provided meaningful information to minimize the spread of infection. Here, we present a combination of analyses using the correlation of viral gene copies with clinical cases, sequencing of wastewater-derived RNA for the viral mutants, and correlative analyses of the viral gene copies with the bacterial biomarkers. Our study provides a unique platform for potentially using the WBE-derived results to predict the spread of COVID-19 and the emergence of new variants of concern. Further, we observed a strong correlation between the presence of SARS-CoV-2 and changes in the microbial community of wastewater, particularly the significant changes in bacterial genera belonging to the families of Lachnospiraceae and Actinomycetaceae. Our study shows that microbial biomarkers could be utilized as prediction tools for future infectious disease surveillance and outbreak responses. Overall, our comprehensive analyses of viral spread, variants, and novel bacterial biomarkers will add significantly to the growing body of literature on WBE and COVID-19.
Genetic analysis of intra-host viral populations provides unique insight into pre-emergent mutations that may contribute to the genotype of future variants. Clinical samples positive for SARS-CoV-2 collected in California during the first months of the pandemic were sequenced to define the dynamics of mutation emergence as the virus became established in the state. Deep sequencing of 90 nasopharyngeal samples showed that many mutations associated with the establishment of SARS-CoV-2 globally were present at varying frequencies in a majority of the samples, even those collected as the virus was first detected in the US. A subset of mutations that emerged months later in consensus sequences were detected as subconsensus members of intra-host populations. Spike mutations P681H, H655Y, and V1104L were detected prior to emergence in variant genotypes, mutations were detected at multiple positions within the furin cleavage site, and pre-emergent mutations were identified in the nucleocapsid and the envelope genes. Because many of the samples had a very high depth of coverage, a bioinformatics pipeline, "Mappgene", was established that uses both iVar and LoFreq variant calling to enable identification of very low-frequency variants. This enabled detection of a spike protein deletion present in many samples at low frequency and associated with a variant of concern.
Eleven Labs within the US Department of Energy (DOE), National Virtual Biotechnology Laboratory (NVBL), came together as a team to address significant R&D gaps in COVID-19 testing. Beginning in March 2020, the NVBL COVID Testing Team developed an R&D agenda, worked with DOE and other agencies to set priorities, and collaborated to deliver timely results. Priority was given to quick implementation as well as development of novel capabilities for immediate and evolving pandemic needs without placing additional burden on operational performers. Priority elements capitalized on DOE National Laboratory strengths and expertise. The Team delivered: testing and evaluation that enabled decisions on testing options, forwardleaning approaches to prepare for future scale-up needs, and models and experiments that supported prioritization of diagnostic and therapeutic candidates.
Probabilistic MRI diffusion tractography is a sophisticated technique to investigate structural connectomes, but its steep computational cost prevents application to broader research and clinical settings. Major speedup can be achieved by reducing the number of tractography streamlines. To ensure this does not degrade connectome quality, we calculate the identifiability of connectomes between test and retest MRI as a proxy for information content. We find that reducing streamline count by up to two orders of magnitude from prevailing levels in literature has no significant impact on identifiability. Incidentally, we also observe that Jaccard similarity is more effective than Pearson correlation in achieving identifiability. This document was prepared as an account of work sponsored by an agency of the United States government. Neither the United States government nor Lawrence Livermore National Security, LLC, nor any of their employees makes any warranty, expressed or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness of any information, apparatus, product, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States government or Lawrence Livermore National Security, LLC. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States government or Lawrence Livermore National Security, LLC, and shall not be used for advertising or product endorsement purposes .
The advancement of machine learning techniques and the heterogeneous architectures of most current supercomputers are propelling the demand for large multiscale simulations that can automatically and autonomously couple diverse components and map them to relevant resources to solve complex problems at multiple scales. Nevertheless, despite the recent progress in workflow technologies, current capabilities are limited to coupling two scales. In the first-ever demonstration of using three scales of resolution, we present a scalable and generalizable framework that couples pairs of models using machine learning and in situ feedback. We expand upon the massively parallel Multiscale Machine -Learned Modeling Infrastructure (MuMMI), a recent, award-winning workflow, and generalize the framework beyond its original design. We discuss the challenges and learnings in executing a massive multiscale simulation campaign that utilized over 600,000 node hours on Summit and achieved more than 98% GPU occupancy for more than 83% of the time. We present innovations to enable several orders of magnitude scaling, including simultaneously coordinating 24,000 jobs, and managing several TBs of new data per day and over a billion files in total. Finally, we describe the generalizability of our framework and, with an upcoming open-source release, discuss how the presented framework may be used for new applications.