Computational predictions of biomolecular structure via artificial intelligence (AI) based approaches, as exemplified by AlphaFold software, have the potential to model of all life's biomolecules. We performed oligonucleotide structure prediction and gauged the accuracy of the AI-generated models via their agreement with experimental solution-state observables. We find parts of these models in good agreement with experimental data, and others falling short of the ground truth. The latter include internal or capping loops, noncanonical base pairings, and regions involving conformational flexibility, all essential for RNA folding, interactions, and function. We estimate root-mean-square (r.m.s.) errors in predicted nucleotide bond vector orientations ranging between 7° and 30°, with higher accuracies for simpler architectures of individual canonically paired helical stems. These mixed results highlight the necessity of experimental validation of AI-based oligonucleotide model predictions and their current tendency to mimic the training data set rather than reproduce the underlying reality.
Paxillin (PXN) and focal adhesion kinase (FAK) are two major components of the focal adhesion complex, a multiprotein structure linking the intracellular cytoskeleton to the cell exterior. The interaction between the disordered amino-terminal domain of PXN and the carboxyl-terminal targeting domain of FAK (FAT) is necessary and sufficient for localizing FAK to focal adhesions. Furthermore, PXN serves as a platform for recruiting other proteins that together control the dynamic changes needed for cell migration and survival. Here, we show that the PXN N-domain undergoes significant compaction upon FAT binding, forming a 48-kilodalton multimodal complex with four major interconverting states. Although the complex is flexible, each state has unique sets of contacts involving disordered regions that are both highly represented in ensembles and conserved. PXN being a hub protein, the results provide a structural basis for understanding how shifts in the multistate equilibrium (e.g., through ligand binding and phosphorylation) may rewire cellular networks leading to phenotypic changes.
Antibody-based pharmaceuticals are the leading biologic drug platform (> $75B/year).[1] Despite a wealth of information collected on them, there is still a lack of knowledge on their inter-domain structural distributions, which impedes innovation and development. To address this measurement gap, we have developed a new methodology to derive biomolecular structure ensembles from distance distribution measurements via a library of tagged proteins bound to an unlabeled and otherwise unmodified target biologic. We have employed the NIST monoclonal antibody (NISTmAb) reference material as our development platform for use with spin-labeled affinity protein (SLAP) reagents. Using double electron-electron resonance (DEER) spectroscopy, we have determined inter-spin distance distributions in SLAP complexes of both the isolated Fc domain and the intact NISTmAb. Our SLAP reagents offer a general and extendable technology, compatible with any non-isotopically labeled immunoglobulin G class mAb. Integrating molecular simulations with the DEER and solution X-ray scattering measurements, we enable simultaneous determination of structural distributions and dynamics of mAb-based biologics.
The solution viscosity and protein-protein interactions (PPIs) as a function of temperature (4-40 °C) were measured at a series of protein concentrations for a monoclonal antibody (mAb) with different formulation conditions, which include NaCl and sucrose. The flow activation energy (Eη) was extracted from the temperature dependence of solution viscosity using the Arrhenius equation. PPIs were quantified via the protein diffusion interaction parameter (kD) measured by dynamic light scattering, together with the osmotic second virial coefficient and the structure factor obtained through small-angle X-ray scattering. Both viscosity and PPIs were found to vary with the formulation conditions. Adding NaCl introduces an attractive interaction but leads to a significant reduction in the viscosity. However, adding sucrose enhances an overall repulsive effect and leads to a slight decrease in viscosity. Thus, the averaged (attractive or repulsive) PPI information is not a good indicator of viscosity at high protein concentrations for the mAb studied here. Instead, a correlation based on the temperature dependence of viscosity (i.e., Eη) and the temperature sensitivity in PPIs was observed for this specific mAb. When kD is more sensitive to the temperature variation, it corresponds to a larger value of Eη and thus a higher viscosity in concentrated protein solutions. When kD is less sensitive to temperature change, it corresponds to a smaller value of Eη and thus a lower viscosity at high protein concentrations. Rather than the absolute value of PPIs at a given temperature, our results show that the temperature sensitivity of PPIs may be a more useful metric for predicting issues with high viscosity of concentrated solutions. In addition, we also demonstrate that caution is required in choosing a proper protein concentration range to extract kD. In some excipient conditions studied here, the appropriate protein concentration range needs to be less than 4 mg/mL, remarkably lower than the typical concentration range used in the literature.
The mRNA technology has emerged as a rapid modality to develop vaccines during pandemic situations with the potential to protect against endemic diseases. The success of mRNA in producing an antigen is dependent on the ability to deliver mRNA to the cells using a vehicle, which typically consists of a lipid nanoparticle (LNP). Self-amplifying mRNA (SAM) is a synthetic mRNA platform that, besides encoding for the antigen of interest, includes the replication machinery for mRNA amplification in the cells. Thus, SAM can generate many antigen encoding mRNA copies and prolong expression of the antigen with lower doses than those required for conventional mRNA. This work describes the morphology of LNPs containing encapsulated SAM (SAM LNPs), with SAM being three to four times larger than conventional mRNA. We show evidence that SAM changes its conformational structure when encapsulated in LNPs, becoming more compact than the free SAM form. A characteristic "bleb" structure is observed in SAM LNPs, which consists of a lipid-rich core and an aqueous RNA-rich core, both surrounded by a DSPC-rich lipid shell. We used SANS and SAXS data to confirm that the prevalent morphology of the LNP consists of two-core compartments where components are heterogeneously distributed between the two cores and the shell. A capped cylinder core-shell model with two interior compartments was built to capture the overall morphology of the LNP. These findings provide evidence that bleb two-compartment structures can be a representative morphology in SAM LNPs and highlight the need for additional studies that elucidate the role of spherical and bleb morphologies, their mechanisms of formation, and the parameters that lead to a particular morphology for a rational design of LNPs for mRNA delivery.
We develop a multiscale coarse-grain model of the NIST Monoclonal Antibody Reference Material 8671 (NISTmAb) to enable systematic computational investigations of high-concentration physical instabilities such as phase separation, clustering, and aggregation. Our multiscale coarse-graining strategy captures atomic-resolution interactions with a computational approach that is orders of magnitude more efficient than atomistic models, assuming the biomolecule can be decomposed into one or more rigid bodies with known, fixed structures. This method reduces interactions between tens of thousands of atoms to a single anisotropic interaction site. The anisotropic interaction between unique pairs of rigid bodies is precomputed over a discrete set of relative orientations and stored, allowing interactions between arbitrarily oriented rigid bodies to be interpolated from the precomputed table during coarse-grained Monte Carlo simulations. We present this approach for lysozyme and lactoferrin as a single rigid body and for the NISTmAb as three rigid bodies bound by a flexible hinge with an implicit solvent model. This coarse-graining strategy predicts experimentally measured radius of gyration and second osmotic virial coefficient data, enabling routine Monte Carlo simulation of medically relevant concentrations of interacting proteins while retaining atomistic detail. All methodologies used in this work are available in the open-source software Free Energy and Advanced Sampling Simulation Toolkit.
Nonspecific protein-protein interactions (PPIs) are key to understanding the behavior of proteins in solutions. However, experimentally measuring anisotropic PPIs as a function of orientation and distance has been challenging. Here, we propose to measure a new parameter, the generalized second virial coefficient, B22(Q), to address this challenge. B22(Q) can be measured by using small-angle X-ray/neutron scattering (SAXS/SANS) at finite Q values, where Q is the magnitude of the scattering wave vector. We develop the analytical theory here to calculate B22(Q) with any known interprotein potentials including anisotropic interaction potentials. This method overcomes the challenges and limitations of commonly used methods for extracting PPI information, namely, using integral approximations to solve the Ornstein-Zernike equation by fitting SAXS/SANS data. The accuracy of this analytical theory is further evaluated with computer simulations using a model system. Not only can our method greatly extend the capability of SAXS/SANS to investigate PPIs of many proteins, but it is also applicable to a wide variety of colloidal systems where anisotropic interaction potentials are important.
We communicate a feasibility study for high-resolution structural characterization of biomacromolecules in aqueous solution from X-ray scattering experiments measured over a range of scattering vectors (q) that is approximately two orders of magnitude wider than used previously for such systems. Scattering data with such an extended q-range enables the recovery of the underlying real-space atomic pair distribution function, which facilitates structure determination. We demonstrate the potential of this method for biomacromolecules using several types of cyclodextrins (CD) as model systems. We successfully identified deviations of the tilting angles for the glycosidic units in CDs in aqueous solutions relative to their values in the crystalline forms of these molecules. Such level of structural detail is inaccessible from standard small angle scattering measurements. Our results call for further exploration of ultra-wide-angle X-ray scattering measurements for biomacromolecules.
We describe the conformational ensemble of the single-stranded r(UCAAUC) oligonucleotide obtained using extensive molecular dynamics (MD) simulations and Rosetta's FARFAR2 algorithm. The conformations observed in MD consist of A-form-like structures and variations thereof. These structures are not present in the pool generated using FARFAR2. By comparing with available nuclear magnetic resonance (NMR) measurements, we show that the presence of both A-form-like and other extended conformations is necessary to quantitatively explain experimental data. To further validate our results, we measure solution X-ray scattering (SAXS) data on the RNA hexamer and find that simulations result in more compact structures than observed from these experiments. The integration of simulations with NMR via a maximum entropy approach shows that small modifications to the MD ensemble lead to an improved description of the conformational ensemble. Nevertheless, we identify persisting discrepancies in matching experimental SAXS data.
Through an expansive international effort that involved data collection on 12 small-angle X-ray scattering (SAXS) and four small-angle neutron scattering (SANS) instruments, 171 SAXS and 76 SANS measurements for five proteins (ribonuclease A, lysozyme, xylanase, urate oxidase and xylose isomerase) were acquired. From these data, the solvent-subtracted protein scattering profiles were shown to be reproducible, with the caveat that an additive constant adjustment was required to account for small errors in solvent subtraction. Further, the major features of the obtained consensus SAXS data over the q measurement range 0–1 Å−1 are consistent with theoretical prediction. The inherently lower statistical precision for SANS limited the reliably measured q-range to <0.5 Å−1, but within the limits of experimental uncertainties the major features of the consensus SANS data were also consistent with prediction for all five proteins measured in H2O and in D2O. Thus, a foundation set of consensus SAS profiles has been obtained for benchmarking scattering-profile prediction from atomic coordinates. Additionally, two sets of SAXS data measured at different facilities to q > 2.2 Å−1 showed good mutual agreement, affirming that this region has interpretable features for structural modelling. SAS measurements with inline size-exclusion chromatography (SEC) proved to be generally superior for eliminating sample heterogeneity, but with unavoidable sample dilution during column elution, while batch SAS data collected at higher concentrations and for longer times provided superior statistical precision. Careful merging of data measured using inline SEC and batch modes, or low- and high-concentration data from batch measurements, was successful in eliminating small amounts of aggregate or interparticle interference from the scattering while providing improved statistical precision overall for the benchmarking data set.
KRas is a small GTPase and membrane-bound signaling protein. Newly synthesized KRas is post-translationally modified with a membrane-anchoring prenyl group. KRas chaperones are therapeutic targets in cancer due to their participation in trafficking oncogenic KRas to membranes. SmgGDS splice variants are chaperones for small GTPases with basic residues in their hypervariable domain (HVR), including KRas. SmgGDS-607 escorts pre-prenylated small GTPases, while SmgGDS-558 escorts prenylated small GTPases. We provide a structural description of farnesylated and fully processed KRas (KRas-FMe) in complex with SmgGDS-558 and define biophysical properties of this interaction. Surface plasmon resonance measurements on biomimetic model membranes quantified the thermodynamics of the interaction of SmgGDS with KRas, and small-angle x-ray scattering was used to characterize complexes of SmgGDS-558 and KRas-FMe structurally. Structural models were refined using Monte Carlo and molecular dynamics simulations. Our results indicate that SmgGDS-558 interacts with the HVR and the farnesylated C-terminus of KRas-FMe, but not its G-domain. Therefore, SmgGDS-558 interacts differently with prenylated KRas than prenylated RhoA, whose G-domain was found in close contact with SmgGDS-558 in a recent crystal structure. Using immunoprecipitation assays, we show that SmgGDS-558 binds the GTP-bound, GDP-bound, and nucleotide-free forms of farnesylated and fully processed KRas in cells, consistent with SmgGDS-558 not engaging the G-domain of KRas. We found that the dissociation constant, Kd, for KRas-FMe binding to SmgGDS-558 is comparable with that for the KRas complex with PDEδ, a well-characterized KRas chaperone that also does not interact with the KRas G-domain. These results suggest that KRas interacts in similar ways with the two chaperones SmgGDS-558 and PDEδ. Therapeutic targeting of the SmgGDS-558/KRas complex might prove as useful as targeting the PDEδ/KRas complex in KRas-driven cancers.
Concerns with current mRNA Lipid Nanoparticle (LNP) systems include dose-limiting reactogenicity, adverse events that may be partly due to systemic off target expression of the immunogen, and a very limited understanding of the mechanisms responsible for the frozen storage requirement. We applied a new rational design process to identify a novel multiprotic ionizable lipid, called C24, as the key component of the mRNA LNP delivery system. We show that the resulting C24 LNP has a multistage protonation behavior resulting in greater endosomal protonation and greater translation of an mRNA-encoded luciferase reporter after intramuscular (IM) administration compared to the standard reference MC3 LNP. Off-target expression in liver after IM administration was reduced 6 fold for the C24 LNP compared to MC3. Neutralizing titers in immunogenicity studies delivering a nucleoside-modified mRNA encoding for the diproline stabilized spike protein immunogen were 10 fold higher for the C24 LNP versus MC3, and protection against viral challenge in a SARS-CoV-2 mouse model occurred at a very low 0.25 µg prime/boost dose of the same immunogen in the C24 LNP. Injection site inflammation was notably reduced for C24 compared to MC3. In addition, we found the C24 LNP to be entirely stable in bioactivity and mRNA integrity when stored at 4 ºC for at least 19 days. Storage at higher temperatures reduced both bioactivity and mRNA integrity, but less so for C24 than MC3, and in a manner consistent with the phosphodiester transesterification reaction mechanism of mRNA cleavage. The higher potency, lower injection site inflammation, and higher stability of the C24 LNP present important advancements in the evolution mRNA vaccine delivery.
Lipid Nanoparticles (LNPs) are used to deliver siRNA and COVID-19 mRNA vaccines. The main factor known to determine their delivery efficiency is the pKa of the LNP containing an ionizable lipid. Herein, we report a method that can predict the LNP pKa from the structure of the ionizable lipid. We used theoretical, NMR, fluorescent-dye binding, and electrophoretic mobility methods to comprehensively measure protonation of both the ionizable lipid and the formulated LNP. The pKa of the ionizable lipid was 2-3 units higher than the pKa of the LNP primarily due to proton solvation energy differences between the LNP and aqueous medium. We exploited these results to explain a wide range of delivery efficiencies in vitro and in vivo for intramuscular (IM) and intravascular (IV) administration of different ionizable lipids at escalating ionizable lipid-to-mRNA ratios in the LNP. In addition, we determined that more negatively charged LNPs exhibit higher off-target systemic expression of mRNA in the liver following IM administration. This undesirable systemic off-target expression of mRNA-LNP vaccines could be minimized through appropriate design of the ionizable lipid and LNP.
The respiratory syncytial virus (RSV) fusion (F) protein/polysorbate 80 (PS80) nanoparticle vaccine is the most clinically advanced vaccine for maternal immunization and protection of newborns against RSV infection. It is composed of a near-full-length RSV F glycoprotein, with an intact membrane domain, formulated into a stable nanoparticle with PS80 detergent. To understand the structural basis for the efficacy of the vaccine, a comprehensive study of its structure and hydrodynamic properties in solution was performed. Small-angle neutron scattering experiments indicate that the nanoparticle contains an average of 350 PS80 molecules, which form a cylindrical micellar core structure and five RSV F trimers that are arranged around the long axis of the PS80 core. All-atom models of full-length RSV F trimers were built from crystal structures of the soluble ectodomain and arranged around the long axis of the PS80 core, allowing for the generation of an ensemble of conformations that agree with small-angle neutron and X-ray scattering data as well as transmission electron microscopy (TEM) images. Furthermore, the hydrodynamic size of the RSV F nanoparticle was found to be modulated by the molar ratio of PS80 to protein, suggesting a mechanism for nanoparticle assembly involving addition of RSV F trimers to and growth along the long axis of the PS80 core. This study provides structural details of antigen presentation and conformation in the RSV F nanoparticle vaccine, helping to explain the induction of broad immunity and observed clinical efficacy. Small-angle scattering methods provide a general strategy to visualize surface glycoproteins from other pathogens and to structurally characterize nanoparticle vaccines.
Targeting Clostridium difficile infection is challenging because treatment options are limited, and high recurrence rates are common. One reason for this is that hypervirulent C. difficile strains often have a binary toxin termed the C. difficile toxin, in addition to the enterotoxins TsdA and TsdB. The C. difficile toxin has an enzymatic component, termed CDTa, and a pore-forming or delivery subunit termed CDTb. CDTb was characterized here using a combination of single-particle cryoelectron microscopy, X-ray crystallography, NMR, and other biophysical methods. In the absence of CDTa, 2 di-heptamer structures for activated CDTb (1.0 MDa) were solved at atomic resolution, including a symmetric (SymCDTb; 3.14 Å) and an asymmetric form (AsymCDTb; 2.84 Å). Roles played by 2 receptor-binding domains of activated CDTb were of particular interest since the receptor-binding domain 1 lacks sequence homology to any other known toxin, and the receptor-binding domain 2 is completely absent in other well-studied heptameric toxins (i.e., anthrax). For AsymCDTb, a Ca2+ binding site was discovered in the first receptor-binding domain that is important for its stability, and the second receptor-binding domain was found to be critical for host cell toxicity and the di-heptamer fold for both forms of activated CDTb. Together, these studies represent a starting point for developing structure-based drug-design strategies to target the most severe strains of C. difficile.
Elicitation of broadly neutralizing Ab (bNAb) responses toward the conserved HIV-1 envelope (Env) CD4 binding site (CD4bs) by vaccination is an important goal for vaccine development and yet to be achieved. The outcome of previous immunogenicity studies suggests that the limited accessibility of the CD4bs and the presence of predominant nonneutralizing determinants (nND) on Env may impede the elicitation of bNAbs and their precursors by vaccination. In this study, we designed a panel of novel immunogens that 1) preferentially expose the CD4bs by selective elimination of glycosylation sites flanking the CD4bs, and 2) minimize the nND immune response by engineering fusion proteins consisting of gp120 Core and one or two CD4-induced (CD4i) mAbs for masking nND epitopes, referred to as gp120-CD4i fusion proteins. As expected, the fusion proteins possess improved antigenicity with retained affinity for VRC01-class, CD4bs-directed bNAbs and dampened affinity for nonneutralizing Abs. We immunized C57BL/6 mice with these fusion proteins and found that overall the fusion proteins elicit more focused CD4bs Ab response than prototypical gp120 Core by serological analysis. Consistently, we found that mice immunized with selected gp120-CD4i fusion proteins have higher frequencies of germinal center-activated B cells and CD4bs-directed memory B cells than those inoculated with parental immunogens. We isolated three mAbs from mice immunized with selected gp120-CD4i fusion proteins and found that their footprints on Env are similar to VRC01-class bNAbs. Thus, using gp120-CD4i fusion proteins with selective glycan deletion as immunogens could focus Ab response toward CD4bs epitope.
Determination of structure of RNA via NMR is complicated in large part by the lack of a precise parameterization linking the observed chemical shifts to the underlying geometric parameters. In contrast to proteins, where numerous high-resolution crystal structures serve as coordinate templates for this mapping, such models are rarely available for smaller oligonucleotides accessible via NMR, or they exhibit crystal packing and counter-ion binding artifacts that prevent their use for the chemical shifts analysis. On the other hand, NMR-determined structures of RNA often are not solved at the density of restraints required to precisely define the variable degrees of freedom. In this study we sidestep the problems of direct parameterization of the RNA chemical shifts/structure relationship and examine the effects of imposing local fragmental coordinate similarity restraints based on similarities of the experimental secondary ribose 13C/1H chemical shifts instead. The effect of such chemical shift similarity (CSS) restraints on the structural accuracy is assessed via residual dipolar coupling (RDC)-based cross-validation. Improvements in the coordinate accuracy are observed for all of the six RNA constructs considered here as test cases, which argues for routine inclusion of these terms during NMR-based oligonucleotide structure determination. Such accuracy improvements are expected to facilitate derivation of the chemical shift/structure relationships for RNA.
AbstractTargetingClostridium difficileinfection (CDI) is challenging because treatment options are limited, and high recurrence rates are common. One reason for this is that hypervirulent CDI often has a binary toxin termed theC. difficiletoxin (CDT), in addition to the enterotoxins TsdA and TsdB. CDT has an enzymatic component, termed CDTa, and a pore-forming or delivery subunit termed CDTb. CDTb was characterized here using a combination of single particle cryoEM, X-ray crystallography, NMR, and other biophysical methods. In the absence of CDTa, two novel di-heptamer structures foractivated CDTb (aCDTb; 1.0 MDa) were solved at atomic resolution including a symmetric (SymCDTb; 3.14 Å) and an asymmetric form (AsymCDTb; 2.84 Å). Roles played by two receptor-binding domains of aCDTb were of particular interest since RBD1 lacks sequence homology to any other known toxin, and the RBD2 domain is completely absent in other well-studied heptameric toxins (i.e. anthrax). ForAsymCDTb, a novel Ca2+binding site was discovered in RBD1 that is important for its stability, and RBD2 was found to be critical for host cell toxicity and the novel di-heptamer fold for both forms of aCDTb. Together, these studies represent a starting point for structure-based drug-discovery strategies to targeting CDT in the most severe strains of CDI.SIGNIFICANCE STATEMENTThere is a high burden fromC. difficileinfection (CDI) throughout the world, and the Center for Disease Control (CDC) reports more than 500,000 cases annually in the United States, resulting in an estimated 15,000 deaths. In addition to the large clostridial toxins, TcdA/TcdB, a thirdC. difficilebinary toxin (CDT) is associated with the most serious outbreaks of drug resistant CDI in the 21stcentury. Here, structural biology and biophysical approaches were used to characterize the cell binding component of CDT, termed CDTb, at atomic resolution. Surprisingly, two novel structures were solved from a single sample that help to explain the molecular underpinnings ofC. difficiletoxicity. These structures will also be important for targeting this human pathogen via structure-based therapeutic design methods.