The emergence of reporting standards represents a tremendous stride toward reproducible science, but significant challenges remain. Preserving data provenance is difficult, from generation through processing and analysis to publication. The accuracy of annotations diminish the longer the annotation is separated from the defining event. Manual annotations are often inaccurate or inconsistent. Federation of data across data resources remains challenging. NMRhub was conceived to address these challenges, and to facilitate the application of machine learning (ML) to unlock latent knowledge. NMRhub links the NSF Network for Advanced NMR (NAN), the NMRbox platform, and the Biological Magnetic Resonance Data Bank (BMRB) to create an integrated ecosystem that sustains data from the instrument source through public deposition, and facilitates surveys, meta-analyses, and ML. NAN harvests data automatically from connected spectrometers and stores it in a secure archive, parses instrument-provided metadata and transmits data to the user NMRbox account. NMRbox virtualizes and archives the complete software environment used for processing and analysis, provides high performance computation resources including access to Open Science Grid, and tools for assembling BMRB depositions. BMRB, a Core Archive of the World-wide Protein Data Bank (wwPDB), curates, annotates, and provides persistent identifiers for all types of biomolecular NMR data. NMRhub provides single sign on across the constituent platforms, ML-enhanced search of assets across the three platforms, as well as across curated external resources (e.g. PubMed, domain-specific YouTube channels, funding databases), and tools for data transfer among the linked resources. The NMRhub ecosystem spans the complete NMR data lifecycle.
Glycans on glycoproteins play roles that range from quality control in protein folding, to mediation of interactions with other proteins, to stabilization of the protein to which they are attached. Computation can suggest structures that underlie these roles, but confidence is limited by the accuracy of energetic calculations and their applicability to the aqueous environment in which proteins function. Experimental validation of suggested structures is therefore of primary importance. Here we use NMR data, including long-range pseudocontact shifts (PCSs) and residual dipolar couplings (RDCs), to screen structures produced by a version of accelerated molecular dynamics (Pep-GaMD). This version was designed to improve the search for peptide-protein interactions, but here it is successfully applied to glycans attached to a target protein. The target protein, the N-terminal domain of human CEACAM1, is expressed with homogeneous GlcNAc2Man5 glycans at its three N-glycosylation sites. One site (N104) is found to have preferred conformations that exploit hydrophobic interactions between its glycans and protein hydrophobic residues, potentially adding to protein stability and protection from adverse interactions.
Human IgG antibodies can oligomerize upon binding to antigens on the cell surface. The HexaBody® technology enhances oligomer formation via a single-point mutation that strengthens Fc:Fc interactions between neighboring antibodies. These oligomers facilitate receptor clustering for outside-in signaling and improve binding to C1q, the first component of the complement system, leading to enhanced complement activation and complement-dependent cytotoxicity (CDC).To validate and study the mechanism of action induced by the HexaBody format, we developed a VHH (HexBlock VHH clone 2E1) that disrupts antibody Fc:Fc interactions. The VHH was generated by immunizing llamas with an antibody variant capable of hexamer formation in solution, followed by phage display library selection.Here, we characterize the biochemical and biophysical properties of HexBlock VHH clone 2E1– IgG complexes using cell surface target binding assays (FACS), biolayer interferometry (BLI), size-exclusion chromatography (HP-SEC), cryogenic electron microscopy (cryo-EM), and nuclear magnetic resonance spectroscopy (NMR). HexBlock VHH clone 2E1 bound to the CH2 domain of IgG antibodies and efficiently interfered with Fc:Fc interactions. In functional assays, this interference caused reduced receptor clustering and suppressed complement activity mediated by IgG1 or HexaBody molecules.
The Network for Advanced NMR (NAN) is a novel distributed resource that connects Nuclear Magnetic Resonance (NMR) facilities via a scalable cyberinfrastructure supporting NMR data harvesting, interactive data management, and the discovery of instruments, methods, and data to enable emerging data standards in biomedicine, chemistry, and material science. Anchored by the first open-access 1.1 GHz instruments in the USA, NAN integrates NMR facilities around a centralized hub for identity management, resource discovery, and access control. The system includes automated data harvesting through the NAN data transport system (NDTS), metadata-rich data archiving, and interactive web-based tools for data and metadata browsing, editing, and publishing, as well as tools for facility and laboratory data management by facility managers and principal investigators. NAN knowledgebases provide vetted, standardized pulse programs, protocols, parameters, and example datasets, along with processed data. Supported by the US National Science Foundation Midscale Research Infrastructure program, NAN helps to democratize access to NMR resources and fosters open, reproducible science.
Not all proteins are amenable to uniform isotopic labeling with 13C and 15N, something needed for the widely used, and largely deductive, triple resonance assignment process. Among them are proteins expressed in mammalian cell culture where native glycosylation can be maintained, and proper formation of disulfide bonds facilitated. Uniform labeling in mammalian cells is prohibitively expensive, but sparse labeling with one or a few isotopically enriched amino acid types is an option for these proteins. However, assignment then relies on accessing the best match between a variety of measured NMR parameters and predictions based on 3D structure, often from X-ray crystallography. Finding this match is a challenging process that has benefitted from many computational tools, including trained neural nets for chemical shift prediction, genetic algorithms for searches through a myriad of assignment possibilities, and now AI-based prediction of high-quality structures for protein targets. AssignSLP_GUI, a new version of a software package for assignment of resonances from sparsely-labeled proteins, uses many of these tools. These tools and new additions to the package are highlighted in an application to a sparsely-labeled domain from a glycoprotein, CEACAM1.
High resolution hydroxyl radical protein footprinting (HR-HRPF) is a mass spectrometry-based method that measures the solvent exposure of multiple amino acids in a single experiment, offering constraints for experimentally informed computational modeling. HR-HRPF-based modeling has previously been used to accurately model the structure of proteins of known structure, but the technique has never been used to determine the structure of a protein of unknown structure. Here, we present the use of HR-HRPF-based modeling to determine the structure of the Ig-like domain of NRG1, a protein with no close homolog of known structure. Independent determination of the protein structure by both HR-HRPF-based modeling and heteronuclear NMR was carried out, with results compared only after both processes were complete. The HR-HRPF-based model was highly similar to the lowest energy NMR model, with a backbone RMSD of 1.6 Å. To our knowledge, this is the first use of HR-HRPF-based modeling to determine a previously uncharacterized protein structure.
Thrombospondin type-1 repeats (TSRs) are small protein motifs containing six conserved cysteines forming three disulfide bonds that can be modified with an O-linked fucose. Protein O-fucosyltransferase 2 (POFUT2) catalyzes the addi-tion of O-fucose to TSRs containing the appropriate consensus sequence, and the O-fucose modification can be elongated to a Glucose-Fucose disaccharide with the addition of glucose by beta 3-glucosyltransferase (B3GLCT). Elimination of Pofut2 in mice results in embryonic lethality in mice, highlighting the biological significance of O-fucose modification on TSRs. Knockout of POFUT2 in HEK293T cells has been shown to cause complete or partial loss of secretion of many proteins containing O-fucosylated TSRs. In addition, POFUT2 is local-ized to the endoplasmic reticulum (ER) and only modifies folded TSRs, stabilizing their structures. These observations suggest that POFUT2 is involved in an ER quality control mechanism for TSR folding and that B3GLCT also participates in quality control by providing additional stabilization to TSRs. However, the mechanisms by which addition of these sugars result in stabilization are poorly understood. Here, we con-ducted molecular dynamics (MD) simulations and provide crystallographic and NMR evidence that the Glucose-Fucose disaccharide interacts with specific amino acids in the TSR3 domain in thrombospondin-1 that are within proximity to the O-fucosylation modification site resulting in protection of a nearby disulfide bond. We also show that mutation of these amino acids reduces the stabilizing effect of the sugars in vitro. These data provide mechanistic details regarding the impor-tance of O-fucosylation and how it participates in quality control mechanisms inside the ER.
Glycans attached to glycoproteins can contribute to stability, mediate interactions with other proteins, and initiate signal transduction. Glycan conformation, which is critical to these processes, is highly variable and often depicted as sampling a multitude of conformers. These conformers can be generated by molecular dynamics simulations, and more inclusively by accelerated molecular dynamics, as well as other extended sampling methods. However, experimental assessments of the contribution that various conformers make to a native ensemble are rare. Here, we use long-range pseudo-contact shifts (PCSs) of NMR resonances from an isotopically labeled glycoprotein to identify preferred conformations of its glycans. The N-terminal domain from human Carcinoembryonic Antigen Cell Adhesion Molecule 1, hCEACAM1-Ig1, was used as the model glycoprotein in this study. It has been engineered to include a lanthanide-ion-binding loop that generates PCSs, as well as a homogeneous set of three 13C-labeled N-glycans. Analysis of the PCSs indicates that preferred glycan conformers have extensive contacts with the protein surface. Factors leading to this preference appear to include interactions between N-acetyl methyls of GlcNAc residues and hydrophobic surface pockets on the protein surface.
Skp1 is an adapter that links F-box proteins to cullin-1 in the Skp1/cullin-1/F-box (SCF) protein family of E3 ubiquitin ligases that targets specific proteins for polyubiquitination and subsequent protein degradation. Skp1 from the amoebozoan Dictyostelium forms a stable homodimer in vitro with a Kd of 2.5 μM as determined by sedimentation velocity studies yet is monomeric in crystal complexes with F-box proteins. To investigate the molecular basis for the difference, we determined the solution NMR structure of a doubly truncated Skp1 homodimer (Skp1ΔΔ). The solution structure of the Skp1ΔΔ dimer reveals a 2-fold symmetry with an interface that buries ∼750 Å2 of predominantly hydrophobic surface. The dimer interface overlaps with subsite 1 of the F-box interaction area, explaining why only the Skp1 monomer binds F-box proteins (FBPs). To confirm the model, Rosetta was used to predict amino acid substitutions that might disrupt the dimer interface, and the F97E substitution was chosen to potentially minimize interference with F-box interactions. A nearly full-length version of Skp1 with this substitution (Skp1ΔF97E) behaved as a stable monomer at concentrations of ≤500 μM and actively bound a model FBP, mammalian Fbs1, which suggests that the dimeric state is not required for Skp1 to carry out a basic biochemical function. Finally, Skp1ΔF97E is expected to serve as a monomer model for high-resolution NMR studies previously hindered by dimerization.
Thrombospondin Type 1 Repeats (TSRs) are small cysteine‐rich domains found in multiple cell‐surface and secreted proteins. TSRs containing the consensus sequence Cxx(S/T)C are typically modified on the serine or threonine with an O‐linked Glcβ1‐3Fuc disaccharide. The O‐fucose is added by Protein O‐fucosyltransferase 2 (POFUT2), which is then elongated by β1‐3glucosyltransferase (B3GLCT). Elimination of Pofut2 in mice results in early embryonic lethality, and human mutations in B3GLCT cause Peters Plus Syndrome (PPS), a Congenital Disorder of Glycosylation. POFUT2 and B3GLCT are localized in the endoplasmic reticulum (ER) ‐ the folding compartment for proteins in the secretory pathway, and POFUT2 only adds fucose to properly folded TSRs. Thus, POFUT2 appears to function as a folding sensor for TSRs, and both POFUT2 and B3GLCT have been proposed to assist in the folding of TSR‐containing proteins. Interestingly, loss of POFUT2 causes secretion defects for most TSR‐containing proteins, while loss of B3GLCT only effects secretion of a subset of these proteins. The mechanisms behind why POFUT2 is required for section of most TSR‐containing proteins but B3GLCT is required for only a few is not understood. We hypothesize that this unique disaccharide on TSRs is interacting with amino acids in close proximity to the sugars thereby stabilizing the folded state of TSRs and ultimately the protein as a whole, and that the fucose has the major stabilizing effect, while the glucose is only required for some TSRs. To test this hypothesis, we designed a Reductive Unfolding Assay to monitor the effects of the sugars on the stability of TSRs. We have identified amino acids affected by the presence of the fucose or glucose using NMR and molecular dynamics simulations. We have mutated several of these amino acids to determine if they reduce the stabilizing effects of the sugars using the reductive unfolding assay. One such mutant, P547A in TSR3 from human thrombospondin‐1, decreases the sugar‐mediated stabilization of TSR3, providing evidence that the disaccharide stabilizes the folded state by interacting with certain amino acids in close proximity. We are examining the stabilizing effects of the fucose and glucose on several other TSRs as well. These results provide further evidence that POFUT2 and B3GLCT act as quality control enzymes inside the ER.Support or Funding InformationThis work was supported by NIH grant HD090156.
As complications associated with antibiotic resistance have intensified, copper (Cu) is attracting attention as an antimicrobial agent. Recent studies have shown that copper surfaces decrease microbial burden, and host macrophages use Cu to increase bacterial killing. Not surprisingly, microbes have evolved mechanisms to tightly control intracellular Cu pools and protect against Cu toxicity. Here, we identified two genes (copB and copL) encoded within the Staphylococcus aureus arginine-catabolic mobile element (ACME) that we hypothesized function in Cu homeostasis. Supporting this hypothesis, mutational inactivation of copB or copL increased copper sensitivity. We found that copBL are co-transcribed and that their transcription is increased during copper stress and in a strain in which csoR, encoding a Cu-responsive transcriptional repressor, was mutated. Moreover, copB displayed genetic synergy with copA, suggesting that CopB functions in Cu export. We further observed that CopL functions independently of CopB or CopA in Cu toxicity protection and that CopL from the S. aureus clone USA300 is a membrane-bound and surface-exposed lipoprotein that binds up to four Cu+ ions. Solution NMR structures of the homologous Bacillus subtilis CopL, together with phylogenetic analysis and chemical-shift perturbation experiments, identified conserved residues potentially involved in Cu+ coordination. The solution NMR structure also revealed a novel Cu-binding architecture. Of note, a CopL variant with defective Cu+ binding did not protect against Cu toxicity in vivo. Taken together, these findings indicate that the ACME-encoded CopB and CopL proteins are additional factors utilized by the highly successful S. aureus USA300 clone to suppress copper toxicity.
Characterization of proteins using NMR methods begins with assignment of resonances to specific residues. This is usually accomplished using sequential connectivities between nuclear pairs in proteins uniformly labeled with NMR active isotopes. This becomes impractical for larger proteins, and especially for proteins that are best expressed in mammalian cells, including glycoproteins. Here an alternate protocol for the assignment of NMR resonances of sparsely labeled proteins, namely, the ones labeled with a single amino acid type, or a limited subset of types, isotopically enriched with 15N or 13C, is described. The protocol is based on comparison of data collected using extensions of simple two-dimensional NMR experiments (correlated chemical shifts, nuclear Overhauser effects, residual dipolar couplings) to predictions from molecular dynamics trajectories that begin with known protein structures. Optimal pairing of predicted and experimental values is facilitated by a software package that employs a genetic algorithm, ASSIGN_SLP_MD. The approach is applied to the 36-kDa luminal domain of the sialyltransferase, rST6Gal1, in which all phenylalanines are labeled with 15N, and the results are validated by elimination of resonances via single-point mutations of selected phenylalanines to tyrosines. Assignment allows the use of previously published paramagnetic relaxation enhancements to evaluate placement of a substrate analog in the active site of this protein. The protocol will open the way to structural characterization of the many glycosylated and other proteins that are best expressed in mammalian cells.
An enzyme- and click chemistry-mediated methodology for the site-specific nitroxide spin labeling of glycoproteins has been developed and applied. The procedure relies on the presence of single N-glycosylation sites that are present natively in proteins or that can be engineered into glycoproteins by mutational elimination of all but one glycosylation site. Recombinantly expressing glycoproteins in HEK293S (GnT1-) cells results in N-glycans with high-mannose structures that can be processed to leave a single GlcNAc residue. This can in turn be modified by enzymatic addition of a GalNAz residue that is subject to reaction with an alkyne-carrying TEMPO moiety using copper(I)-catalyzed click chemistry. To illustrate the procedure, we have made an application to a two-domain construct of Robo1, a protein that carries a single N-glycosylation site in its N-terminal domains. The construct has also been labeled with 15N at amide nitrogens of lysine residues to provide a set of sites that are used to derive an effective location of the paramagnetic nitroxide moiety of the TEMPO group. This, in turn, allowed measurements of paramagnetic perturbations to the spectra of a new high affinity heparan sulfate ligand. Calculation of distance constraints from these data facilitated determination of an atomic level model for the docked complex.