The NMR Exchange Format (NEF) is a community-driven standard for representing NMR experimental data in a consistent, interoperable, and machine-readable form. Built on the STAR syntax, NEF provides a structured framework for storing and exchanging chemical shifts, peak lists, various types of structural restraints, and related metadata, thus allowing for data exchange across software platforms. By enabling direct, lossless transfer of information, NEF simplifies multi-software workflows, improves reproducibility, and supports FAIR (Findable, Accessible, Interoperable, Reusable) data principles. We describe the NEF specification, its current implementation across commonly used NMR software packages, and its application in areas including biomolecular structure determination, metabolomics, and ligand screening. Testing demonstrates that NEF can be used to exchange complete datasets between programs without loss of information or functionality. We also outline recent developments and future directions, such as inclusion of NMR relaxation data and support for non-standard residue topologies. NEFs growing adoption highlights its potential as a unifying standard for NMR data, enabling more efficient, transparent and collaborative research.
Energetically favorable local interactions can overcome the entropic cost of chain ordering and cause otherwise flexible polymers to adopt regularly repeating backbone conformations. A prominent example is the α helix present in many protein structures, which is stabilized by i, i + 4 hydrogen bonds between backbone peptide units. With the increased chemical diversity offered by unnatural amino acids and backbones, it has been possible to identify regularly repeating structures not present in proteins, but to date, there has been no systematic approach for identifying new polymers likely to have such structures despite their considerable potential for molecular engineering. Here we describe a systematic approach to search through dipeptide combinations of 130 chemically diverse amino acids to identify those predicted to populate unique low-energy states. We characterize ten newly identified dipeptide repeating structures using circular dichroism spectroscopy and comparison with calculated spectra. NMR and X-ray crystallographic structures of two of these dipeptide-repeat polymers are similar to the computational models. Our approach is readily generalizable to identify low-energy repeating structures for a wide variety of polymers, and our ordered dipeptide repeats provide new building blocks for molecular engineering.
Biomolecular structure analysis from experimental NMR studies generally relies on restraints derived from a combination of experimental and knowledge-based data. A challenge for the structural biology community has been a lack of standards for representing these restraints, preventing the establishment of uniform methods of model-vs-data structure validation against restraints and limiting interoperability between restraint-based structure modeling programs. The NEF and NMR-STAR formats provide a standardized approach for representing commonly used NMR restraints. Using these restraint formats, a standardized validation system for assessing structural models of biopolymers against restraints has been developed and implemented in the wwPDB OneDep data deposition-validation-biocuration system. The resulting wwPDB restraint violation report provides a model vs. data assessment of biomolecule structures determined using distance and dihedral restraints, with extensions to other restraint types currently being implemented. These tools are useful for assessing NMR models, as well as for assessing biomolecular structure predictions based on distance restraints.
Recent advances in molecular modeling of protein structures are changing the field of structural biology. AlphaFold-2 (AF2), an AI system developed by DeepMind, Inc., utilizes attention-based deep learning to predict models of protein structures with high accuracy relative to structures determined by X-ray crystallography and cryo-electron microscopy (cryoEM). Comparing AF2 models to structures determined using solution NMR data, both high similarities and distinct differences have been observed. Since AF2 was trained on X-ray crystal and cryoEM structures, we assessed how accurately AF2 can model small, monomeric, solution protein NMR structures which (i) were not used in the AF2 training data set, and (ii) did not have homologous structures in the Protein Data Bank at the time of AF2 training. We identified nine open source protein NMR data sets for such “blind” targets, including chemical shift, raw NMR FID data, NOESY peak lists, and (for 1 case) 15 N- 1 H residual dipolar coupling data. For these nine small (70 - 108 residues) monomeric proteins, we generated AF2 prediction models and assessed how well these models fit to these experimental NMR data, using several well-established NMR structure validation tools. In most of these cases, the AF2 models fit the NMR data nearly as well, or sometimes better than, the corresponding NMR structure models previously deposited in the Protein Data Bank. These results provide benchmark NMR data for assessing new NMR data analysis and protein structure prediction methods. They also document the potential for using AF2 as a guiding tool in protein NMR data analysis, and more generally for hypothesis generation in structural biology research. Highlights AF2 models assessed against NMR data for 9 monomeric proteins not used in training. AF2 models fit NMR data almost as well as the experimentally-determined structures. RPF-DP, PSVS , and PDBStat software provide structure quality and RDC assessment. RPF-DP analysis using AF2 models suggests multiple conformational states.
Biomolecules exhibit dynamic behavior that single-state models of their structures cannot fully capture. We review some recent advances for investigating multiple conformations of biomolecules, including experimental methods, molecular dynamics simulations, and machine learning. We also address the challenges associated with representing single- and multiple-state models in data archives, with a particular focus on NMR structures. Establishing standardized representations and annotations will facilitate effective communication and understanding of these complex models to the broader scientific community.
Biomolecules exhibit dynamic behavior that single-state models of their structures cannot fully capture. We review some recent advances for investigating multiple conformations of biomolecules, including experimental methods, molecular dynamics simulations, and machine learning. We also address the challenges associated with representing single- and multiple-state models in data archives, with a particular focus on NMR structures. Establishing standardized representations and annotations will facilitate effective communication and understanding of these complex models to the broader scientific community.
Recent advances in molecular modeling using deep learning have the potential to revolutionize the field of structural biology. In particular, AlphaFold has been observed to provide models of protein structures with accuracies rivaling medium-resolution X-ray crystal structures, and with excellent atomic coordinate matches to experimental protein NMR and cryo-electron microscopy structures. Here we assess the hypothesis that AlphaFold models of small, relatively rigid proteins have accuracies (based on comparison against experimental data) similar to experimental solution NMR structures. We selected six representative small proteins with structures determined by both NMR and X-ray crystallography, and modeled each of them using AlphaFold. Using several structure validation tools integrated under the Protein Structure Validation Software suite (PSVS), we then assessed how well these models fit to experimental NMR data, including NOESY peak lists (RPF-DP scores), comparisons between predicted rigidity and chemical shift data (ANSURR scores), and 15N-1H residual dipolar coupling data (RDC Q factors) analyzed by software tools integrated in the PSVS suite. Remarkably, the fits to NMR data for the protein structure models predicted with AlphaFold are generally similar, or better, than for the corresponding experimental NMR or X-ray crystal structures. Similar conclusions were reached in comparing AlphaFold2 predictions and NMR structures for three targets from the Critical Assessment of Protein Structure Prediction (CASP). These results contradict the widely held misperception that AlphaFold cannot accurately model solution NMR structures. They also document the value of PSVS for model vs. data assessment of protein NMR structures, and the potential for using AlphaFold models for guiding analysis of experimental NMR data and more generally in structural biology.
We use computational design coupled with experimental characterization to systematically investigate the design principles for macrocycle membrane permeability and oral bioavailability. We designed 184 6-12 residue macrocycles with a wide range of predicted structures containing noncanonical backbone modifications and experimentally determined structures of 35; 29 are very close to the computational models. With such control, we show that membrane permeability can be systematically achieved by ensuring all amide (NH) groups are engaged in internal hydrogen bonding interactions. 84 designs over the 6-12 residue size range cross membranes with an apparent permeability greater than 1 3 10(-6) cm/s. Designs with exposed NH groups can be made membrane permeable through the design of an alternative isoenergetic fully hydrogen-bonded state favored in the lipid membrane. The ability to robustly design membrane-permeable and orally bioavailable peptides with high structural accuracy should contribute to the next generation of designed macrocycle therapeutics.
Currently, significant efforts are devoted to designing small molecules able to bind selectively to guanine quadruplexes (G4s). These noncanonical DNA structures are implicated in various important biological processes and have been identified as potential targets for drug development. Previously, a series of triphenylamine (TPA)-based compounds, including macrocyclic polyamines, that displayed high affinity towards G4 DNA were reported. Following this initial work, herein a series of second-generation compounds, in which the central TPA has been functionalised with flexible and adaptive linear polyamines, are presented with the aim of maximising the selectivity towards G4 DNA. The acid-base properties of the new derivatives have been studied by means of potentiometric titrations, UV/Vis and fluorescence emission spectroscopy. The interaction with G4s and duplex DNA has been explored by using FRET melting assays, fluorescence spectroscopy and circular dichroism. Compared with previous TPA derivatives with macrocyclic substituents, the new ligands reported herein retain the G4 affinity, but display two orders of magnitude higher selectivity for G4 versus duplex DNA; this is most likely due to the ability of the linear substituents to embrace the G4 structure.
CASP13 has investigated the impact of sparse NMR data on the accuracy of protein structure prediction. NOESY and 15 N-1 H residual dipolar coupling data, typical of that obtained for 15 N,13 C-enriched, perdeuterated proteins up to about 40 kDa, were simulated for 11 CASP13 targets ranging in size from 80 to 326 residues. For several targets, two prediction groups generated models that are more accurate than those produced using baseline methods. Real NMR data collected for a de novo designed protein were also provided to predictors, including one data set in which only backbone resonance assignments were available. Some NMR-assisted prediction groups also did very well with these data. CASP13 also assessed whether incorporation of sparse NMR data improves the accuracy of protein structure prediction relative to nonassisted regular methods. In most cases, incorporation of sparse, noisy NMR data results in models with higher accuracy. The best NMR-assisted models were also compared with the best regular predictions of any CASP13 group for the same target. For six of 13 targets, the most accurate model provided by any NMR-assisted prediction group was more accurate than the most accurate model provided by any regular prediction group; however, for the remaining seven targets, one or more regular prediction method provided a more accurate model than even the best NMR-assisted model. These results suggest a novel approach for protein structure determination, in which advanced prediction methods are first used to generate structural models, and sparse NMR data is then used to validate and/or refine these models.