Nuclear magnetic resonance (NMR) spectroscopy is an indispensable analytical technique in chemistry, biology, and medicine, offering molecular-level structural and dynamic insights for metabolite identification and biomarker discovery in biological and clinical research. However, its broader application in metabolite research is constrained by inherent limitations such as low sensitivity, severe peak overlap, and the trade-off between portability of low-field instruments and spectral resolution and sensitivity, as well as the difficulty of mining valid biological information from high-throughput metabolite big data. Recent advancements in artificial intelligence (AI), particularly deep learning (DL), have contributed to addressing these core challenges, reshaping the capabilities of NMR technology in metabolite research. This perspective overviews NMR's core role and existing metabolite-specific bottlenecks, and traces the evolutionary progress of AI in NMR. Then it focuses on three interrelated and technologically innovative directions of AI-powered NMR for metabolite, including (1) improving the spectral resolution of low-field spectra to match that of high-field spectra, or directly extracting high-field-like NMR information from low-field spectra; (2) mitigating spectral congestion via optimizing pure shift spectra of complex metabolic mixtures, which simplifies complex spectra by converting multiplets into singlets; and (3) identifying metabolites and biomarkers (including lipoprotein-related metabolic markers in hyperlipidemia) in metabolomics. The synergy of AI and NMR is poised to unlock unprecedented analytical capabilities for metabolomics, extending its impact across scientific discovery, precision medicine, food science, and other related fields.
Oxidative modification of cytochrome c (Cyt c) may influence the local conformation of protein, yet the mechanism by which structural alterations of Cyt c affect its degree of oxidative modification remains unclear. In this study, Girard’s reagent T (GRT) was employed as a nuclear magnetic resonance (NMR) probe to investigate the oxidative modification levels of human Cyt c under varying environmental conditions. Experimental results demonstrated that protecting lysine residues through reductive methylation effectively reduced protein oxidation. Partial unfolding of Cyt c was found to enhance its oxidative modification, while binding Cyt c with cardiolipin significantly increased the extent of oxidation. Additionally, other factors such as protein aggregation exhibited inhibitory effects on oxidative modification.
SMILES-based molecular language models are scalable but lack explicit threedimensional structural modeling, while existing 3D molecular pretraining methods usually do not use chemically informed SMILES tokens as an interface to atom-level geometry. Here, we propose AIS-RoChem3D, an Atom-in-SMILES-guided 3D denoising pretraining framework that aligns AIS token sequences with atom-level edge types, interatomic distances, and coordinates through an explicit token-to-atom mapping. AIS-RoChem3D uses a pair-bias Transformer encoder with ProbMix attention to jointly model sequence semantics, AIS-aware topology, and conformational geometry, and is pretrained with masked language modeling, coordinate recovery, and pairwise distance recovery. On MoleculeNet benchmarks, AISRoChem3D achieved competitive performance across classification and regression tasks, with strong improvements on ESOL and FreeSolv and task-dependent gains over the sequenceonlyAIS-RoChem reference model. On experimental ^13C NMR chemical shift prediction, AIS-RoChem3D also achieved strong atom-level regression performance using a simple feedforward readout, supporting the effectiveness of token-to-atom alignment for learning transferable atomic representations. These results show that AIS-guided 3D denoising pretraining provides a practical bridge between SMILES language modeling and atom-level structural representation learning.
Low-populated protein excited states can play decisive roles in function, ligand recognition, and misfolding, but their structural characterization remains difficult because these conformations are transient and sparsely populated. Here we present CS-SelectFold, a training-free framework that combines generative conformational sampling with NMR chemical-shift-guided post hoc selection to recover low-populated protein conformations. CS-SelectFold uses SimpleFold to sample diverse candidate structures from sequence, applies structural pre-screening to enrich models that escape the dominant ground-state basin, predicts chemical shifts with UCBShift 2.0, and ranks candidates by agreement with experimentally determined excited-state chemical shifts.We benchmarked CS-SelectFold on two proteins with experimentally characterized low-populated states. For the T4 lysozyme L99A cavity mutant, a canonical excited-state system, the method recovered the defining excited-state-like local rearrangement around the engineered cavity, including inward placement of Phe114, consistent with the NMR/CS-Rosetta-derived excited-state model. A two-reference hotspot dRMSD analysis further showed that chemical-shift-based ranking enriched conformations that moved away from the ground-state basin and toward the experimentally defined excited-state basin. For the A39V/N53P/V55L Fyn SH3 domain, which populates an aggregation-prone hidden folding intermediate, CS-SelectFold recovered the key topological signature of the experimentally characterized intermediate: loss of the C-terminal β-strand present in the native state, as confirmed by both three-dimensional structural comparison and secondary-structure analysis.These results establish CS-SelectFold as a practical framework for excited-state structure recovery and suggest that the same sample-and-select principle may extend to broader low-populated hidden conformations, including folding intermediates.
Cytochrome c is a multifunctional protein involved in electron transport and apoptosis. However, its conformational landscape is complex and heterogeneous, which has obscured the functional understanding-particularly regarding its long-observed dimerized state. In this study, we characterized the conformational distributions of yeast iso-1 cytochrome c (ycyt c) by exploiting the native trimethylated K72 residue (K72me3) as an NMR molecular probe. Our analysis revealed that C102-mediated dimerization disrupts the M80-heme iron coordination, shifting the equilibrium toward conformations featuring an exposed heme and enhanced peroxidase activity. Furthermore, molecular dynamics simulations suggest that the oxidized monomer samples an "open" conformation, which increases the active site accessibility, likely contributing to its elevated peroxidase activity. These findings shed light on the biological significance of C102 in ycyt c and propose a novel mechanism by which dimerization activates the protein as a peroxidase, potentially protecting cells against apoptosis or other oxidative damages.
Protein excited-state conformations are often low-populated and transient, yet can play critical roles in ligand binding, allostery, and catalysis. While NMR relaxation-dispersion and CEST experiments can provide chemical shifts for these otherwise invisible states, existing structure-determination strategies typically incorporate such data directly into restrained reconstruction workflows. Here we introduce CS-SelectFold, a training-free framework for discovering protein excited-state structures by chemical-shift-guided conformational sampling with SimpleFold. CS-SelectFold first generates a diverse ensemble of candidate conformations with SimpleFold, then predicts chemical shifts for each structure using UCBShift 2.0, and finally ranks conformations by agreement with experimentally determined excited-state chemical shifts. Using the L99A mutant of T4 lysozyme as a canonical benchmark, CS-SelectFold recovered excited-state-like conformations that reproduced the defining cavity-centered rearrangement of the known hidden state, including inward placement of Phe114 into the engineered cavity. Hotspot-region dRMSD analysis further showed that chemical-shift-guided post hoc selection enriched structures that escaped the dominant ground-state basin and converged toward the experimentally established excited-state structural basin. These results establish CS-SelectFold as a conceptually distinct sample-and-select strategy for excited-state structure discovery and suggest that experimentally determined chemical shifts can serve as an effective post hoc signal for identifying rare, functionally relevant conformations within generative protein structure ensembles.
The coating level of core-shell metal-organic framework (MOF) composites plays a pivotal role in the application fields of gas adsorption, chemical separation, and molecular catalysis. Quantitative characterization of the shell coverage in MOF@MOF systems is essential for elucidating structural features and establishing structure-property relationships across diverse applications. In this study, a combination of solid-state nuclear magnetic resonance (NMR), transmission electron microscopy (TEM), scanning electron microscopy (SEM), X-ray diffraction (XRD), and water contact angle measurements were employed to systematically investigate the relationship between shell coating levels and surface hydrophobicity in two representative hydrophobic MOF@MOF composites, MIL-53@ZIF-8 and ZIF-90@ZIF-71. Solid-state NMR, in conjunction with XRD, enabled structural identification and quantitative analysis of the shell-to-core ratios in these composites with varying coating levels. The formation of core-shell architecture was directly visualized by TEM and confirmed by EDX elemental mapping, which also allowed for a rough estimation of shell thickness at different shell-to-core ratios. Notably, solid-state NMR and contact angle tests revealed a gradual enhancement in the surface hydrophobicity with increasing shell coverage in both MIL-53@ZIF-8 and ZIF-90@ZIF-71. The critical shell-to-core ratios required to achieve complete shell coverage with surface hydrophobicity were determined to be 2.5 for MIL-53@ZIF-8 and 2.9 for ZIF-90@ZIF-71, respectively. This work demonstrates that spectroscopic techniques enable quantitative evaluation of shell coverage and direct assessment of surface hydrophobicity in MOF@MOF systems, providing valuable insights into the rational design of hydrophobic MOF-based core-shell materials for moisture-sensitive applications.
B-factor is a measure of ray attenuation or scattering caused by atomic thermal motion during X-ray diffraction of protein crystal structure. B-factor reflects the vibration of atoms; hence, it is the most common experimental descriptor of protein flexibility and has been extensively applied in the studies of protein dynamics, screening of bioactive small molecules, and protein engineering. The prediction of B-factor profiles has considerable significance for analyzing the dynamic properties of unknown proteins. Deep learning technology has developed rapidly in recent years and has been widely implemented in many research fields, especially structural biology. In this paper, a deep neural network model based on bidirectional long short-term memory (biLSTM) network is proposed to predict the B-factor profile of a protein by combining its sequence-based features and structure-based features. Based on a large dataset of high-resolution proteins, our method predicts the B-factor profiles with an average Pearson correlation coefficient (PCC) of 0.71, and 85% of the B-factor profiles have a PCC greater than 0.6, which indicates a strong correlation between predicted and experimental values. In addition, our method remarkably outperforms the existing methods on four test datasets with different protein sizes.
The description and understanding of protein structure rely on secondary structure heavily. Secondary structure determination and prediction are widely used in protein structure related research. The secondary structure prediction methods based on NMR chemical shifts are convenient to use, so they are popular in protein NMR research. In recent years, there is significant improvement in deep neural network, which is consequently applied in many search fields. Here we proposed a deep neural network based on bidirectional long short term memory (biLSTM) to predict protein 3-state secondary structure using NMR chemical shifts of backbone nuclei. Compared with the existing methods of the same sort, the accuracy of the proposed method was improved. And a web server was built to provide secondary structure prediction service using this method.
Understanding and predicting individual responses to immune checkpoint inhibitors (ICIs) are of great importance for advancing precise cancer immunotherapy. This study utilizes 1H NMR-based plasma phenotyping to reveal the metabolic characteristics and predictive signatures of tumor-bearing mice that are sensitive and nonsensitive to the PD-1 inhibitor RMP1-14. Metabolite set enrichment analysis showed that the high-sensitivity (HS) group may exhibit a more robust immune profile by upregulating sphingolipid metabolism and lactose metabolism before treatment, while the enrichment of branched-chain amino acids and disorder of glucose and lipid metabolism in the low-sensitivity (LS) group may facilitate tumor growth and immune escape. 1H NMR information on glycoproteins and lipoproteins in predose plasma was significantly correlated with differences in treatment effects. This correlation was validated by the random forest prediction model constructed from the NMR information, which demonstrated a lower out-of-bag error rate. The model further suggested that acetyl-signaling GlycA and GlycB from N-acetyl glycoproteins of predose plasma are predictive signatures for the HS group, whereas lipid-related signals were predictive for the LS group. This study presents an NMR-based framework for investigating and predicting individual variability in responses to ICIs, providing a methodological and data-driven foundation for enhancing precision cancer immunotherapy.
Separating alcohol from water using highly efficient adsorbents with exceptional flux and separation factors remains a significant challenge, garnering widespread interest in the field of biofuel production. In this study, we developed several nanoscale hydrophobic zeolitic imidazolate framework (ZIF) composites, including core-shell structures like ZIF-90@ZIF-8 and ZIF-90@ZIF-71, as well as polydimethylsiloxane (PDMS)-coated ZIFs (ZIF-8, ZIF-71, and ZIF-90), to separate tert-butanol from water. We assessed the external hydrophobicity of these materials by using water contact angle tests. Comprehensive characterization of the core-shell ZIF composites was conducted by using X-ray diffraction (XRD), transmission electron microscopy (TEM), and solid-state nuclear magnetic resonance (NMR). ZIF-8 and ZIF-71 demonstrated hydrophobic properties, whereas ZIF-90 was found to be hydrophilic. Notably, ZIF-90 exhibited a significantly higher uptake capacity for tert-butanol compared with ZIF-71 and ZIF-8. Two-dimensional (2D) 1H-1H spin diffusion homonuclear correlation NMR was employed to distinguish in-pore and ex-pore water and tert-butanol in various core-shell ZIF composites. Additionally, 1H magic angle spinning (MAS) NMR experiments were conducted to determine the separation factor of tert-butanol/water mixtures and the uptake capacity of tert-butanol. Our solid-state NMR measurements revealed that the core-shell ZIF-90@ZIF-71 composite exhibited outstanding separation selectivity exceeding 500 for tert-butanol/water mixtures and a superior uptake capacity of 1.14 mmol/g for tert-butanol among the tested ZIF composites. The combination of a hydrophilic core and a hydrophobic shell in ZIF-90@ZIF-71 enhances the separation efficiency of tert-butanol from water.
MOFs (metal-organic frameworks) have large specific surface area, abundant active sites, and tunable pore structures, making them highly attractive anode materials for lithium-ion batteries. However, their inadequate electrical conductivity and rate performance greatly hinder the further applications of MOFs in field of energy storage. In this study, we present a robust Mn-MOF/HA (humic acid) composite anode material with multilayer nano-sheet structures, synthesized via in-situ compositing HA with Mn-MOF in general hydrothermal reaction. The Mn-MOF derived from hydrothermal reaction shows poor electrochemical stability and low electrical conductivity. The incorporation of HA notably elevates electrochemical stability of materials, boosts both electrical and lithium-ion conductivity, and also modulates the microstructure to form ultrathin multilayer nanosheets by rationally adjusting the feed ratio of HA. The resulting HA20-Mn-MOF (with a HA feed ratio of 20 %) demonstrates significant improvements in cycling stability and rate performance, with a reversible specific capacity of 1318.7mAh/g at 0.1 A/g after 100 cycles and a substantial capacity of 657.0mAh/g even after 1000 cycles at 1 A/g. An extraordinary V-shaped capacity reversal is observed for HA20-Mn-MOF during cycling. Ex-situ EPR (Electron Paramagnetic Resonance) investigation reveals this capacity growth is associated with the reoxidation of active manganese during cycling, in which a notable correlation between the specific capacity and the Mn2+ signal intensity in EPR spectra is found. These results suggest in-situ compositing HA with MOFs can be a costeffective strategy in regulating MOFs morphology and in achieving an optimized electrochemical performance for MOFs-based electrode materials in energy storage field.
Metabolite analysis is essential for understanding the biochemical processes and pathways that sustain life, providing insights into the complex interactions within cellular systems and clinical examinations. This review explores recent applications of nuclear magnetic resonance (NMR) spectroscopy in metabolite studies. Various methods enhancing analytical accuracy for metabolome profiling and metabolic pathway studies, including spectral simplification techniques, quantitative NMR, high-resolution MAS NMR, and isotopic labeling, are discussed. The application of NMR in in situ and in vivo studies is also covered, highlighting in-cell NMR and in vivo MRS techniques. Last but not least, we discuss recent advancements in NMR hyperpolarization, with a focus on dynamic nuclear polarization (DNP), chemically induced dynamic nuclear polarization (CIDNP), para-hydrogen-induced polarization (PHIP), and signal amplification by reversible exchange (SABRE). These advancements offer significant potential for enhancing the sensitivity and accuracy of metabolite studies and are expected to further deepen the study and understanding of metabolites and metabolic pathways.
Glycolysis is a fundamental process for the generation of cellular energy, and its dysregulation has been linked to a number of diseases, including obesity, neurological disorders, and cancer. Targeting glycolysis is therefore a promising therapeutic strategy. However, effective drugs that specifically target glycolysis are still lacking. In this study, we introduce a novel approach utilizing 19F NMR to monitor early glycolysis in living MCF-7 cells. By tracking metabolites downstream of the glucose analogue 2-fluoro-2-deoxyglucose (2-FDG), we successfully observed the activity of key glycolytic components, such as glucose transporters (GLUTs) and hexokinase (HKs). Our results reveal distinct metabolic profiles upon inhibition of these targets, advancing the understanding of glycolytic regulation. In addition, we applied this approach to screen traditional Chinese medicines for their effects on glycolysis and identified Salvia miltiorrhiza and Fructus evodiae as modulators with contrasting effects on glycolytic metabolism. This dual modulation highlights their potential as valuable tools for therapeutic intervention. Our study provides an innovative methodology for both the exploration of glycolytic pathways and the discovery of novel therapeutics, offering new perspectives for drug development targeting metabolic diseases.
Protein dynamics are pivotal to biological function, and elucidating these dynamic properties is essential for understanding their behavior in cellular processes. Nuclear magnetic resonance (NMR) spectroscopy quantifies residue-specific motional freedoms through order parameters (S²), thereby providing critical insights into local structural flexibility and conformational dynamics. Nevertheless, accurate prediction of NMR order parameters remains a critical challenge in structural biology. Traditional approaches often rely on experimental structures or suffer from limited accuracy when using sequence-based methods. To address this challenge, we introduce SOPPCL, a novel deep learning framework that significantly improves sequence-based prediction of order parameters by integrating advanced protein sequence representations and contrastive learning. Our method leverages the ESM-2 protein language model to capture high-dimensional semantic features from amino acid sequences and employs HHblits-derived HMM profiles to encode evolutionary information. To further enhance feature discriminability, SOPPCL introduces a contrastive learning module that aligns and optimizes the fused representations of ESM-2 and HMM features by maximizing mutual information between positive pairs while minimizing similarity among negative pairs. Subsequently, the refined features are fed into a regression network to predict order parameter values. Evaluation on a benchmark dataset containing 10 proteins shows that SOPPCL achieves a Pearson correlation coefficient of 0.845 and a root mean square error of 0.132, surpassing sequence-only baselines such as DynaMine (PCC=0.464, RMSE=0.178) by 82% and 26% relative improvements, respectively. By eliminating reliance on structural data, SOPPCL provides a powerful computational approach for investigating protein dynamics directly from sequence information, with potential applications in NMR-assisted drug design and functional annotation. Altogether, our study highlights the synergistic benefits of integrating protein language models, evolutionary information, and contrastive learning for advancing the prediction of biomolecular properties.
Cytochrome c (cyt c) is released from mitochondria into the cytosol upon apoptotic stimulation, ultimately triggering programmed cell death. Recent studies have revealed that transfer RNA (tRNA) interacts with cyt c, impeding the formation of the apoptosome complex and thereby suppressing apoptosis. To elucidate the molecular mechanism underlying the interaction between cyt c and tRNA, nuclear magnetic resonance (NMR)-based chemical shift perturbation and intensity analysis were employed to characterize the binding interface between cyt c and tRNAphe. The findings demonstrate that cyt c primarily engages with tRNAphe through its 70-85 Ω-loop and N-terminal α-helix. This interaction sterically hinders the accessibility of small molecules, such as H2O2, to the hydrophobic pocket of cyt c, consequently attenuating its peroxidase activity. Furthermore, oxidative modification of cyt c, particularly the carbonylation of positively charged lysine residues, weakens this interaction.
High-resolution nuclear magnetic resonance (NMR) spectroscopy is a powerful analytical tool with wide application. However, the conventional shim technique may not guarantee the homogeneity of the magnetic field when the experimental conditions are unfavorable. In this study, we proposed a data post-processing method called Restore High-resolution Unet (RH-Unet), which uses a convolutional neural network to restore distorted NMR spectra that have been acquired in inhomogeneous magnetic fields. The method generates feature-label pairs from singlet peak regions and ideal Lorentzian line shape and trains a RH-Unet model to map low resolution spectra to high resolution spectra. The method was applied to different samples, and showed superior performance than the REFDCON method incorporated in Bruker Topspin software. The proposed method provides a simple and fast way to obtain high resolution NMR spectra in inhomogeneous fields, which can facilitate the application of NMR spectroscopy in various fields.
As a crucial parameter obtained through NMR spectroscopy for elucidating the dynamic behavior of proteins in solution, backbone N-H order parameters indicate the flexibility of atoms or groups on the picosecond to nanosecond timescale, thereby facilitating a comprehensive understanding of protein function and mechanisms. In NMR experiments, obtaining order parameters involves measuring relaxation parameters and steady-state NOE, followed by conducting Lipari-Szabo model-free analysis, which is a complex and challenging process. In this paper, we employ the self-training strategy, a semi-supervised learning (SSL) approach, to improve the capability of machine learning algorithms in predicting backbone N-H order parameters through iterative training. Throughout the iterative process, we gradually enrich the training set by selecting unlabeled data with high-quality pseudo labels, allowing the student model to progressively enhance its generalization performance under the guidance of the teacher model. The final student model is evaluated on a test set comprising 10 proteins, achieving an average Pearson correlation coefficient of 0.854 between the predicted and experimental order parameters, representing a 4.5% improvement over the original teacher model. It demonstrates that self-training effectively leverages unlabeled structural ensemble data to improve protein fast-dynamics prediction.
Herbal extracts are rich sources of active compounds that can be used for drug screening due to their diverse and unique chemical structures. However, traditional methods for screening these compounds are notably laborious and time-consuming. In this manuscript, we introduce a new high-throughput approach that combines nuclear magnetic resonance (NMR) spectroscopy with a tailored database and algorithm to rapidly identify bioactive components in herbal extracts. This method distinguishes characteristic signals and structural motifs of active constituents in the raw extracts through a relaxation-weighted technique, particularly utilizing the perfect echo Carr-Purcell-Meiboom-Gill (peCPMG) sequence, complemented by precise 2D spectroscopic strategies. The cornerstone of our approach is a customized database designed to filter potential compounds based on defined parameters, such as the presence of CH n segments and unique chemical shifts, thereby expediting the identification of promising compounds. This innovative technique was applied to identifying substances interacting with choline kinase alpha (ChoK alpha 1), resulting in the discovery of four new inhibitors. Our findings demonstrate a powerful tool for unraveling the complex chemical landscape of herbal extracts, considerably facilitating the search for new pharmaceutical candidates. This approach offers an efficient alternative to traditional methods in the quest for drug discovery from natural sources.
Background: In-cell NMR is a valuable technique for investigating protein structure and function in cellular environments. However, challenges arise due to highly crowded cellular environment, where nonspecific interactions between the target protein and other cellular components can lead to signals broadening or disappearance in NMR spectra. Results: We implemented chemical reduction methylation to selectively modify lysine residues on protein surfaces aiming to weaken charge interactions and recover obscured NMR signals. This method was tested on six proteins varying in molecular size and lysine content. While methylation did not disrupt the protein's native conformation, it successful restored some previously obscured in-cell NMR signals, particularly for proteins with high isoelectric points that decreased post-methylation. Significance: This study affirms lysine methylation as a feasible approach to enhance the sensitivity of in-cell NMR spectra for protein studies. By mitigating signal loss due to nonspecific interactions, this method expands the utility of in-cell NMR for investigating proteins in their natural cellular environment, potentially leading to more accurate structural and functional insights.