Bridged azobenzene derivatives are key photo-responsive molecular switches. However, probing and interpreting their microscopic Z ↔ E isomerization mechanisms remain challenging as isolated spectroscopic and computational efforts struggle to establish clear structure-spectrum relationships. We report an integrated, large-language-model (LLM) agent-driven workflow that links literature-guided planning, ab initio molecular dynamics (AIMD) sampling, density functional theory spectral calculations, robotic infrared/Raman measurements, and interpretable machine learning for structural-spectral analysis of bridged azobenzenes. Central to the analysis is an attention-based convolutional neural network (ATT-CNN) that predicts the C-N[double bond, length as m-dash]N-C dihedral angle directly from vibrational spectra with r = 0.99 and MAE = 5°. Attention maps highlight mechanistically informative bands and support holistic (non-marker-dependent) interpretation; transfer learning extends performance across chemical environments and experimental datasets. LLM agents formulated the research plan and coordinated automated simulations and measurements, whereas neural-network architecture design, training, and comparative benchmarking were performed by human researchers to retain full flexibility for model exploration and ensure rigorous interpretation. To our knowledge, this is the first LLM-agent-planned and -orchestrated mechanistic study unifying literature synthesis, theory, experiment, and machine learning. The resulting strategy advances quantitative insight into azobenzene photoisomerization and provides a generalizable blueprint for AI-driven investigations of dynamic molecular systems.
The vast compositional space of high-entropy materials offers exceptional opportunities for the development of powerful catalysts. However, the inverse design of these materials remains unfeasible due to the lack of robust theoretical frameworks and high-throughput experimental tools. This study demonstrates a practical inverse design approach that integrates spectroscopic descriptors, generative machine learning, and a robotic experimental platform to synthesize and optimize catalysts for the oxygen evolution reaction. The automated system substantially accelerated catalyst design and experimental validation, reducing the time required for synthesis, characterization and performance testing from approximately 20 h to only 78 min per sample. Following a rapid screen for efficient senary high-entropy catalysts, the spectroscopic generative model further optimized the top-performing candidate, lowering its overpotential at 10 mA cm-2 by an additional 32.0 mV. Our findings demonstrate the potential of an inverse design approach that incorporates spectroscopic descriptors into generative machine learning to accelerate catalyst discovery. Moreover, this approach is expected to drive the intelligent design of high-performance complex materials.
Exciton coherence is a phenomenon involving collective electronic transition dipole moments of molecular aggregates with interesting photophysical behaviors such as superradiance and subradiance in J- and H-type aggregates, respectively. Although singlet aggregate excitons have been extensively studied, the understanding of triplet exciton coherence in terms of specific conditions is still lacking due to limited experimental observations. Here, by synthesizing model organic compounds of fluorene monomer, dimer, trimer, and polymer, we systematically studied their photoluminescence in both solution and aggregation states. Strong triplet exciton coherence is present in polyfluorene aggregates, including nanoparticles, microparticles, and thin films, and is manifested either as HJ-aggregate phosphorescence or delayed fluorescence depending on the size of these aggregates. Notably, in the phosphorescence state, the polyfluorene polymer exhibits an unusually sharp atomic-spectrum-like emission band with a narrow full width at half maximum (FWHM) of only 0.05 eV (14 nm) and the longest lifetime of 0.63 s. This study shows that the size of organic aggregates is essential in dictating organic exciton dynamics in the solid state and holds importance for the development of advanced optoelectronic technologies.
The recognition and differentiation of organic amines are crucial for applications in drug analysis, food spoilage, biomedical assays, and clinical diagnostics. Existing luminescence-based recognition methods for amines predominantly rely on fluorescence quenching, limiting the scope of sensitive and selective detection. Here, we present a fluorochromic approach for rapidly distinguishing different organic amines based on their unique excited-state and ground-state interactions with a naphthalimide derivative under ultraviolet light. Our findings reveal that the photoluminescence quantum yield and emission color are significantly influenced by the substituent group and the molecular flexibility of the amine. Specifically, primary amines, together with other common lone-pair donors, such as alcohol, ether, thiol, thioether, and phosphine, did not exhibit photoluminescence changes, while secondary amines exhibited only weak emission. For tertiary amines, however, bright green photoluminescence activation was rapidly produced for molecules containing at least one methyl group; red-shifted yellow emission was observed for ones with bulkier side groups other than methyl; and for conformationally locked bicycloamines, no emission was observed. In addition, this fluorochromic process of the naphthalimide derivative not only depends on tertiary amine substituent groups but also shows distinctly different ground- and excited-state photoluminescence dynamics in time-resolved spectroscopy. Based on these differences, a qualitative method is developed for visual recognition of natural and synthetic opioids, including heroin, fentanyl, and metonitazene, which is more facile and rapid compared to current methods such as the Marquis reagent kit, and could facilitate onsite testing, real-time monitoring, and streamlined workflows in both laboratory and field settings.
The investigation of material properties based on atomic structure is a commonly used approach. However, in the study of complex systems such as high-entropy alloys, atomic structure not only covers an excessively vast chemical space, but also has an imprecise correspondence to chemical properties. Herein, we present a label-free machine learning (ML) model based on physics-based spectroscopic descriptors to study the catalytic properties of AgAuCuPdPt high-entropy alloy catalysts. Even if the atomic structures of two such alloys are different, these alloys may have similar catalytic properties if their spectral characteristics match closely. One cluster with the strongest CO adsorption exhibited high selectivity for C2+ product generation, indicating that the spectra-based ML model can provide deeper chemical insight than one based on atomic structure. Moreover, such a model can be extended to other systems with consistent results, thus demonstrating its transferability and versatility. This not only underscores the potential of spectral analysis in identifying high-performance alloy catalysts, but facilitates the formation of a new spectra-based modeling approach and research theory in materials science.
Theoretical analyses of small-molecule adsorption on heterogeneous catalyst surfaces often rely on simplified models of molecular adsorption with the most favorable configuration. Given that real-world experimental tests frequently entail multiple molecules interacting with the surface, there is a pressing need for a comprehensive multimolecule adsorption model to bridge the gap between theory and experiment. Using machine learning, we predict the average values of important adsorption properties from conformationally averaged, calculated infrared and Raman spectra and compare these values to those theoretically derived from the conformationally averaged ensemble. Remarkably, our approach yields excellent predictions even when faced with large and indeterminate numbers of surface molecules. These quantitative spectra-averaged property relationships provide a theoretical framework for extracting key interaction properties from the spectra of real chemical environments.
Geometric information of molecules is closely related to their properties, and vibrational spectroscopy, as a common and powerful analytical tool for determining molecular structure, can assist in gaining precise geometric information. Traditional methods used to delineate spectrum-structure correlations are often expensive, time-consuming, and require extensive professional expertise. In this work, we used a machine learning protocol to construct a map from spectra to molecular geometric structures, and employed Grad-CAM, a convolutional network interpretation technology, to analyze which kinds of chemical information are important for determining our model’s results. The results obtained for six small molecules of differing structures demonstrate that the model is capable of (1) extracting the crucial spectral features that are vital to downstream tasks without necessitating any manual preprocessing, and (2) enabling retrieval of molecular structural information with high precision.
Machine learning (ML) is causing profound changes to chemical research through its powerful statistical and mathematical methodological capabilities. However, the nature of chemistry experiments often sets very high hurdles to collect high-quality data that are deficiency free, contradicting the need of ML to learn from big data. Even worse, the black-box nature of most ML methods requires more abundant data to ensure good transferability. Herein, we combine physics-based spectral descriptors with a symbolic regression method to establish interpretable spectra-property relationship. Using the machine-learned mathematical formulas, we have predicted the adsorption energy and charge transfer of the CO-adsorbed Cu-based MOF systems from their infrared and Raman spectra. The explicit prediction models are robust, allowing them to be transferrable to small and low-quality dataset containing partial errors. Surprisingly, they can be used to identify and clean error data, which are common data scenarios in real experiments. Such robust learning protocol will significantly enhance the applicability of machine-learned spectroscopy for chemical science.
Traditional trial-and-error experiments and theoretical simulations have difficulty optimizing catalytic processes and developing new, better-performing catalysts. Machine learning (ML) provides a promising approach for accelerating catalysis research due to its powerful learning and predictive abilities. The selection of appropriate input features (descriptors) plays a decisive role in improving the predictive accuracy of ML models and uncovering the key factors that influence catalytic activity and selectivity. This review introduces tactics for the utilization and extraction of catalytic descriptors in ML-assisted experimental and theoretical research. In addition to the effectiveness and advantages of various descriptors, their limitations are also discussed. Highlighted are both 1) newly developed spectral descriptors for catalytic performance prediction and 2) a novel research paradigm combining computational and experimental ML models through suitable intermediate descriptors. Current challenges and future perspectives on the application of descriptors and ML techniques to catalysis are also presented.
Background: DNA repair capacity (DRC) is the cell's ability to repair DNA damage by oxidants, carcinogens, radiation and other DNA damaging agents.DRC may be decreased inpatients with lung cancer, and DRC affects cancer cell response to several therapies.Methods: Peripheral blood mononuclear cells (PBMCs) from healthy volunteers and patients with and without lung cancer were cryopreserved using 40% FBS and 10% DMSO until transfection.We developed a host cell reactivation (HCR) assay using PBMCs transfected with modified and unmodified plasmids (pMax-GFP) and a transfection control (pCMV-E2-Crimson) to measure patient-specific DRC of the nucleotide excision repair (NER) and non-homologous end joining (NHEJ) pathways.GFP is produced only if the cell is able to repair the plasmid, measured by flow cytometry and DRC calculated by calculated by the ratio of green to red fluorescence.Results: Titratable measurement of NER and NHEJ activity were confirmed by mixing assays and treatment with a DNA-PK inhibitor, NU7441.DRC measurements in PBMCs from humans without (control) and diagnosed with lung cancer are reproducibility with low inter-individual and inter-assay variability in cancer-free patients (controls).Average NHEJ repair in healthy volunteers was 20.8+/-2.16 at 24 hours.Average NER repair in same volunteers (using a UV-modified plasmid) was 34.3+/-3.4at 24 hours.We are collecting PBMCs from patients with and without lung cancer as well as healthy volunteers.We will measure DRC in these samples and compare the relationship between DRC in patients with and without lung cancer, and comparing other risk factors including gender, race and tobacco smoking history.Conclusions: We can quantify DRC in human PBMCs with highly reproducible results.Measurement of NER and NHEJ DRC by this high throughput, flow cytometry-based HCR assay may be an effective way of identifying DNA repair deficits, lung cancer risk, complex drivers of lung and other cancer development, and possibly in predicting response to DNAdamaging lung cancer therapeutics.
Huntington's disease is a neurodegenerative disorder resulting from an expanded polyglutamine (polyQ) repeat of the Huntingtin (Htt) protein. Affected tissues often contain aggregates of the N-terminal Htt exon 1 (Htt-Ex1) fragment. The N-terminal N17 domain proximal to the polyQ tract is key to enhance aggregation and modulate Htt toxicity. Htt-Ex1 is intrinsically disordered, yet it has been postulated that under physiological conditions membranes induce the N17 to adopt an α-helical structure, which then plays a key role in regulating Htt protein aggregation. The present study leverages the recently available assignment of NMR peaks in an N17Q17 construct, in order to provide a look into the changes occurring in vitro upon exposing this fragment to various brain extract fragments as well as to synthetic bilayers. Residue-specific changes were observed by 3D HNCO NMR, whose nature was further clarified with ancillary CD and aggregation studies, as well as with molecular dynamic calculations. From this combination of measurements and computations, a unified picture emerges, whereby transient structures consisting of α-helices spanning a fraction of the N17 residues form during N17Q17-membrane interactions. These interactions are fairly dynamic, but they qualitatively mimic more rigid variants that have been discussed in the literature. The nature of these interactions and their potential influence on the aggregation process of these kinds of constructs under physiological conditions are briefly assessed.
In cellular environments, proteins not only interact with their specific partners but also encounter a high concentration of bystander macromolecules, or crowders. Nonspecific interactions with macromolecular crowders modulate the activities of proteins, but our knowledge about the rules of nonspecific interactions is still very limited. In previous work, we presented experimental evidence that macromolecular crowders acted competitively in inhibiting the binding of maltose binding protein (MBP) with its ligand maltose. Competition between a ligand and an inhibitor may result from binding to either the same site or different conformations of the protein. Maltose binds to the cleft between two lobes of MBP, and in a series of mutants, the affinities increased with an increase in the extent of lobe closure. Here we investigated whether macromolecular crowders also have a conformational or site preference when binding to MBP. The affinities of a polymer crowder, Ficoll70, measured by monitoring tryptophan fluorescence were 3-6-fold higher for closure mutants than for wild-type MBP. Competition between the ligand and crowder, as indicated by fitting of titration data and directly by nuclear magnetic resonance spectroscopy, and their similar preferences for closed MBP conformations further suggest the scenario in which the crowder, like maltose, preferentially binds to the interlobe cleft of MBP. Similar observations were made for bovine serum albumin as a protein crowder. Conformational and site preferences in MBP-crowder binding allude to the paradigm that nonspecific interactions can possess hallmarks of molecular recognition, which may be essential for intracellular organizations including colocalization of proteins and liquid-liquid phase separation.
Objective: A method is proposed to obtain high-resolution 2-D $J$-resolved nuclear magnetic resonance (NMR) spectra in inhomogeneous magnetic fields. Methods: The proposed experiment enables the acquisition of an entire 2-D spectrum in a single scan by utilizing intermolecular double-quantum coherences and the spatial encoding of NMR observables. Results: Chemical shifts, $J$ coupling constants, and multiplet patterns are recovered even when field inhomogeneities are severe enough to completely obscure conventional NMR spectra. After intentional deshimming to yield inhomogeneous magnetic fields, the method was demonstrated on ethyl 3-bromoproprionate in acetone and on a complex mixture of organic compounds. To illustrate the technique's applicability to biological samples with intrinsic magnetic field inhomogeneities arising from macroscopic magnetic susceptibility variations, we performed the experiment on a pig bone marrow sample. Conclusion: Our results show that the new method is a fast and effective tool for studying complex chemical mixtures and biological tissues. Significance: The method could potentially be useful for real-time in vivo NMR studies.
Many neurodegenerative diseases are characterized by misfolding and aggregation of an expanded polyglutamine tract (polyQ). Huntington's Disease, caused by expansion of the polyQ tract in exon 1 of the Huntingtin protein (Htt), is associated with aggregation and neuronal toxicity. Despite recent structural progress in understanding the structures of amyloid fibrils, little is known about the solution states of Htt in general, and about molecular details of their transition from soluble to aggregation-prone conformations in particular. This is an important question, given the increasing realization that toxicity may reside in soluble conformers. This study presents an approach that combines NMR with computational methods to elucidate the structural conformations of Htt Exon 1 in solution. Of particular focus was Htt's N17 domain sited N-terminal to the polyQ tract, which is key to enhancing aggregation and modulate Htt toxicity. Such in-depth structural study of Htt presents a number of unique challenges: the long homopolymeric polyQ tract contains nearly identical residues, exon 1 displays a high degree of conformational flexibility leading to a scaling of the NMR chemical shift dispersion, and a large portion of the backbone amide groups are solvent-exposed leading to fast hydrogen exchange and causing extensive line broadening. To deal with these problems, NMR assignment was achieved on a minimal Htt exon 1, comprising the N17 domain, a polyQ tract of 17 glutamines, and a short hexameric polyProline region that does not contribute to the spectrum. A pH titration method enhanced this polypeptide's solubility and, with the aid of ≤5D NMR, permitted the full assignment of N17 and the entire polyQ tract. Structural predictions were then derived using the experimental chemical shifts of the Htt peptide at low and neutral pH, together with various different computational approaches. All these methods concurred in indicating that low-pH protonation stabilizes a soluble conformation where a helical region of N17 propagates into the polyQ region, while at neutral pH both N17 and the polyQ become largely unstructured—thereby suggesting a mechanism for how N17 regulates Htt aggregation.
The majority of biochemical and biophysical studies are performed using dilute solutions of macromolecules. In contrast, the cellular environment contains a high total concentration (up to 400 g/L) of biomacromolecules. Concentrated bystander molecules (i.e. crowders) can exert important effects on the dynamics and conformational ensembles of proteins, especially intrinsically disordered proteins (IDPs). In particular, FlgM, a regulator of the ordered synthesis of flagellar proteins, is an IDP with transient helices in the C-terminal region and becomes structured upon binding its sigma factor target. Structure formation can be induced by high concentrations of glucose and bovine serum albumin. To investigate the conformations and exchange dynamics of FlgM in dilute and crowded conditions, we carried out backbone 15N NMR relaxation and CPMG relaxation dispersion experiments. In a dilute condition, elevated transverse relaxation rates (R2) in the C-terminal region are consistent with transient secondary structure. CPMG relaxation dispersion data indicate conformational exchange throughout the protein sequence, possibly between extended conformations with few tertiary contacts and more compact conformations with some tertiary contacts. The addition of 100 g/L dextran resulted in elevation of R2 throughout the protein, suggesting an increase in helical content in both the C-terminal and N-terminal regions. Moreover, dextran led to an increase in the amplitude of dispersion, suggesting an increase in the population of compact conformations. The NMR studies lay the foundation for quantitative characterization of an IDP in dilute and crowded conditions.
Multidimensional Nuclear Magnetic Resonance (NMR) provides a unique window into structure and dynamics at an atomic level. Traditionally, given the scan-by-scan time modulation involved in these experiments, the duration of nD NMR increases exponentially with spectral dimensionality. In addition, acquisition times increase as the number of spectral elements being sought in each indirect domain - given by the ratio between the spectral bandwidth being targeted and the resolution desired. These long sampling times can be substantially reduced by exploiting information that is often available from lower-dimensionality acquisitions. This work presents a novel approach that exploits previous 2D information to speed up the acquisition of 3D spectra, based on what we denote as a Time-Optimized FouriEr Encoding (TOFEE) of pre-targeted peaks. Such 3D TOFEE experiments, which present points in common with Hadamard-encoded 3D acquisitions, do not necessarily require more scans than their 2D counterparts. This is here demonstrated based on extensions of 2D Heteronuclear Single-quantum Coherence (HSQC) experiments, to 3D HSQC-TOCSY or 3D HSQC-NOESY acquisitions. The theoretical basis of this new approach is given, and experimental demonstrations are presented on small molecule and protein-based model systems.
A half-century quest for higher magnetic fields has been an integral part of the progress undergone in the Nuclear Magnetic Resonance (NMR) study of materials’ structure and dynamics. Because 2D NMR relies on systematic changes in coherences’ phases as a function of an encoding time varied over a series of independent experiments, it generally cannot be applied in temporally unstable fields. This precludes most NMR methods from being used to characterize samples situated in hybrid or resistive magnets that are capable of achieving extremely high magnetic field strength. Recently, “ultrafast” NMR has been developed into an effective and widely applicable methodology enabling the acquisition of a multidimensional NMR spectrum in a single scan; it can therefore be used to partially mitigate the effects of temporally varying magnetic fields. Nevertheless, the strong interference of fluctuating fields with the spatial encoding of ultrafast NMR still severely restricts measurement sensitivity and resolution. Here, we introduce a strategy for obtaining high resolution NMR spectra that exploits the immunity of intermolecular zero-quantum coherences (iZQCs) to field instabilities and inhomogeneities. The spatial encoding of iZQCs is combined with a J-modulated detection scheme that removes the influence of arbitrary field inhomogeneities during acquisition. This new method can acquire high-resolution one-dimensional NMR spectra in large inhomogeneous and fluctuating fields, and it is tested with fields experimentally modeled to mimic those of resistive and resistive-superconducting hybrid magnets.