Abstract Rubisco is the entry point of nearly all organic carbon into the biosphere and is present in all domains of life. Despite its global importance, biochemical studies of this enzyme superfamily have been limited to a relatively narrow set of subclades. Recent advances in metagenomics have dramatically reshaped our understanding of both microbial and rubisco diversity; however, biochemical characterization of these sequences has not kept pace with the exponential growth in sequence data. To better survey the functional and structural diversity of rubisco, we systematically sample and synthesize a library of diverse rubisco sequences with an emphasis on clades that are sparsely represented in the biochemical literature. Our updated phylogenetic analysis reveals that many deep‑branching rubiscos assemble as dimers, supporting a dimeric origin for the superfamily — in contrast to the ecologically dominant hexadecameric form I. Additionally, we discover and structurally characterize an unusually large catalytic subunit among characterized rubiscos, originating from a early-branching subclade with secondary structural elements not present in canonical rubisco architectures.
As the largest available freshwater resource, groundwater provides water for drinking and agricultural irrigation to billions of people worldwide. Subsurface environments are estimated to host over 30% of all microorganisms on Earth, with microbial communities having important roles in various groundwater processes. Discoveries of diverse and novel lineages of archaea, bacteria, microeukaryotes and viruses in groundwater have highlighted the key contributions of groundwater microbiomes to elemental cycling, contaminant degradation, pathogenicity and antimicrobial resistance. In this Review, we summarize the diversity and composition of prokaryotic, microeukaryotic and viral communities in groundwater, and describe the complex interactions between these different microbial groups. Groundwater microbiomes exhibit distinct biogeographic patterns, with key differences and similarities across groundwater types and other habitats. The community assembly of microorganisms across different groundwater environments is driven primarily by stochastic processes, whereas deterministic processes caused by environmental stress further modulate community structure and function. With regard to their function, groundwater microbiomes have crucial roles in shaping ecosystem functioning, including biogeochemical cycling, water quality and One Health. Finally, we highlight several future prospects for research to harness microbiomes as eco-sustainable solutions for groundwater restoration and protection in a changing world. Groundwater hosts a vast array of microbiomes that are essential to global water quality and ecosystem function. This Review discusses the diversity, biogeography and community assembly of the groundwater microbiome, and explores its role in biogeochemical cycling and sustainable groundwater management.
The NMR Exchange Format (NEF) is a community-driven standard for representing NMR experimental data in a consistent, interoperable, and machine-readable form. Built on the STAR syntax, NEF provides a structured framework for storing and exchanging chemical shifts, peak lists, various types of structural restraints, and related metadata, thus allowing for data exchange across software platforms. By enabling direct, lossless transfer of information, NEF simplifies multi-software workflows, improves reproducibility, and supports FAIR (Findable, Accessible, Interoperable, Reusable) data principles. We describe the NEF specification, its current implementation across commonly used NMR software packages, and its application in areas including biomolecular structure determination, metabolomics, and ligand screening. Testing demonstrates that NEF can be used to exchange complete datasets between programs without loss of information or functionality. We also outline recent developments and future directions, such as inclusion of NMR relaxation data and support for non-standard residue topologies. NEFs growing adoption highlights its potential as a unifying standard for NMR data, enabling more efficient, transparent and collaborative research.
During photosynthetic water oxidation, the Mn4Ca cluster in Photosystem II progresses through five intermediate Si (i = 0-4) states. X-ray crystallography studies have reported the insertion of one new O ligand during the formation of the S3 state, but recent studies question the presence of this additional ligand based on cryo-EM and earlier room-temperature crystallography data. There is also controversy about whether the O-O bond interaction already occurs in the S3 state or in the subsequent S3 to S0 transition. Here we report conventional high-resolution data for the S1, S2, and S3 states to a resolution of ~1.9 Å, and anomalous diffraction data at two energies (9.5 keV and 7 keV), that was used to model the Mn positions, followed by determination of oxygen positions using the high-resolution maps. We show that the new oxygen atom, OX (or O6), in the S3 state is observable as a distinct peak without any restraints, confirming its ligation to Mn1 and Ca. The OX-O5 distance is ~2.1 Å, supporting no strong interaction between them in the S3 state, suggesting that if this is the O-O bond formation site, it is formed during the S3 to S0 transition initiated by the final oxidation of the cluster.
Multipolar scattering models, such as the transferable aspherical atom model, account for atomic chemical interactions and provide a more accurate representation of experimental data. However, the simpler independent atom model (IAM), which assumes non-interacting atoms, is the only model available in the most widely used macromolecular refinement programs. This is primarily because IAM offers a hard-to-beat combination of computational efficiency and modelling power at typical macromolecular resolutions. By contrast, more accurate multipolar modelling has historically been limited due to its computational cost and the absence of an interface between software capable of calculating structure factors and gradients based on multipolar models and software designed for macromolecular refinement. This work introduces pyDiSCaMB, a Python software package designed to integrate between the computational crystallography toolbox (cctbx) and the quantum crystallography library DiSCaMB (Densities in Structural Chemistry and Molecular Biology), thus enabling multipolar scattering models in Phenix's toolkit. The implementation, features and capabilities of pyDiSCaMB are presented, the runtimes for the calculation of structure factor and target gradients with respect to atomic parameters are explored, and Fourier images of electrostatic potential, electron density and deformation maps are computed as illustrative examples. The pyDiSCaMB library will make multipolar modelling widely available to the structural biology community, potentially transforming refinement and model-building for both crystallography and cryogenic electron microscopy (cryoEM).
In macromolecular structure refinement the low observation-to-parameter ratio and the lack of high-resolution data is countered by using a priori information in the form of restraints. Having accurate geometries of the chemical entities in the sample is paramount for generating accurate chemical restraints and, therefore, accurate macromolecular structures. In particular, it is desirable to have accurate restraints for known and novel ligand entities. Quantum Mechanics (QM) can minimise the energy of a ligand by adjusting its geometry, and these geometries can be used to generate restraints macromolecular refinement. We describe here a library of 37,000 small molecules extracted from the Chemical Components Dictionary in the Protein Data Bank and minimized by density functional QM. The library includes restraint files for use in crystallography or cryo-EM refinement, along with files suitable for molecular dynamics simulation. Because the geometries are validated, the restraints library provides users with both functional restraints and minimised geometries. This work also provides procedures for generating new and accurate restraints.
Calculation of density maps from atomic models is essential for structural studies using crystallography and electron cryo-microscopy (cryoEM). These maps serve various purposes, including atomic model building, refinement, visualization, and validation. However, accurately comparing model-calculated maps to experimental data poses challenges, particularly because the resolution of cryoEM experimental maps varies across the map. Traditional crystallography methods generate finite-resolution maps with uniform resolution throughout the unit cell volume, while most modern software in cryoEM employ Gaussian-like functions to generate these maps, which does not adequately account for atomic model parameters and resolution. Recent work by Urzhumtsev & Lunin (2022, IUCr Journal, 9, 728-734) introduces a novel method for computing atomic model maps that incorporate local resolution and can be expressed as analytically differentiable functions of all atomic parameters. This approach enhances the accuracy of matching atomic models to experimental maps. In this paper, we detail the implementation of this method in CCTBX and Phenix.
Synthetic biology generates vast combinatorial designs, yet high-throughput analytical methods to screen them are poorly matched to interrogate this search space. We address this challenge by developing a biosensor-driven, growth-coupled selection strategy in Pseudomonas putida for isoprenol, a potential aviation fuel precursor. We found and characterized a noncanonical signaling pathway, revealing a functional and physical complex between a hybrid histidine kinase and an alcohol dehydrogenase, whose activity is tuned by heterodimerization. Leveraging this biosensor in a pooled CRISPRi library selection, we identified key host limitations. Iterative combinatorial strain engineering derived from these hits yielded a 36-fold titer increase to ~900 milligrams per liter. Integrated omics analysis revealed that metabolic rewiring toward amino acid catabolism was crucial for this improvement. This observation was found to be beneficial by technoeconomic analysis. Our modular workflow provides a powerful strategy for optimizing complex heterologous pathways and uncovering emergent host biology.
Cabrerite (IMA2023-123), NiMg2(AsO4)(2)8H(2)O, is a newly approved mineral species from the Nickel mine, Cottonwood Canyon, Table Mountain district, Churchill County, Nevada, U.S.A., that was originally described in 1863 from Sierra Cabrera, Almer & iacute;a, Andalusia, Spain. At the Nickel mine, cabrerite occurs in divergent groups of green blades up to 1 mm long. Blades are elongated and striated parallel to [001], flattened on {010}, and exhibit the forms {010}, {110}, and {201}. The mineral is transparent with vitreous luster and white streak. The Mohs hardness is similar to 2 1/2. The mineral has moderately sectile tenacity, irregular and stepped fracture, and three cleavages: perfect on {010}, fair on {100}, and poor on {102}. The measured density is 2.93(2) gcm(-3). The mineral dissolves slowly in RT dilute HCl. The mineral is optically biaxial (+), alpha = 1.609(2), beta = 1.633(2), gamma = 1.667(2) (white light); 2V(meas) = 82(2)degrees; slight r < v dispersion; orientation X = b, Z <<^>> c = 37 degrees in obtuse beta; nonpleochroic. Electron probe microanalysis provided the empirical formula (Mg1.46Ni1.55)(Sigma 3.01)(As1.00O4)(2)8H(2)O. Cabrerite is monoclinic, C2/m, a = 10.2054(11), b = 13.3772(13), c = 4.7382(4) & Aring;, beta = 105.057(7)degrees, V = 624.66(11) & Aring;(3), and Z = 2. The mineral has a vivianite-type structure (R-1 = 0.0353 for 668 I > 2 sigma(I) reflections) in which MlO(2)(H2O)(4) octahedra and M2(2)O(6)(H2O)(4) edge-sharing octahedral dimers are linked together via TO4 tetrahedra (where T = P or As), and hydrogen bonds to form layers parallel to {010}; successive layers are linked by hydrogen bonds only. Cabrerite is the ordered intermediate between hornesite and annabergite, with Ni dominant at M1, Mg dominant at M2, and T = As.
Quantum Mechanical methods provide geometries and energies of molecules using just the atomic and electronic positions. Being independent of the experimental data, they provide complementary information. Furthermore, allowing the experimental and QM methods to share information leads to better results. The Quantum Interface (QI) in Phenix (Liebschner et al., 2019) provides close integration with MOPAC (Moussa & Stewart, 2024) allowing the calculation of in situ restraints for drug candidates. Known as QM Restraints (QMR) (Liebschner et al., 2023), this method will provide protein binding pocket specific restraints during a refinement or in a stand-alone program. Another QI procedure is a novel approach for predicting histidine protonation states using QM methods. Historically, determining histidine protonation has been challenging due to limited resolution in X-ray crystallography and the inherent difficulty in detecting hydrogen atoms. Previous methods relied on empirical or geometric models, which provided some insights but had limitations. The proposed method, Quantum Mechanical Flipping (QMF) (Moriarty et al., In review), employs quantum mechanical calculations to predict protonation states based on the molecular environment. This approach considers all possible configurations of histidine protonation and assesses their feasibility by minimising geometry and energy calculations. QMF accounts for factors such as hydrogen bonding and molecular interactions providing a more accurate prediction of the most likely protonation state particularly in the binding pocket. The study demonstrates the effectiveness of QMF through comparisons with existing methods and validation using high- resolution protein structures. Results show that QMF can accurately predict histidine protonation states, even in cases with limited experimental data. Additionally, QMF's versatility allows it to be applied to various macromolecular environments, including ligand interactions and non-standard amino acids. Other applications of QI include metal coordination and ligand strain energies, both of which are being pursued.
Hoperanchite (IMA2024-017), (NH4)2(S2O3), is a newly approved mineral species from an active vent in a burning bituminous shale at Hope Ranch, Santa Barbara County, California, U.S.A. Hoperanchite occurs as tabular crystals up to about 0.3 mm in diameter. Tablets are flattened on {001} and exhibit the forms {100}, {010}, {010}, {110}, and {110}. The mineral is colorless and transparent with a vitreous luster and white streak. The Mohs hardness is similar to 2 1/2. The mineral has brittle tenacity, irregular fracture, and one good cleavage on {001}. The measured density is 1.68(2) gcm-3. The mineral dissolves instantly in H2O at room temperature. The mineral is optically biaxial (+), alpha = 1.602(2), beta = 1.616(2), gamma = 1.634(2) (white light); 2Vmeas = 84(2)degrees; orientation Y = b, X<^>a = 20 degrees in obtuse beta; nonpleochroic. Electron microprobe analysis provided the empirical formula N1.89H7.40S1.97O3. Hoperanchite is monoclinic, C2, a = 10.2313(5), b = 6.4998(3), c = 8.8098(6) & Aring;, beta = 94.611(7)degrees, V = 583.97(6) & Aring;3, and Z = 4. The crystal structure (R1 = 0.0325 for 1176 I > 2 sigma I reflections) is the same as that of synthetic (NH4)2(S2O3). Hoperanchite is the first thiosulfate mineral that does not contain essential Pb.
Synthetic biology tools have accelerated the generation of simple mutants, but combinatorial testing remains a major hurdle. High-throughput methods struggle translating from proof-of-principle molecules to advanced bioproducts. We address this challenge with a biosensor-driven strategy for enhanced isoprenol production in Pseudomonas putida , a key precursor for sustainable aviation fuel and platform chemicals. This biosensor leverages P. putida's native response to short-chain alcohols via a previously uncharacterized hybrid histidine kinase signaling cascade. Refactoring the biosensor for a conditional growth-based selection enabled identification of competing cellular processes with a ~16,500-member CRISPRi-library. An iterative combinatorial strain engineering approach yielded an integrated P. putida strain producing ~900 mg/L isoprenol in glucose minimal medium, a 36-fold increase. Ensemble -omics analysis revealed metabolic rewiring, including amino acid accumulation as key drivers of enhanced production. Techno-economic analysis elucidated the path to economic viability and confirmed the benefits of adding amino acids outweigh the additional costs. This study establishes a robust biosensor driven approach for optimizing other heterologous pathways, accelerating microbial cell factory development. ### Competing Interest Statement TE, JM, and AM are inventors on a patent application related to the workflows described in this report (LBNL Docket 2024-064-01; US Patent Application No. 63/761,819). NRB has a financial interest in Erg Bio. CDS has a financial interest in Cyklos Materials.
Microbial taxonomic diversity declines with increased environmental stress. Yet, few studies have explored whether phylogenetic and functional diversities track taxonomic diversity along the stress gradient. Here, we investigated microbial communities within an aquifer in Oak Ridge, Tennessee, USA, which is characterized by a broad spectrum of stressors, including extremely high levels of nitrate, heavy metals like cadmium and chromium, radionuclides such as uranium, and extremely low pH (< 3). Both taxonomic and phylogenetic α-diversities were reduced in the most impacted wells, while the decline in functional α-diversity was modest and statistically insignificant, indicating a more robust buffering capacity to environmental stress. Differences in functional gene composition (i.e., functional β-diversity) were pronounced in highly contaminated wells, while convergent functional gene composition was observed in uncontaminated wells. The relative abundances of most carbon degradation genes were decreased in contaminated wells, but genes associated with denitrification, adenylylsulfate reduction, and sulfite reduction were increased. Compared to taxonomic and phylogenetic compositions, environmental variables played a more significant role in shaping functional gene composition, suggesting that niche selection could be more closely related to microbial functionality than taxonomy. Overall, we demonstrated that despite a reduced taxonomic α-diversity, microbial communities under stress maintained functionality underpinned by environmental selection.
Cryo-EM and X-ray crystallography provide crucial experimental data for obtaining atomic-detail models of biomacromolecules. Refining these models relies on library-based stereochemical data, which, in addition to being limited to known chemical entities, do not include meaningful noncovalent interactions. Quantum mechanical (QM) calculations could alleviate these issues but are too expensive for large molecules. Here we present a novel AI-enabled Quantum Refinement (AQuaRef) based on AIMNet2 machine learned interatomic potential (MLIP) mimicking QM at substantially lower computational costs. By refining 41 cryo-EM and 30 X-ray structures, we show that this approach yields atomic models with superior geometric quality compared to standard techniques, while maintaining an equal or better fit to experimental data. Notably, AQuaRef aids in determining proton positions, as illustrated in the challenging case of short hydrogen bonds in the parkinsonism-associated human protein DJ-1 and its bacterial homolog YajL.
Rubisco is the entry point of nearly all organic carbon into the biosphere and is present in all domains of life. Despite its global importance, biochemical studies of this enzyme superfamily have been limited to a relatively narrow set of subclades. Recent advances in metagenomics have dramatically reshaped our understanding of both microbial and rubisco diversity; however, biochemical characterization of these sequences has not kept pace with the exponential growth in sequence data. To better survey the functional and structural diversity of rubisco, we systematically sampled and synthesized a library of diverse rubisco sequences with an emphasis on clades that have previously not been characterized. Our updated phylogenetic analysis reveals that many deep‑branching rubiscos assemble as dimers, supporting a dimeric origin for the superfamily -- in contrast to the ecologically dominant hexadecameric form I. Additionally, we discover and structurally characterize the largest rubisco described to date, originating from a cryptic, early-branching subclade with novel structural folds that have previously not been observed in the rubisco superfamily. By integrating biochemical data with an updated phylogenetic framework, we propose a revised nomenclature for the rubisco protein family that reflects current insights and will better accommodate future discoveries.
Xylan, the most abundant non-cellulosic polymer in plant cell walls, is structurally diverse, especially in grasses where it is heavily substituted with arabinofuranose and further modified by various residues. Common substitutions across species include glucuronic and 4- O -methyl-glucuronic acid. Arabinose and xylose sidechains are synthesized by glycosyltransferase family 61 (GT61) proteins, many of which remain uncharacterized in plants, with limited structural and mechanistic understanding. In this study, we identified two novel GT61 enzymes in Sorghum bicolor , functioning as xylan arabinosyltransferase (SbXAT) and xylan xylosyltransferase (SbXXT). We resolved the crystal structure of SbXAT, which exhibits a GT-B fold with two Rossmann-like domains linked by a cleft that accommodates the catalytic site. Structural comparison with a predicted SbXXT model revealed a substrate-binding residue critical for sugar donor specificity, validated through site-directed mutagenesis and enzymatic assays. These findings enhance understanding of xylan biosynthesis and provide a foundation for engineering glycosyltransferases and predicting their functions.
Advances in genome engineering have improved our ability to perturb microbial metabolic networks, yet bioproduction campaigns often struggle with parsing complex metabolic datasets to efficiently enhance product titers. We address this challenge by coupling laboratory automation with machine learning to systematically optimize the production of isoprenol, a sustainable aviation fuel precursor, in Pseudomonas putida. The simultaneous downregulation through CRISPR interference of combinations of up to four gene targets, guided by machine learning, permitted us to increase isoprenol titer 5-fold in six consecutive design-build-test-learn cycles. Moreover, machine learning enabled us to swiftly explore a vast experimental design space of 800,000 possible combinations by strategically recommending approximately 400 priority constructs. High-throughput proteomics allowed us to validate CRISPRi downregulation and identify biological mechanisms driving production increases. Our work demonstrates that ML-driven automated design-build-test-learn cycles, when combined with rigorous data validation, can rapidly enhance titers without specific biological knowledge, suggesting that it can be applied to any host, product, or pathway. Laboratory automation, machine learning, and metabolic engineering may be combined to quickly and efficiently build productive microbial strains. Here the authors used these techniques in P. putida to boost isoprenol titers 5-fold over six DBTL cycles while sampling a reduced design space.
The interpretation of cryo-EM maps often includes the docking of known or predicted structures of the components, which is particularly useful when the map resolution is worse than 4 Å. Although it can be effective to search the entire map to find the best placement of a component, the process can be slow when the maps are large. However, frequently there is a well-founded hypothesis about where particular components are located. In such cases, a local search using a map subvolume will be much faster because the search volume is smaller, and more sensitive because optimizing the search volume for the rotation-search step enhances the signal to noise. A Fourier-space likelihood-based local search approach, based on the previously published em_placement software, has been implemented in the new emplace_local program. Tests confirm that the local search approach enhances the speed and sensitivity of the computations. An interactive graphical interface in the ChimeraX molecular-graphics program provides a convenient way to set up and evaluate docking calculations, particularly in defining the part of the map into which the components should be placed.
Advances in machine learning have enabled sufficiently accurate predictions of protein structure to be used in macromolecular structure determination with crystallography and cryo-electron microscopy data. The Phenix software suite has AlphaFold predictions integrated into an automated pipeline that can start with an amino acid sequence and data, and automatically perform model-building and refinement to return a protein model fitted into the data. Due to the steep technical requirements of running AlphaFold efficiently, we have implemented a Phenix-AlphaFold webservice that enables all Phenix users to run AlphaFold predictions remotely from the Phenix GUI starting with the official 1.21 release. This webservice will be improved based on how it is used by the research community and the future research directions for Phenix.