Cyanophycin is a nitrogen- and carbon-rich reserve biopolymer conserved across diverse microbial taxa and ecological habitats, yet its in situ distribution and ecological role remain poorly understood due to limitations in existing detection methods. Here, we present a high-resolution, label-free, and extraction-free method for detecting and semiquantifying intracellular cyanophycin using single-cell Raman spectroscopy (SCRS) integrated with explainable machine learning. Genetically engineered cyanophycin-producing and nonproducing strains of Acinetobacter baylyi and Escherichia coli are used to provide robust positive and negative controls. Our explainable machine learning approach uncovered that Raman peaks at 898.0, 982.0, 1230.0, and 1674.0 cm-1 should be used to identify intracellular cyanophycin. A linear dose-response relationship (R 2 = 0.9924) confirmed the semiquantitative capability of SCRS for intracellular cyanophycin detection. In a proof-of-concept study using full-scale enhanced biological phosphorus removal system biomass spiked with engineered cyanophycin-producing A. baylyi ADP1-ISx, SCRS successfully detected cyanophycin-positive cells within complex microbial matrices. These findings establish SCRS as a powerful tool for noninvasive monitoring of nitrogen polymer storage in environmental microbiomes at the single-cell level, offering new opportunities for understanding and managing microbial nitrogen cycling in engineered ecosystems.
The design of pathways to synthesize valuable molecules remains a central challenge in chemistry and biotechnology. Several computational retrosynthesis tools have been developed to address this problem, but their scope is often confined only to reactions in either synthetic organic chemistry or monofunctional enzymatic chemistry. We present TridentSynth, a web-based retrosynthesis tool (https://tridentsynth.lbl.gov) to scale synthesis planning up to three different routes by also incorporating multifunctional Type I polyketide synthase (PKS) enzymes into our reaction toolkit along with organic chemistry and monofunctional enzymes. Unlike monofunctional enzymes that catalyze single transformations, PKSs function as molecular assembly lines that catalyze multiple carbon-carbon bond formation reactions between acyl-coenzyme A substrates to construct elongated carbon scaffolds. PKSs follow a modular, programmable logic that allows them to be reconfigured to make new molecules in a predictable way. These scaffolds can then be chemoenzymatically modified to eventually access a wider array of molecular targets than would be possible with just synthetic chemistry or monofunctional enzymes alone, in a manner that mimics the evolved biosynthesis routes of many useful natural products. TridentSynth assists synthetic biologists by suggesting routes to synthesize a desired molecule through an intuitive web interface that requires no local installation or programming expertise.
Uncharacterized functions of enzymes represent an untapped opportunity to develop therapeutics, unlock the sustainable synthesis of materials, and understand the evolution of life-sustaining metabolic networks. Uncharacterized enzymes and reactions, generated by protein language models and computer-aided synthesis tools, respectively, make up a large part of this opportunity. Given the technical complexity of high-throughput enzymatic activity screens, predictive models are needed that can prescreen enzyme-reaction pairs in silico. We present (1) a high-quality data set of enzyme-reaction pairs, (2) a rigorous battery of model evaluations varying in their approaches to data splitting and negative sampling, (3) a comprehensive benchmarking of enzyme-reaction models, and (4) a pair of parameter-efficient, data-efficient, high-performing models called Reaction-Center Graph Neural Networks (RC-GNNs) capable of predicting whether an enzyme, represented by an amino acid sequence, can significantly catalyze a given reaction, represented by its full set of reactants and products. In the most difficult conditions, where the query reactions were highly dissimilar from those present in the training data set, our models achieved 0.88 and 0.84 ROC-AUC on classification tasks featuring globally selected and synthetic negatives, respectively. On a time-based split, an RC-GNN achieved 0.91 ROC-AUC. The ability to successfully make predictions on enzymes and reactions distinct from those used during training makes the RC-GNNs especially useful for both metabolic engineers and evolutionary biologists who need to reason about uncharacterized enzymatic reactions.
ABSTRACT In this work, we build a mechanistic kinetic model for a phosphate‐based enzyme cascade for high‐yield fructose production from waste starch, a potentially industrially impactful cascade whose kinetics have remained underexplored. The cascade proceeds through sugar phosphate intermediates catalyzed by alpha‐glucan phosphorylase (AGP), phosphoglucomutase (PGM), phosphoglucoisomerase (PGI), transaldolase (TRA), and 3‐phosphoglycerate phosphatase (3‐PGP), thus enabling a kinetically driven route that overcomes equilibrium limitations of conventional glucose isomerization. The intrinsic kinetic behavior of each enzyme was incorporated to build the overall deterministic kinetic model for the cascade. Thermodynamic and kinetic parameters were obtained or estimated from literature and theory, and a small number of highly sensitive parameters were fit to batch experimental data. Metabolic control analysis and net‐rate analysis revealed transaldolase as the primary driver of fructose yield. We also highlight the power of this model by providing microscopic details such as coverages of various species bound to enzymes and using these insights in proposing experimental conditions that can improve fructose yields. Through this work, we emphasize that this modeling framework not only has the power to propose experimental conditions that improve the yield of the desired fructose product but also provide fundamental insights into the behavior of enzyme cascades.
Enzymes catalyze reactions with remarkable specificity and can unlock recalcitrant feedstocks that are dilute, complex, and variable in their constituent molecules. While characterized enzymatic reactions cover a wide range of chemistries, there are an undetermined number of cryptic activities for every known one. These cryptic activities can be elicited through rational design, adaptive laboratory evolution, and increasingly, generative models of proteins. However, prior to tuning a catalyst, one must efficiently predict viable novel reactions. In this work, we leverage the growing amount of mechanistic enzyme information, specifically the Mechanism and Catalytic Site Atlas, to construct a set of reaction rules that can meet this demand. By explicitly utilizing mechanistic information, the rule sets developed here more accurately identify molecular structures required for catalysis compared to existing curated and heuristically constructed rules. The 899 Distilled rules are constructed directly from characterized mechanisms and recapitulate 62.5% of atom-mapped reactions from Rhea. The Learned rule set is generated from a classifier trained on structural patterns putatively required for catalytic mechanisms. The Learned rules recapitulate all atom-mapped Rhea reactions and precisely predict mechanism-required atoms (ROC-AUC = 0.98). Additionally, our Learned rules exhibit a more favorable trade-off between novelty and feasibility and provide users with fine-grained control over this trade-off. The rules are compatible with all SMARTS-based reaction network expansion and retrosynthesis software.
Recovering nitrogen (N) from wastewater is a potential avenue to reduce reliance on energy-intensive synthetic nitrogen fixation via Haber-Bosch and subsequent treatment of N-laden wastewaters through nitrification-denitrification. However, many technical and economic factors hinder widespread application of N recovery, particularly low N concentrations in municipal wastewater, paucity of high-efficiency separations technologies compatible with biological treatment, and suitable products and markets for recovered N. In this perspective, we contextualize the challenges of N recovery today, propose integrated biological and physicochemical technologies to improve selective and tunable N recovery, and propose an expanded product portfolio for recovered N products beyond fertilizers. We highlight cyanophycin, an N-rich biopolymer produced by a diverse range of bacteria, as a potential target for N bioconcentration and downstream recovery from municipal wastewater. This perspective emphasizes the equal importance of integrated biological systems, physicochemical separations, and market assessment in advancing nitrogen recovery from wastewater.
Uncharacterized functions of enzymes represent untapped opportunity to develop therapeutics, unlock the sustainable synthesis of materials, and understand the evolution of life-sustaining metabolic networks. Enzymes and de novo reactions (i.e., non-native, promiscuous reactions), generated by protein language models and computer-aided synthesis tools, respectively, make up a large part of this opportunity. Given the technical complexity of high-throughput enzymatic activity screens, predictive models are needed that can pre-screen de novo enzyme-reaction pairs in silico. We present Reaction-Center Graph Neural Network, (RC-GNN) a model capable of predicting whether an enzyme, represented by an amino acid sequence, can significantly catalyze a given reaction, represented by its full set of reactants and products. We explicitly evaluated RC-GNN's generalization to de novo queries. In the most difficult conditions tested, where difficulty is measured by the level of dissimilarity between training and test data points, the model achieves 78.0% and 94.8% accuracy when reaction and enzyme similarity were respectively controlled. The ability to successfully make predictions on enzymes and reactions distinct from those used during training make RC-GNN especially useful for both metabolic engineers and evolutionary biologists who need to reason about uncharacterized enzymatic reactions.
Synthetic biology offers the promise of manufacturing chemicals more sustainably than petrochemistry. Yet, both the rate at which biomanufacturing can synthesize these molecules and the net chemical accessible space are limited by existing pathway discovery methods, which can often rely on arduous literature searches. Here, we introduce BioPKS pipeline, an automated retrobiosynthesis tool combining multifunctional type I polyketide synthases (PKSs) and monofunctional enzymes via two complementary tools: RetroTide and DORAnet. Monofunctional enzymes are valuable for carefully decorating a substrate's carbon backbone while PKSs are unique in their ability to iteratively catalyze carbon-carbon bond formation reactions, thereby expanding carbon backbones in a predictable fashion. We evaluate the performance of BioPKS pipeline using a previously reported set of 155 biomanufacturing candidates, achieving exact synthetic designs for 93 compounds and generating chemically similar pathways for most remaining targets. Furthermore, BioPKS pipeline can propose pathways for the complex therapeutic natural products cryptofolione and basidalin.
Harnessing DNA as a high-density storage medium for information storage and molecular recording of signals has been of increasing interest in the biotechnology field. Recently, progress in enzymatic DNA synthesis, DNA digital data storage, and DNA-based molecular recording has been made by leveraging the activity of the template-independent DNA polymerase, terminal deoxynucleotidyl transferase (TdT). TdT adds deoxyribonucleotides to the 3' end of single-stranded DNA, generating random sequences of single-stranded DNA. TdT can use several divalent cations for its enzymatic activity and exhibits shifts in deoxyribonucleotide incorporation frequencies in response to changes in its reaction environment. However, there is limited understanding of sequence-structure-function relationships regarding these properties, which in turn limits our ability to modulate TdT to further advance TdT-based tools. Most TdT literature to-date explores the activity of murine, bovine or human TdTs; studies probing TdT sequence and structure largely focus on strictly conserved residues that are functionally critical to TdT activity. Here, we explore non-conserved TdT sequence space by surveying the natural diversity of TdT. We characterize a diverse set of TdT homologs from different organisms and identify several TdT residues/regions that confer differences in TdT behavior between homologs. The observations in this study can design rules for targeted TdT libraries, in tandem with a screening assay, to modulate TdT properties. Moreover, the data can be useful in guiding further studies of potential residues of interest. Overall, we characterize TdTs that have not been previously studied in the literature, and we provide new insights into TdT sequence-function relationships.
High resolution cellular signal encoding is critical for better understanding of complex biological phenomena. DNA-based biosignal encoders alter genomic or plasmid DNA in a signal dependent manner. Current approaches involve the signal of interest affecting a DNA edit by interacting with a signal specific promoter which then results in expression of the effector molecule (DNA altering enzyme). Here, we present the proof of concept of a biosignal encoding system where the enzyme terminal deoxynucleotidyl transferase (TdT) acts as the effector molecule upon directly interacting with the signal of interest. A template independent DNA polymerase (DNAp), TdT incorporates nucleotides at the 3' OH ends of DNA substrate in a signal dependent manner. By employing CRISPR-Cas9 to create double stranded breaks in genomic DNA, we make 3'OH ends available to act as substrate for TdT. We show that this system can successfully resolve and encode different concentrations of various biosignals into the genomic DNA of HEK-293T cells. Finally, we develop a simple encoding scheme associated with the tested biosignals and encode the message "HELLO WORLD" into the genomic DNA of HEK-293T cells at a population level with 91% accuracy. This work demonstrates a simple and engineerable system that can reliably store local biosignal information into the genomes of mammalian cell populations.
Retrobiosynthesis tools harness the inherent promiscuities of enzymes for the de novo design of novel biosynthetic pathways to key small molecules. Many existing pathway search algorithms rely on exhaustively enumerating the space of all possible enzymatic reactions using generalized rules, followed by an extensive analysis of the ensuing reaction network to extract candidate pathways for experimental validation. While this approach is comprehensive, many false positive reactions are often generated given the permissiveness of such reaction rules. Here, we have developed DORA-XGB, a enzymatic reaction feasibility classifier. DORA-XGB can be used within our DORAnet framework to assess whether newly enumerated enzymatic reactions and pathways would be feasible. To curate a training dataset for our model, we extracted enzymatic reactions from public databases and screened them for their general thermodynamic feasibility. We then considered alternate reaction centers on known substrates to strategically generate infeasible reactions with high confidence, thereby circumventing the lack of negative data in the literature. In training our model, we also experimented with various molecular fingerprinting techniques and configurations for assembling reaction fingerprints, taking into account not just primary substrate and primary product structures, but cofactor structures as well. Our model's utility is demonstrated through favorable benchmarking against a previously published classifier, the successful recovery of newly published reactions, and the ranking of previously predicted pathways for the biosynthesis of propionic acid from pyruvate.
Enzymatic DNA writing technologies based on the template-independent DNA polymerase terminal deoxynucleotidyl transferase (TdT) have the potential to advance DNA information storage. TdT is unique in its ability to synthesize single-stranded DNA de novo but has limitations, including catalytic inhibition by ribonucleotide presence and slower incorporation rates compared to replicative polymerases. We anticipate that protein engineering can improve, modulate, and tailor the enzyme's properties, but there is limited information on TdT sequence-structure-function relationships to facilitate rational approaches. Therefore, we developed an easily modifiable screening assay that can measure the TdT activity in high-throughput to evaluate large TdT mutant libraries. We demonstrated the assay's capabilities by engineering TdT mutants that exhibit both improved catalytic efficiency and improved activity in the presence of an inhibitor. We screened for and identified TdT variants with greater catalytic efficiency in both selectively incorporating deoxyribonucleotides and in the presence of deoxyribonucleotide/ribonucleotide mixes. Using this information from the screening assay, we rationally engineered other TdT homologues with the same properties. The emulsion-based assay we developed is, to the best of our knowledge, the first high-throughput screening assay that can measure TdT activity quantitatively and without the need for protein purification.
Terminal deoxynucleotidyl transferase (TdT) is a unique DNA polymerase capable of template-independent extension of DNA. TdT's de novo DNA synthesis ability has found utility in DNA recording, DNA data storage, oligonucleotide synthesis, and nucleic acid labeling, but TdT's intrinsic nucleotide biases limit its versatility in such applications. Here, we describe a multiplexed assay for profiling and engineering the bias and overall activity of TdT variants with high throughput. In our assay, a library of TdTs is encoded next to a CRISPR-Cas9 target site in HEK293T cells. Upon transfection of Cas9 and sgRNA, the target site is cut, allowing TdT to intercept the double-strand break and add nucleotides. Each resulting insertion is sequenced alongside the identity of the TdT variant that generated it. Using this assay, 25,623 unique TdT variants, constructed by site-saturation mutagenesis at strategic positions, were profiled. This resulted in the isolation of several altered-bias TdTs that expanded the capabilities of our TdT-based DNA recording system, Cell HistorY Recording by Ordered InsertioN (CHYRON), by increasing the information density of recording through an unbiased TdT and achieving dual-channel recording of two distinct inducers (hypoxia and Wnt) through two differently biased TdTs. Select TdT variants were also tested in vitro, revealing concordance between each variant's in vitro bias and the in vivo bias determined from the multiplexed high throughput assay. Overall, our work and the multiplex assay it features should support the continued development of TdT-based DNA recorders, in vitro applications of TdT, and further study of the biology of TdT.
Synthetic biology allows us to reuse, repurpose, and reconfigure biological systems to address society's most pressing challenges. Developing biotechnologies in this way requires integrating concepts across disciplines, posing challenges to educating students with diverse expertise. We created a framework for synthetic biology training that deconstructs biotechnologies across scales-molecular, circuit/network, cell/cell-free systems, biological communities, and societal-giving students a holistic toolkit to integrate cross-disciplinary concepts towards responsible innovation of successful biotechnologies. We present this framework, lessons learned, and inclusive teaching materials to allow its adaption to train the next generation of synthetic biologists.
AbstractAchieving sustainable chemical synthesis and a circular economy will require process innovation to minimize or recover existing waste streams. Valorization of lignin biomass has the ability to advance this goal. While lignin has proved a recalcitrant feedstock for upgrading, biological approaches can leverage native microbial metabolism to simplify complex and heterogeneous feedstocks to tractable starting points for biochemical upgrading. Recently, we demonstrated that one microbe with lignin relevant metabolism,Acinetobacter baylyiADP1, is both highly engineerable and capable of undergoing rapid design-build-test-learn cycles, making it an ideal candidate for these applications. Here, we utilize these genetic traits and ADP1’s native β-ketoadipate metabolism to convert mock alkali pretreated liquor lignin (APL) to two valuable natural products, vanillin-glucoside and resveratrol. En route, we create strains with up to 22 genetic modifications, including up to 8 heterologously expressed enzymes. Our approach takes advantage of preexisting aromatic species in APL (vanillate, ferulate, andp-coumarate) to create shortened biochemical routes to end products. Together, this work demonstrates ADP1’s potential as a platform for upgrading lignin waste streams and highlights the potential for biosynthetic methods to maximize the existing chemical potential of lignin aromatic monomers.
Recovering nitrogen (N) from municipal wastewater is a promising approach to prevent nutrient pollution, reduce energy use, and transition toward a circular N bioeconomy, but remains a technologically challenging endeavor. Existing N recovery techniques are optimized for high-strength, low-volume wastewater. Therefore, developing methods to concentrate dilute N from mainstream wastewater will bridge the gap between existing technologies and practical implementation. The N-rich biopolymer cyanophycin is a promising candidate for N bioconcentration due to its pH-tunable solubility characteristics and potential for high levels of accumulation. However, the cyanophycin synthesis pathway is poorly explored in engineered microbiomes. In this study, we analyzed over 3,700 publicly available metagenome assembled genomes (MAGs) and found that the cyanophycin synthesis gene cphA was ubiquitous across common activated sludge bacteria. We found that cphA was present in common phosphorus accumulating organisms (PAO) Ca. ‘Accumulibacter’ and Tetrasphaera, suggesting potential for simultaneous N and P bioconcentration in the same organisms. Using metatranscriptomic data, we confirmed the expression of cphA in lab-scale bioreactors enriched with PAO. Our findings suggest that cyanophycin synthesis is a ubiquitous metabolic activity in activated sludge microbiomes. The possibility of combined N and P bioconcentration could lower barriers to entry for N recovery, since P concentration by PAO is already a widespread biotechnology in municipal wastewater treatment. We anticipate this work to be a starting point for future evaluations of combined N and P bioaccumulation, with the ultimate goal of advancing widespread adoption of N recovery from municipal wastewater.
Achieving sustainable chemical synthesis and a circular economy will require process innovation to minimize or recover existing waste streams. Valorization of lignin biomass has the ability to advance this goal. While lignin has proved a recalcitrant feedstock for upgrading, biological approaches can leverage native microbial metabolism to simplify complex and heterogeneous feedstocks to tractable starting points for biochemical upgrading. Recently, we demonstrated that one microbe with lignin relevant metabolism, Acinetobacter baylyi ADP1, is both highly engineerable and capable of undergoing rapid design-build-test-learn cycles, making it an ideal candidate for these applications. Here, we utilize these genetic traits and ADP1’s native β-ketoadipate metabolism to convert mock alkali pretreated liquor lignin (APL) to two valuable natural products, vanillin-glucoside and resveratrol. En route, we create strains with up to 22 genetic modifications, including up to 8 heterologously expressed enzymes. Our approach takes advantage of preexisting aromatic species in APL (vanillate, ferulate, and p-coumarate) to create shortened biochemical routes to end products. Together, this work demonstrates ADP1’s potential as a platform for upgrading lignin waste streams and highlights the potential for biosynthetic methods to maximize the existing chemical potential of lignin aromatic monomers.
BACKGROUND:Biochemical reaction prediction tools leverage enzymatic promiscuity rules to generate reaction networks containing novel compounds and reactions. The resulting reaction networks can be used for multiple applications such as designing novel biosynthetic pathways and annotating untargeted metabolomics data. It is vital for these tools to provide a robust, user-friendly method to generate networks for a given application. However, existing tools lack the flexibility to easily generate networks that are tailor-fit for a user's application due to lack of exhaustive reaction rules, restriction to pre-computed networks, and difficulty in using the software due to lack of documentation.RESULTS:Here we present Pickaxe, an open-source, flexible software that provides a user-friendly method to generate novel reaction networks. This software iteratively applies reaction rules to a set of metabolites to generate novel reactions. Users can select rules from the prepackaged JN1224min ruleset, derived from MetaCyc, or define their own custom rules. Additionally, filters are provided which allow for the pruning of a network on-the-fly based on compound and reaction properties. The filters include chemical similarity to target molecules, metabolomics, thermodynamics, and reaction feasibility filters. Example applications are given to highlight the capabilities of Pickaxe: the expansion of common biological databases with novel reactions, the generation of industrially useful chemicals from a yeast metabolome database, and the annotation of untargeted metabolomics peaks from an E. coli dataset.CONCLUSION:Pickaxe predicts novel metabolic reactions and compounds, which can be used for a variety of applications. This software is open-source and available as part of the MINE Database python package ( https://pypi.org/project/minedatabase/ ) or on GitHub ( https://github.com/tyo-nu/MINE-Database ). Documentation and examples can be found on Read the Docs ( https://mine-database.readthedocs.io/en/latest/ ). Through its documentation, pre-packaged features, and customizable nature, Pickaxe allows users to generate novel reaction networks tailored to their application.
Cell-free systems are useful tools for prototyping metabolic pathways and optimizing the production of various bioproducts. Mechanistically-based kinetic models are uniquely suited to analyze dynamic experimental data collected from cell-free systems and provide vital qualitative insight. However, to date, dynamic kinetic models have not been applied with rigorous biological constraints or trained on adequate experimental data to the degree that they would give high confidence in predictions and broadly demonstrate the potential for widespread use of such kinetic models. In this work, we construct a large-scale dynamic model of cell-free metabolism with the goal of understanding and optimizing butanol production in a cell-free system. Using a combination of parameterization methods, the resultant model captures experimental metabolite measurements across two experimental conditions for nine metabolites at timepoints between 0 and 24 h. We present analysis of the model predictions, provide recommendations for butanol optimization, and identify the aldehyde/alcohol dehydrogenase as the primary bottleneck in butanol production. Sensitivity analysis further reveals the extent to which various parameters are constrained, and our approach for probing valid parameter ranges can be applied to other modeling efforts.