The precise recognition of specific peptide-major histocompatibility complex (pMHC) complexes by T cell receptors (TCRs) plays a key role in infectious disease, cancer, and autoimmunity. A critical step in many immunobiological studies is the identification of T cells expressing TCRs specific to a given pMHC antigen. However, the intrinsic instability of empty class-I MHCs limits their soluble expression in Escherichia coli and makes it very difficult to characterize even a small fraction of possible pMHC/TCR interactions. To overcome this limitation, we designed small proteins which buttress the peptide binding groove of class I MHCs, replacing β2-microglobulin (β2m) and the heavy chain α3 domain, and enable soluble and partially soluble expression in E. coli of H-2Db and A*02:01, respectively. We demonstrate that these soluble, monomeric, antigen-receptive, truncated (SMART) MHCs retain both peptide- and TCR-binding specificity and that peptide-bound structures of both allomorphs are similar to their full-length, native counterparts. With extension to the majority of HLA alleles, SMART MHCs should be broadly useful for probing the T cell repertoire in approaches ranging from yeast display to T cell staining.
Enzyme engineering is limited by the challenge of rapidly generating and using large datasets of sequence-function relationships for predictive design. To address this challenge, we develop a machine learning (ML)-guided platform that integrates cell-free DNA assembly, cell-free gene expression, and functional assays to rapidly map fitness landscapes across protein sequence space and optimize enzymes for multiple, distinct chemical reactions. We apply this platform to engineer amide synthetases by evaluating substrate preference for 1217 enzyme variants in 10,953 unique reactions. We use these data to build augmented ridge regression ML models for predicting amide synthetase variants capable of making 9 small molecule pharmaceuticals. Over these nine compounds, ML-predicted enzyme variants demonstrate 1.6- to 42-fold improved activity relative to the parent. Our ML-guided, cell-free framework promises to accelerate enzyme engineering by enabling iterative exploration of protein sequence space to build specialized biocatalysts in parallel. While machine learning shows promise in expanding protein engineering efforts, its potential is limited by the challenge of gathering large datasets of sequence-function relationships. Here, authors introduce a platform that integrates cell-free DNA assembly and gene expression to accelerate enzyme engineering.
Post-translational modifications (PTMs) are important for the stability and function of many therapeutic proteins and peptides. Current methods for studying and engineering PTMs are often limited by low-throughput experimental techniques. Here we describe a generalizable, in vitro workflow coupling cell-free gene expression (CFE) with AlphaLISA for the rapid expression and testing of PTM installing proteins. We apply our workflow to two representative classes of peptide and protein therapeutics: ribosomally synthesized and post-translationally modified peptides (RiPPs) and glycoproteins. First, we demonstrate how our workflow can be used to characterize the binding activity of RiPP recognition elements, an important first step in RiPP biosynthesis, and be integrated into a biodiscovery pipeline for computationally predicted RiPP products. Then, we adapt our workflow to study and engineer oligosaccharyltransferases (OSTs) involved in protein glycan coupling technology, leading to the identification of mutant OSTs and sites within a model vaccine carrier protein that enable high efficiency production of glycosylated proteins. We expect that our workflow will accelerate design-build-test-learn cycles for engineering PTMs.
The production of N-linked glycoproteins in genetically tractable bacterial hosts and their cell-free extracts holds great promise for low-cost, customizable, and distributed biomanufacturing of glycoconjugate vaccines and glycoprotein therapeutics. In nearly all bacterial N-linked protein glycosylation systems described so far, a single-subunit, transmembrane oligosaccharyltransferase (OST) is employed which favors acceptor sites in flexible, solvent-exposed motifs of the glycoprotein substrate. Yet despite this preference, acceptor sites in structured domains can also be glycosylated in living bacteria, presumably by a mechanism where the site is presented to the OST in a flexible form during or after the membrane translocation step but prior to folding being completed. While N-glycoprotein biosynthesis can also be accomplished using cell-free extracts derived from glycosylation-competent bacteria, it remains to be determined whether the cell-free reaction environment involves a similar mechanism for glycosylation of structured domains. Using an Escherichia coli-based cell-free glycoprotein synthesis (CFGpS) system, we observed efficient glycosylation of two eukaryotic glycoproteins, namely ribonuclease A (RNase A) and the fragment crystallizable (Fc) region of human immunoglobulin G (IgG), whose acceptor sites occur in structurally constrained regions that were not glycosylated when the proteins were already folded. Because this cell-free glycosylation depended on ribosomal translation but not on signal peptide-mediated translocation, we propose the existence of a unique co-translational, but not co-translocational, glycosylation mechanism in CFGpS. Collectively, these findings reveal the potential for CFGpS to become a viable platform for producing complex eukaryotic glycoprotein targets.
Many industrial chemicals that are produced from fossil resources could be manufactured more sustainably through fermentation. Here we describe the development of a carbon-negative fermentation route to producing the industrially important chemicals acetone and isopropanol from abundant, low-cost waste gas feedstocks, such as industrial emissions and syngas. Using a combinatorial pathway library approach, we first mined a historical industrial strain collection for superior enzymes that we used to engineer the autotrophic acetogen Clostridium autoethanogenum . Next, we used omics analysis, kinetic modeling and cell-free prototyping to optimize flux. Finally, we scaled-up our optimized strains for continuous production at rates of up to ~3 g/L/h and ~90% selectivity. Life cycle analysis confirmed a negative carbon footprint for the products. Unlike traditional production processes, which result in release of greenhouse gases, our process fixes carbon. These results show that engineered acetogens enable sustainable, high-efficiency, high-selectivity chemicals production. We expect that our approach can be readily adapted to a wide range of commodity chemicals.
The SARS-CoV-2 pandemic highlighted the urgent need for biomanufacturing paradigms that are robust and fast. Here, we demonstrate the rapid process development and scalable cell-free production of T7 RNA polymerase, a critical component in mRNA vaccine synthesis. We carry out a 1-L cell-free gene expression (CFE) reaction that achieves over 90% purity, low endotoxin levels, and enhanced activity relative to commercial T7 RNA polymerase. To achieve this demonstration, we implement rolling circle amplification to circumvent difficulties in DNA template generation, and tune cell-free reaction conditions, such as temperature, additives, purification tags, and agitation, to boost yields. We achieve production of a similar quality and titer of T7 RNA polymerase over more than four orders of magnitude in reaction volume. This proof of principle positions CFE as a viable solution for decentralized biotherapeutic manufacturing, enhancing preparedness for future public health crises or emergent threats.
Brochosomes are proteinaceous nanostructures produced by leafhopper insects with superhydrophobic and antireflective properties. Unfortunately, the production and study of brochosome-based materials has been limited by poor understanding of their major constituent subunit proteins, known as brochosomins, as well as their sensitivity to redox conditions due to essential disulfide bonds. Here, we used cell-free gene expression (CFE) to achieve recombinant production and analysis of brochosomin proteins. Through the optimization of redox environment, reaction temperature, and disulfide bond isomerase concentration, we achieved soluble brochosomin yields of up to 341 ± 30 μg/mL. Analysis using dynamic light scattering and transmission electron microscopy revealed distinct aggregation patterns among cell-free mixtures with different expressed brochosomins. We anticipate that the CFE methods developed here will accelerate the ability to change the geometries and properties of natural and modified brochosomes, as well as facilitate the expression and structural analysis of other poorly understood protein complexes.
Human immunoglobulin G (IgG) antibodies are a major class of biotherapeutics and undergo N-linked glycosylation in their Fc domain, which is critical for immune functions and therapeutic activity. Hence, technologies for producing authentically glycosylated IgGs are in high demand. Previous attempts to engineer Escherichia coli for this purpose have met limited success due in part to the lack of oligosaccharyltransferase (OST) enzymes that can install N-glycans at the conserved N297 site in the Fc region. Here, we identify a single-subunit OST from Desulfovibrio marinus with relaxed substrate specificity that catalyzes glycosylation of native Fc acceptor sites. By chemoenzymatic remodeling the attached bacterial glycans to homogeneous, asialo complex-type G2 N-glycans, the E. coli-derived Fc binds human FcγRIIIa/CD16a, a key receptor for antibody-dependent cellular cytotoxicity (ADCC). Overall, the discovery of D. marinus OST provides previously unavailable biocatalytic capabilities and sets the stage for using E. coli to produce fully human antibodies.
As part of the first anniversary issue of Nature Chemical Engineering, we present a collection of opinions from 40 researchers within the field on what they think are the most exciting opportunities that lie ahead for their respective topics.
Industrialization and failing infrastructure have led to a growing number of irreversible health conditions resulting from chronic lead exposure. While state-of-the-art analytical chemistry methods provide accurate and sensitive detection of lead, they are too slow, expensive, and centralized to be accessible to many. Cell-free biosensors based on allosteric transcription factors (aTFs) can address the need for accessible, on-demand lead detection at the point of use. However, known aTFs, such as PbrR, are unable to detect lead at concentrations regulated by the Environmental Protection Agency (24-72 nM). Here, we develop a rapid cell-free platform for engineering aTF biosensors with improved sensitivity, selectivity, and dynamic range characteristics. We apply this platform to engineer PbrR mutants for a shift in limit of detection from 10 μM to 50 nM lead and demonstrate use of PbrR as a cell-free biosensor. We envision that our workflow could be applied to engineer any aTF.
Climate change poses a significant threat to global agriculture, necessitating innovative solutions. Plant synthetic biology, particularly chloroplast engineering, holds promise as a viable approach to this challenge. Chloroplasts present a variety of advantageous traits for genetic engineering, but the development of genetic tools and genetic part characterization in these organelles is hindered by the lengthy time scales required to generate transplastomic organisms. To address these challenges, we have established a versatile protocol for generating highly active chloroplast-based cell-free gene expression (CFE) systems derived from a diverse range of plant species, including wheat (monocot), spinach, and poplar trees (dicots). We show that these systems work with conventionally used T7 RNA polymerase as well as the endogenous chloroplast polymerases, allowing for detailed characterization and prototyping of regulatory sequences at both transcription and translation levels. To demonstrate the platform for characterization of promoters and 5' and 3' untranslated regions (UTRs) in higher plant chloroplast gene expression, we analyze a collection of 23 5'UTRs, 10 3'UTRs, and 6 chloroplast promoters, assessed their expression in spinach and wheat extracts, and found consistency in expression patterns, suggesting cross-species compatibility. Looking forward, our chloroplast CFE systems open new avenues for plant synthetic biology, offering prototyping tools for both understanding gene expression and developing engineered plants, which could help meet the demands of a changing global climate.
Cell-free protein synthesis (CFPS), whereby cell lysates are used to produce proteins from a genetic template, has matured as an attractive alternative to standard biomanufacturing modalities due to its high volumetric productivity contained within a distributable platform. Initially, cell-free lysates produced from Escherichia coli, which are both simple to produce and cost-effective for the production of a wide variety of proteins, were unable to produce glycosylated proteins as E. coli lacks native glycosylation machinery. With many important therapeutic proteins possessing asparagine-linked glycans that are critical for structure and function, this gap in CFPS production capabilities was addressed with the development of cell-free expression of glycoproteins (glycoCFE), which uses the supplementation of extracted lipid-linked oligosaccharides and purified oligosaccharyltransferases to enable glycoprotein production in the CFPS reaction environment. In this chapter, we highlight the basic methods for the preparation of reagents for glycoCFE and the protocol for expression and glycosylation of a model protein using a more productive, yet simplified, glycoCFE setup. Beyond this initial protocol, we also highlight how this protocol can be extended to a wide range of alternative glycan structures, oligosaccharyltransferases, and acceptor proteins as well as to a one-pot cell-free glycoprotein synthesis reaction.
Recent years have seen intense interest in the development of point-of-care nucleic acid diagnostic technologies to address the scaling limitations of laboratory-based approaches. Chief among these are combinations of isothermal amplification approaches with CRISPR-based detection and readouts of target products. Here, we contribute to the growing body of rapid, programmable point-of-care pathogen tests by developing and optimizing a one-pot NASBA-Cas13a nucleic acid detection assay. This test uses the isothermal amplification technique NASBA to amplify target viral nucleic acids, followed by the Cas13a-based detection of amplified sequences. We first demonstrate an in-house formulation of NASBA that enables the optimization of individual NASBA components. We then present design rules for NASBA primer sets and LbuCas13a guide RNAs for the fast and sensitive detection of SARS-CoV-2 viral RNA fragments, resulting in 20-200 aM sensitivity. Finally, we explore the combination of high-throughput assay condition screening with mechanistic ordinary differential equation modeling of the reaction scheme to gain a deeper understanding of the NASBA-Cas13a system. This work presents a framework for developing a mechanistic understanding of reaction performance and optimization that uses both experiments and modeling, which we anticipate will be useful in developing future nucleic acid detection technologies.
The important roles that protein glycosylation plays in modulating the activities and efficacies of protein therapeutics have motivated the development of synthetic glycosylation systems in living bacteria and in vitro. A key challenge is the lack of glycosyltransferases that can efficiently and site-specifically glycosylate desired target proteins without the need to alter primary amino acid sequences at the acceptor site. Here, we report an efficient and systematic method to screen a library of glycosyltransferases capable of modifying comprehensive sets of acceptor peptide sequences in parallel. This approach is enabled by cell-free protein synthesis and mass spectrometry of self-assembled monolayers and is used to engineer a recently discovered prokaryotic N-glycosyltransferase (NGT). We screened 26 pools of site-saturated NGT libraries to identify relevant residues that determine polypeptide specificity and then characterized 122 NGT mutants, using 1052 unique peptides and 52,894 unique reaction conditions. We define a panel of 14 NGTs that can modify 93% of all sequences within the canonical X-1-N-X+1-S/T eukaryotic glycosylation sequences as well as another panel for many noncanonical sequences (with 10 of 17 non-S/T amino acids at the X+2 position). We then successfully applied our panel of NGTs to increase the efficiency of glycosylation for three protein therapeutics. Our work promises to significantly expand the substrates amenable to in vitro and bacterial glycoengineering.
Plastid engineering offers the potential to carry multi-gene traits in plants, however, it requires reliable genetic parts to balance expression. The difficulty of chloroplast transformation and slow plant growth make it challenging to build plants just to characterize genetic parts. To address these limitations, we developed a cell-free system from Nicotiana tabacum chloroplast extracts for prototyping genetic parts. Our cell-free system uses combined transcription and translation driven by T7 RNA polymerase and works with plasmid or linear template DNA. To develop our system, we optimized lysis, extract preparation procedures (e.g., runoff reaction, centrifugation, and dialysis), and the physiochemical reaction conditions. Our cell-free system can synthesize 34 +/- 1 micrograms per milliliter luciferase in batch reactions. We apply our system to test a library of 104 ribosome binding site (RBS) variants and rank them based on cell-free gene expression. We observe a 1300-fold range of luciferase expression normalized by mRNA expression, as assessed by the malachite green aptamer (relative luminescence units per relative fluorescence units). We also find a positive correlation between the observed expression in chloroplast extracts and the predictions made by the RBS calculator. We anticipate that chloroplast cell-free systems will increase the speed and reliability of building genetic programs in plant chloroplasts for diverse applications.
The de novo construction of a living organism is a compelling vision. Despite the astonishing technologies developed to modify living cells, building a functioning cell "from scratch" has yet to be accomplished. The pursuit of this goal alone has─and will─yield scientific insights affecting fields as diverse as cell biology, biotechnology, medicine, and astrobiology. Multiple approaches have aimed to create biochemical systems manifesting common characteristics of life, such as compartmentalization, metabolism, and replication and the derived features, evolution, responsiveness to stimuli, and directed movement. Significant achievements in synthesizing each of these criteria have been made, individually and in limited combinations. Here, we review these efforts, distinguish different approaches, and highlight bottlenecks in the current research. We look ahead at what work remains to be accomplished and propose a "roadmap" with key milestones to achieve the vision of building cells from molecular parts.
Cell-free protein synthesis (CFPS) is a rapidly maturing in vitro gene expression platform that can be used to transcribe and translate nucleic acids at the point of need, enabling on-demand synthesis of peptide-based vaccines and biotherapeutics as well as the development of diagnostic tests for environmental contaminants and infectious agents. Unlike traditional cell-based systems, CFPS platforms do not require the maintenance of living cells and can be deployed with minimal equipment; therefore, they hold promise for applications in low-resource contexts, including spaceflight. Here, we evaluate the performance of the cell-free platform BioBits aboard the International Space Station by expressing RNA-based aptamers and fluorescent proteins that can serve as biological indicators. We validate two classes of biological sensors that detect either the small-molecule DFHBI or a specific RNA sequence. Upon detection of their respective analytes, both biological sensors produce fluorescent readouts that are visually confirmed using a hand-held fluorescence viewer and imaged for quantitative analysis. Our findings provide insights into the kinetics of cell-free transcription and translation in a microgravity environment and reveal that both biosensors perform robustly in space. Our findings lay the groundwork for portable, low-cost applications ranging from point-of-care health monitoring to on-demand detection of environmental hazards in low-resource communities both on Earth and beyond.
The genus Clostridium is a large and diverse group within the Bacillota (formerly Firmicutes), whose members can encode useful complex traits such as solvent production, gas-fermentation, and lignocellulose breakdown. We describe 270 genome sequences of solventogenic clostridia from a comprehensive industrial strain collection assembled by Professor David Jones that includes 194 C. beijerinckii, 57 C. saccharobutylicum, 4 C. saccharoperbutylacetonicum, 5 C. butyricum, 7 C. acetobutylicum, and 3 C. tetanomorphum genomes. We report methods, analyses and characterization for phylogeny, key attributes, core biosynthetic genes, secondary metabolites, plasmids, prophage/CRISPR diversity, cellulosomes and quorum sensing for the 6 species. The expanded genomic data described here will facilitate engineering of solvent-producing clostridia as well as non-model microorganisms with innately desirable traits. Sequences could be applied in conventional platform biocatalysts such as yeast or Escherichia coli for enhanced chemical production. Recently, gene sequences from this collection were used to engineer Clostridium autoethanogenum, a gas-fermenting autotrophic acetogen, for continuous acetone or isopropanol production, as well as butanol, butanoic acid, hexanol and hexanoic acid production.
Plastid engineering offers the potential to carry multigene traits in plants; however, it requires reliable genetic parts to balance expression. The difficulty of chloroplast transformation and slow plant growth makes it challenging to build plants just to characterize genetic parts. To address these limitations, we developed a high-yield cell-free system from Nicotiana tabacum chloroplast extracts for prototyping genetic parts. Our cell-free system uses combined transcription and translation driven by T7 RNA polymerase and works with plasmid or linear template DNA. To develop our system, we optimized lysis, extract preparation procedures (e.g., runoff reaction, centrifugation, and dialysis), and the physiochemical reaction conditions. Our cell-free system can synthesize 34 ± 1 μg/mL luciferase in batch reactions and 60 ± 4 μg/mL in semicontinuous reactions. We apply our batch reaction system to test a library of 103 ribosome binding site (RBS) variants and rank them based on cell-free gene expression. We observe a 1300-fold dynamic range of luciferase expression when normalized by maximum mRNA expression, as assessed by the malachite green aptamer. We also find that the observed normalized gene expression in chloroplast extracts and the predictions made by the RBS Calculator are correlated. We anticipate that chloroplast cell-free systems will increase the speed and reliability of building genetic systems in plant chloroplasts for diverse applications.
The design and optimization of metabolic pathways, genetic systems, and engineered proteins rely on high-throughput assays to streamline design-build-test-learn cycles. However, assay development is a time-consuming and laborious process. Here, we create a generalizable approach for the tailored optimization of automated cell-free gene expression (CFE)-based workflows, which offers distinct advantages over in vivo assays in reaction flexibility, control, and time to data. Centered around designing highly accurate and precise transfers on the Echo Acoustic Liquid Handler, we introduce pilot assays and validation strategies for each stage of protocol development. We then demonstrate the efficacy of our platform by engineering transcription factor-based biosensors. As a model, we rapidly generate and assay libraries of 127 MerR and 134 CadR transcription factor variants in 3682 unique CFE reactions in less than 48 h to improve limit of detection, selectivity, and dynamic range for mercury and cadmium detection. This was achieved by assessing a panel of ligand conditions for sensitivity (to 0.1, 1, 10 μM Hg and 0, 1, 10, 100 μM Cd for MerR and CadR, respectively) and selectivity (against Ag, As, Cd, Co, Cu, Hg, Ni, Pb, and Zn). We anticipate that our Echo-based, cell-free approach can be used to accelerate multiple design workflows in synthetic biology.