Base editing enables precise genome modification without double-strand breaks but remains limited by narrow editing windows, DNA repair pathway biases, and restricted nucleotide diversity. Here, we report MUTATOR, a MUlTiplexAble and self-iTerative ORthogonal base-editing platform that enables N-to-N diversification in Escherichia coli. MUTATOR combines CWBE and ABE with iterative editing on two complementary DNA strands, thereby overcoming endogenous DNA repair constraints and expanding A-to-N and C-to-N editing outcomes across both strands. This strategy substantially expands accessible nucleotide outcomes, codon variants, and amino-acid diversity within existing editing windows relative to conventional editors. Using four gRNAs, MUTATOR facilitated four-site editing of ompR, generating 84 distinct amino-acid combinations and 252 codon combinations, with the synonymous OmpR_P160P variant increasing isobutanol production by up to 56.2%. We further applied MUTATOR to a 151-gene library encompassing transcriptional regulators, translation factors, DNA repair proteins, ribosomal components, and NAD(P)H-associated metabolic genes, identifying single and combinatorial mutations that markedly enhanced cell growth and ethanol utilization when ethanol was used as the sole carbon source. Together, these results establish MUTATOR as a broadly applicable platform for genome-wide diversification, functional dissection, and rapid engineering of industrial microbial chassis.
Abstract The strong cytotoxicity of wild-type CcdB limits its use as a counter-selection marker, as leaky expression can cause unintended cell death. Here, we report the rational engineering of a tunable CcdB variant with minimal basal toxicity while retaining inducible lethality. Using a structure-aware deep learning-guided strategy, we prioritized mutations predicted to alter CcdB−DNA gyrase binding energetics and experimentally identified variants with reduced basal toxicity and retained inducible killing. Targeted validation identified CcdB_L96P, which significantly reduces basal toxicity while preserving effective killing upon induction. This variant exhibits minimal leakage and robust inducible lethality across multiple Escherichia coli strains without requiring specialized hosts. It enables antibiotic-free CRISPR gRNA plasmid construction with high accuracy and efficient enrichment of correctly assembled clones. In a succinic acid-producing chassis, CcdB_L96P imposes no detectable metabolic burden and supports effective removal of residual cells using low concentrations of anhydrotetracycline. Together, this work provides a generalizable protein-level strategy for tuning toxin activity in synthetic biology.
Stilbenoids are plant-derived polyphenols with potent bioactivities, which makes them promising for pharmaceutical and nutraceutical use. Their production remains challenging due to their low natural abundance in plants and the inherent toxicity associated with chemical synthesis. Here, we developed an integrated strategy to engineer Escherichia coli for efficient biosynthesis of multiple stilbenoids. A p-coumaric acid-responsive CRISPRi system dynamically regulated malonyl-CoA allocation, balancing host metabolism and stilbenoid production. We systematically explored enzyme co-localization strategies and used structure-informed in silico analysis to support dual-enzyme fusion design. Orthogonal SpyTag/SpyCatcher and SnoopTag/SnoopCatcher systems enabled optimal intermediate channeling and significantly enhanced pathway flux. As a result, we achieved the highest reported titers of piceatannol (494.30 mg/L from glycerol, increased to 583.10 mg/L with l-tyrosine supplementation) and pterostilbene (563.89 mg/L from glycerol, reaching 1110.92 mg/L with l-tyrosine feeding). Techno-economic analysis indicated that combined optimization of metabolic regulation, enzyme co-localization, and process design contributes substantially to the overall feasibility of stilbenoid manufacturing. Altogether, this study demonstrates the potential of stilbenoid biosynthesis and provides a valuable paradigm for the efficient synthesis of other complex natural products.
3-Hydroxypropionic acid is an important malonyl-CoA-derived platform chemical whose efficient biosynthesis is constrained by limited utilization of malonyl-CoA across subcellular compartments. Using a biosensor-guided transcription factor mutagenesis screen combined with transcriptomic and functional validation, we identify four positive mutants NRG1_R224H, STB3_L52Y, PDR1_T820C and PGD1_V243D that increase cytosolic malonyl-CoA accumulation by globally reprogramming transcription to reinforce central carbon flux and acetyl-CoA precursor supply, and remodel amino acid, redox, and lipid metabolism to favor malonyl-CoA accumulation. We further utilize mitochondrial malonyl-CoA for 3-HP production through dynamic control of HFA1 and optimized POS5 expression, and further develop a dual-compartment coordination strategy to efficiently exploit cytosolic and mitochondrial malonyl-CoA pools. Integration of optimized pathways in diploid strains enables coordinated precursor utilization, achieving 81.8 g/L 3-HP in 5-L fed-batch fermentation, the highest titer reported to date in Saccharomyces cerevisiae. This work establishes a generalizable framework for multi-compartment malonyl-CoA utilization in eukaryotic cell factories.
Succinate is a biobased platform chemical with wide applications in food, pharmaceuticals, and biodegradable polymers such as polybutylene succinate. Despite advances in microbial fermentation, cost-effective production remains limited by inefficient utilization of lignocellulosic hydrolysates, where glucose and xylose are the predominant sugars. In this study, we systematically engineered Escherichia coli C600 to enhance succinate biosynthesis from mixed sugars and hydrolysates. Competing by-product pathways were eliminated, the phosphotransferase system was modified to relieve carbon catabolite repression, and the pck gene from Bacillus subtilis was introduced to alleviate the ATP burden in xylose metabolism. To further improve xylose utilization, heterologous oxidative pathways (Weimberg and Dahms) from Caulobacter crescentus were integrated and fine-tuned using ribosome binding site libraries. The optimized strain exhibited flexible glucose-xylose co-utilization across varying sugar ratios, maintaining high succinate yields. A global transcriptional regulator library was then applied, and a crp mutant ESC6crp-W68+ was identified and enabled efficient growth and succinate production using inorganic nitrogen as the sole nitrogen source. Scale-up fermentation in a 5-L bioreactor confirmed the industrial relevance of the engineered strain: ESC6crp-W68+ produced 87.7 g/L succinate from synthetic mixed sugars with a yield of 1.15 mol/mol, and 77.3 g/L from corn stover hydrolysate with a yield of 1.02 mol/mol. This multi-layered engineering framework established a metabolically robust and cost-efficient E. coli platform, enabling high-titer succinate production directly from lignocellulosic hydrolysates.
Prime editing enables precise genome modifications without DNA double-strand breaks, yet bacterial applications are limited by low efficiency and small edit sizes. Here, we develop PE-STAR, Prime Editing with SOS-Triggered and RecJ-Augmented Repair, to enhance prime editing in Escherichia coli. Removing three inhibitory 3'→5' exonucleases (SbcB, ExoX, and XseA) improved edited-strand retention, and extending post-transformation outgrowth increased editing efficiency. RecJ overexpression strengthened 5'-directed processing during flap resolution and gap expansion, biasing repair toward incorporation of the reverse-transcribed edited strand. To enrich edited cells, we integrated an SOS-responsive counter-selection circuit that links PE3-associated dual nicking to LexA-dependent gRNA expression targeting a plasmid encoding the toxin CcdB, thereby eliminating unedited cells. PE-STAR achieved up to 80%-90% editing efficiency for short-fragment modifications, representing up to 16-fold improvement across loci. The platform supported insertions, deletions, and replacements of up to 46 bp with high efficiency. Furthermore, installing an attB site by prime editing, followed by Bxb1 integrase recombination, enabled chromosomal integration of 3.2 and 8.0 kb cassettes with 100% recombination efficiency among screened colonies, including GFP reporter and riboflavin biosynthetic pathway. PE-STAR expands both the efficiency and functional scope of bacterial prime editing for programmable genome engineering.
Protein-nucleic acid interactions play central roles in gene regulation and cellular function, and extensive efforts have been devoted to predicting nucleic acid binding sites from protein structures. However, protein-nucleic acid recognition is inherently dynamic, whereas most existing computational approaches rely on single static conformations, limiting their ability to capture conformational heterogeneity underlying binding. Here, we present DyProL, an ensemble-based conformational representation learning framework that models proteins as ensembles of conformations sampled from equilibrium-like structural distributions. DyProL learns dynamic structural features through iterative aggregation of intra- and inter-conformation geometric information, enabling representation of both local structural context and global conformational variability. Across multiple benchmarks, DyProL consistently outperforms state-of-the-art methods in nucleic acid binding site prediction, with particularly pronounced improvements under realistic settings using predicted or apo-like structures, where static methods degrade substantially. These results establish dynamic ensemble-based representations as a general and scalable paradigm for structure-based protein modeling, providing a foundation for improving a broad range of protein function prediction tasks.
CRISPR (clustered regularly interspaced short palindromic repeats)-Cas (CRISPR-associated protein) nucleases enable precise genome editing, but off-target cleavage remains a critical challenge. Here, we report the development of MAD7_HF, a high-fidelity variant of the MAD7 nuclease engineered through a bacterial screening system leveraging the DNA gyrase-targeting toxic gene ccdB. This system couples survival to efficient on-target cleavage and minimal off-target activity, mimicking the transient action required for high-precision editing. Through iterative selection and sequencing validation, we identified MAD7_HF, harboring three substitutions (R187C, S350T, K1019N) that enhanced discrimination between on- and off-target sites. In Escherichia coli assays, MAD7_HF exhibited a >20-fold reduction in off-target cleavage across multiple mismatch contexts while maintaining on-target efficiency comparable to wild-type MAD7. Structural modeling revealed that these mutations stabilize the guide RNA-DNA hybrid at on-target sites and weaken interactions with mismatched sequences. This work establishes a high-throughput bacterial screening strategy that allows the identification of Cas12a variants with improved specificity at a given target site, providing a useful framework for future efforts to develop precision genome-editing tools.
CRISPR-based methods enable genome modifications for diverse applications but often face challenges, such as inconsistent efficiencies, reduced performance in iterative modifications, and difficulties generating high-quality datasets for high-throughput genome engineering. Here, we present SELECT (SOS Enhanced programmabLE CRISPR-Cas ediTing), a novel strategy integrating the CRISPR–Cas system with the DNA damage response. By employing designed and optimized double-strand break induced promoters that are activated upon genome editing, SELECT enables a counter-selection process to eliminate unedited cells, ensuring high-fidelity editing. This approach achieves up to 100% efficiency for point mutations, iterative knockouts, and insertions. In high-throughput library editing, SELECT achieved up to 94.2% efficiency and preserved higher library diversity compared with conventional methods. Application of SELECT in flaviolin biosynthesis resulted in a 3.97-fold increase in production. Furthermore, integration with machine learning tools allowed rapid mapping of genotype–phenotype relationships. SELECT provides a versatile platform for precision genome engineering in Escherichia coli and Saccharomyces cerevisiae.
The CRISPR-Cas9 system has been widely applied for industrial microbiology but is not effective in certain microorganisms. This forum explores the strategies aimed at overcoming these challenges, including the use of the Cas12a system, Cas9 variants, and non-CRISPR techniques, to provide more effective strategies for expanding applications in microbial engineering.
The CRISPR/Cas systems comprising the clustered regularly interspaced short palindromic repeats(CRISPR)and its associated Cas protein is an acquired immune system unique to archaea or bacteria.Since its development as a gene editing tool,it has rapidly become a popular research direction in the field of synthetic biology due to its advantages of high efficiency,precision,and versatility.This technique has since revolutionized the research of many fields including life sciences,bioengineering technology,food science,and crop breeding.Currently,the single gene editing and regulation techniques based on CRISPR/Cas systems have been increasingly improved,but challenges still exist in the multiplex gene editing and regulation.This review focuses on the development and application of multiplex gene editing and regulation techniques based on the CRISPR/Cas systems,and summarizes the techniques for multiplex gene editing or regulation within a single cell or within a cell population.This includes the multiplex gene editing techniques developed based on the CRISPR/Cas systems with double-strand breaks;or with single-strand breaks;or with multiple gene regulation techniques,etc.These works have enriched the tools for the multiplex gene editing and regulation and contributed to the application of CRISPR/Cas systems in the multiple fields.
With different types of nucleases, genome editing technologies have opened up the possibility for targeting and modifying specific gene sequences, which show potential applications in basic and applied aspects of biotechnology research. Zinc-finger nucleases(ZFNs) and transcription activator-like effector nucleases(TALENs) are artificial proteins generated by fusing a specific DNA-binding domain with a restriction enzyme FokI DNA-cleavage domain, which arise from their ability to customize the DNA-binding domain for recognition of targeting sequences and cleaving them by the FokI domain. However, the design and construction of such a system are time consuming,laborious and costly. Clustered regularly interspaced short palindromic repeats(CRISPR)/CRISPR associated(Cas)protein is a unique acquired immune system of bacteria or archaea. Since researchers constructed the CRISPR/Cas system for gene editing, its high efficiency has revolutionized a variety of fields such as life sciences, bioengineering, biomedicine, food, and agricultural sciences. However, the CRISPR/Cas system still has some challenges, such as offtarget effect and limited PAM site recognition range, which limit its further applications. In order to solve these problems, molecular engineering of Cas proteins has become an important strategy for developing and optimizing CRISPR/Cas systems. In this study, with CRISPR/Cas9 and CRISPR/Cas12a selected as representative examples for DNA-targeting Class II CRISPR/Cas systems, we focus on the optimization and modification methods, and progress of Cas9 and Cas12a proteins achieved within recent years, such as Cas protein engineering for improved on-target specificity and expanded PAM scopes, developing new functions using CRISPR/Cas systems as gene targeting tools, and introducing exogenous protein domains to regulate CRISPR/Cas functions. These studies have generated a series of high-specificity and high-precision CRISPR/Cas systems, which have greatly expanded their functions and scopes, and made important contributions to the wide-range applications of CRISPR/Cas systems.
Protein engineering has been used successfully in fields ranging from medicine to food science to biofuels. Applications of protein engineering include developing antiviral peptides or other protein therapeutics, antibody engineering, designing protein-based logic circuits, engineering enzymes to be more specific or to function under industrially relevant conditions such as at higher temperatures or high/low pH, modifying cell signaling or regulatory functions, and so on. Advances in recombinant DNA, "omics," and CRISPR-Cas (clustered regularly interspaced short palindromic repeats and its associated proteins) technologies, combined with high-throughput screening facilities, will lead to improved methods for protein engineering, enabling easy modification of more proteins/enzymes for new specific applications. New methods for rational design, directed evolution, and computer-aided protein design will further accelerate the speed of protein evolution and expand the scope for protein engineering. In this chapter we discuss general protein engineering strategies and advances in engineering proteins with desired functions, focusing on the "design" and "build" part of the design-build-test-learn cycle.
The development of microbial chassis for the production of a variety of biochemicals and biofuels is a growing area of research. How to efficiently manipulate genetic information to achieve optimal production of these compounds is a key area of focus in the field. In recent years, clustered regularly interspaced palindromic repeats (CRISPR) and its associated proteins (Cas) have become a popular strategy for gene editing and regulation in many organisms due to its versatility and efficacy. Here, we describe methods developed utilizing CRISPR-Cas systems for engineering microbial cell factories (e.g., bacteria and yeast) at the single gene to genome scale for mutagenesis and transcriptional regulation of target genes. Finally, we provide a perspective on the challenges and opportunities for the applications of advanced CRISPR-Cas-based tools for engineering microbial cell factories.
Alcohol toxicity significantly impacts the titer and productivity of industrially produced biofuels. To overcome this limitation, we must find and use strategies to improve stress tolerance in production strains. Previously, we developed a multiplex navigation of a global regulatory network (MINR) library that targeted 25 regulatory genes that are predicted to modify global regulation in yeast under different stress conditions. In this study, we expanded this concept to target the active sites of 47 transcriptional regulators using a saturation mutagenesis library. The 47 targeted regulators interact with more than half of all yeast genes. We then screened and selected for C3-C4 alcohol tolerance. We identified specific mutants that have resistance to isopropanol and isobutanol. Notably, the WAR1_K110N variant improved tolerance to both isopropanol and isobutanol. In addition, we investigated the mechanisms for improvement of isopropanol and isobutanol stress tolerance and found that genes related to glycolysis play a role in tolerance to isobutanol, while changes in ATP synthesis and mitochondrial respiration play a role in tolerance to both isobutanol and isopropanol. Overall, this work sheds light on basic mechanisms for isopropanol and isobutanol toxicity and demonstrates a promising strategy to improve tolerance to C3-C4 alcohols by perturbing the transcriptional regulatory network.
CRISPR technology is a universal tool for genome engineering that has revolutionized biotechnology. Recently identified unique CRISPR/Cas systems, as well as re-engineered Cas proteins, have rapidly expanded the functions and applications of CRISPR/Cas systems. The structures of Cas proteins are complex, containing multiple functional domains. These protein domains are evolutionarily conserved polypeptide units that generally show independent structural or functional properties. In this review, we propose using protein domains as a new way to classify protein engineering strategies for these proteins and discuss common ways to engineer key domains to modify the functions of CRISPR/Cas systems.
Regulatory networks describe the hierarchical relationship between transcription factors, associated proteins, and their target genes. Regulatory networks respond to environmental and genetic perturbations by reprogramming cellular metabolism. Here we design, construct, and map a comprehensive regulatory network library containing 110,120 specific mutations in 82 regulators expected to perturb metabolism. We screen the library for different targeted phenotypes, and identify mutants that confer strong resistance to various inhibitors, and/or enhanced production of target compounds. These improvements are identified in a single round of selection, showing that the regulatory network library is universally applicable and is convenient and effective for engineering targeted phenotypes. The facile construction and mapping of the regulatory network library provides a path for developing a more detailed understanding of global regulation in E. coli , with potential for adaptation and use in less-understood organisms, expanding toolkits for future strain engineering, synthetic biology, and broader efforts.