
Abstract The rapid growth of genomic and metagenomic data has created a large gap between protein sequence discovery and enzyme functional characterization. Deep learning has become an important tool for enzyme-related prediction tasks, including EC and GO annotation, substrate specificity, catalytic-residue identification, kinetic-parameter estimation, thermostability, pH optima, and candidate discovery. This review summarizes recent progress with an emphasis on benchmark design rather than reported scores alone. We compare representative methods by input modality, training data, split strategy, leakage control, evaluation metric, and validation evidence. Across tasks, apparent performance gains can arise from dataset scale, annotation quality, homolog overlap, negative sampling, or endpoint definition rather than architecture alone. Current models are valuable for annotation and candidate prioritization, but robust discovery of remote homologs, rare functions, or new catalytic activities still requires low-identity or out-of-distribution evaluation and experimental validation.
Abstract Pichia pastoris is a widely used host for recombinant protein production because it combines the advantages of microbial cultivation with eukaryotic protein folding and secretion. However, secretion efficiency is often limited by the folding capacity of the endoplasmic reticulum (ER), where recombinant proteins must be translocated, folded, and processed prior to export. When ER folding capacity is exceeded, proteins may be retained, degraded, or secreted in non-native conformations, reducing both yield and product quality. Chaperone engineering and codon optimization represent two promising strategies to address these limitations. Here, we generated stable Pichia strains expressing four model secreted proteins (human serum albumin, interleukin-2, thaumatin-I, and thaumatin-II) using either conventional codon optimization or Epi-MAX codon engineering, which adapts transgene codon usage to stress-responsive translational programs. We also engineered strains containing an additional chromosomal copy of either the ER Hsp70 chaperone Kar2 or protein disulfide isomerase (Pdi1). To assess protein quality, we applied limited proteolysis mass spectrometry (LiP-MS), a structural proteomics approach that can detect subtle conformational differences to secreted proteins. Increased Pdi1 levels improved secretion of all four proteins tested, whereas Kar2 overexpression generally reduced yield. For thaumatin-II, Pdi1 enhanced secretion but promoted release of a non-native conformation, which we could correct through codon engineering. Together, these results demonstrate that maximizing recombinant protein production requires optimization of both yield and structural quality and establish complementary strategies for improving secreted protein expression in Pichia.
Abstract Improving the reconstituted translation system is a key requirement for bottom-up synthetic biology. Here, we developed a two-step in vitro evolutionary method that can be used for improving translational proteins. In this method, two distinct conditions were sequentially applied while maintaining genotype-phenotype linkage in water-in-oil droplets. Using this method, we performed in vitro evolution of four translation factors, IleRS, PheRS, EF-G, and EF-Tu, and identified mutations that modestly enhanced translation activity in in vitro expression assays. One of the EF-G mutations (P610S) increased activity per protein approximately 2-fold for the recombinant protein purified from E. coli. This selection method is useful for improving translational proteins for bottom-up synthetic biology.
Abstract The interleukin-1 receptor antagonist (IL-1Ra) is a key anti-inflammatory cytokine that inhibits interleukin-1 signaling by competitively binding to IL-1 receptors. However, recombinant IL-1Ra, currently produced in Escherichia coli, exhibits limited pharmacokinetics and requires frequent administration due to its short plasma half-life and absence of glycosylation. To overcome these limitations, we designed a novel IL-1Ra-Fc fusion protein consisting of human IL-1Ra fused to a human IgG1 Fc fragment via a flexible linker and transiently expressed in Nicotiana benthamiana to enhance protein yield, stability, and enable glycosylation. The recombinant protein accumulated to approximately 340 mg/kg fresh leaf biomass and achieved an overall DSP recovery of 45% protein yield, with a characteristic high α-helical secondary structure, as confirmed by CD spectroscopy. Receptor-binding analysis shows that rhIL-1Ra-Fc exhibited significantly enhanced IL-1 receptor binding activity, showing approximately 17% higher activity than that of commercial IL-1Ra. Overall, this is the first study to demonstrate the feasibility of producing functionally active IL-1Ra-Fc in plants for potential therapeutic applications in inflammatory diseases.
Abstract Noncanonical amino acid (ncAA) incorporation at the protein N-terminus enables defined protein functionalization. However, initiation-based ncAA incorporation systems generally suppress the native methionine pathway, limiting their use for proteins containing internal methionine residues. Here, we developed an orthogonal initiation system for N-terminal ncAA incorporation in cell-free translation systems while retaining native methionine incorporation. We profiled background initiation from all 64 codons in cell-free translation systems and identified low-background artificial initiation codons. Engineered initiator tRNAs were designed to introduce ncAA at the protein N-terminus using artificial initiation codons. This system enabled efficient incorporation of N-biotinyl-l-phenylalanine into protein, reaching over 90% incorporation. This system was further extended to p-azido-l-phenylalanine and to an Escherichia coli extract-based cell-free translation system. Finally, N-biotinylated proteins synthesized in this system were utilized for purification-free biolayer interferometry analysis. This work establishes an orthogonal initiation strategy for N-terminal protein functionalization while preserving the native methionine incorporation.
Abstract l-arabinose, a valuable C5 pentose sugar in xylose mother liquor (a low-cost cellulose hydrolysis byproduct), remains underutilized due to costly separation requirements, resulting in significant waste of fermentable carbon resources. In this study, Escherichia coli W3110 was systematically engineered to efficiently convert l-arabinose into xylitol, a low-calorie sweetener with significant commercial value. Firstly, the l-arabinose metabolic network was reconstructed and glucose catabolic repression was alleviated through coordinated pathway modifications, enabling simultaneous utilization of both l-arabinose and glucose. Subsequently, a xylitol synthesis module consisting of l-arabinose isomerase (AraA), l-xylulose reductase (LXR), and d-psicose-3-epimerase (DPE) was systematically optimized via gene arrangement, promoter and RBS engineering. The optimized pathway was integrated into the E. coli W3110 genome at the IS5 locus using MUCICAT technology, generating a plasmid-free production strain and reducing plasmid-segregation concerns. Fed-batch fermentation in a 3 L bioreactor yielded 64.07 g/L xylitol at a productivity of 1.46 g/L/h with 90.77% l-arabinose conversion in 44 h. This achievement overcomes two critical metabolic bottlenecks: (1) glucose catabolite repression, which normally prevents pentose utilization in the presence of glucose, and (2) the successful stoichiometric balancing of three enzymatic steps (AraA, LXR, DPE). The engineered strain achieves simultaneous glucose-arabinose co-metabolism, and glucose co-utilization supports xylitol formation in a manner consistent with an endogenous reducing-power contribution, thereby eliminating the requirement for exogenous glycerol supplementation in the optimized process. This work establishes a defined-substrate engineering platform for l-arabinose-to-xylitol conversion and provides a strategic basis for future evaluation using complex industrial carbohydrate streams.
Abstract l-Ornithine is a versatile amino acid with broad applications in the pharmaceutical, nutraceutical, and food industries. However, its efficient microbial production is constrained by a critical trade-off: excessive blockage of the downstream l-arginine pathway severely impairs cell growth, while intracellular l-ornithine accumulation causes metabolic stress. Here, we report an integrated metabolic engineering strategy in Escherichia coli that resolves these challenges through three key innovations. First, instead of complete pathway disruption, a fine-tuned attenuation of l-arginine biosynthesis was implemented to balance cellular growth with l-ornithine production. Second, a bidirectional dynamic control system was established to simultaneously repress the competing l-proline pathway and facilitate l-ornithine export, thereby coordinating intracellular synthesis with extracellular secretion. Third, fed-batch fermentation identified N-acetyl-l-ornithine accumulation as a major metabolic bottleneck, which was effectively alleviated by introducing an alternative deacetylation route to bypass feedback inhibition. The final engineered strain produced 65.55 g/L l-ornithine with a yield of 0.40 g/g glucose in a 5 L fed-batch fermentation, the highest de novo l-ornithine titer reported in E. coli to date. This work provides a generalizable framework for balancing growth, metabolic flux, and product secretion in amino acid biomanufacturing.
Abstract Polyhydroxyalkanoates (PHAs) are bio-based polyesters with the potential to replace petroleum-based products. However, their thermal and mechanical properties are limited by the composition of the monomers used to make them. Current approaches to expanding monomer diversity require individualized pathway engineering for each targeted monomer, limiting the scalability of PHA design. In this work, we establish polyketide synthases (PKSs) as a modular and programmable platform for expanding PHA monomer diversity with variable chain length and branching. Using Pseudomonas putidaas a host, we integrate four engineered PKSs and their variants with heterologous CoA-activating enzymes and short-chain-length PHA polymerases to produce diverse (R)-3-hydroxyacid monomers with defined substitution and branching at the ɑ- and β-positions and enable their in vivo polymerization. We also apply active learning-assisted directed evolution to improve the heterodimeric PHA polymerase CapPhaEC, achieving a 3.6-fold increase in incorporation of α-branched monomers. Leveraging the modularity of PKSs, we demonstrate the biosynthesis and polymerization of linear, branched, and aromatic monomers by exchanging acyltransferase (AT) and ketoreductase (KR) domains to control chain length, branching, and stereochemistry. This study illustrates the utility of using a flexible, programmable PKS platform for rapidly designing and testing new-to-nature PHAs.
Abstract Clustered regularly interspaced short palindromic repeats (CRISPR) and CRISPR-associated (Cas) proteins constitute adaptive immune systems in prokaryotes and have transformed life sciences, precision medicine, and synthetic biology as programmable genome-editing tools. Despite their broad utility, naturally occurring DNA-targeting Cas effectors remain constrained by several intrinsic limitations, including large protein size that complicates delivery, stringent protospacer adjacent motif (PAM) requirements that restrict targetable genomic space, and mismatch tolerance that can lead to off-target activity and potential genotoxicity. These challenges have made Cas protein engineering and the discovery of novel CRISPR and CRISPR-like systems from metagenomic resources central to the development of next-generation genome-editing platforms. This Review places recent advances within an integrated synthetic biology engineering continuum that links natural effector discovery, structure-guided hypothesis generation, high-throughput functional screening, machine learning-enabled model construction, and iterative redesign. This Review summarizes progress in the screening, optimization, and functional engineering of DNA-targeting CRISPR and CRISPR-like effectors, with emphasis on structure-guided rational design, directed evolution coupled with high-throughput screening, bioinformatics- and evolution-guided mining of novel systems from large-scale sequence databases, and artificial intelligence-assisted development. By integrating these strategies, we highlight how CRISPR effector engineering is moving toward design-build-test-learn (DBTL)-inspired workflows that expand the functional landscape of genome-editing technologies and advance genome editing toward improved efficiency, safety, and programmability.
Abstract Type A carbohydrate-binding modules (CBMs) preferentially recognize crystalline polysaccharides such as cellulose and chitin, yet some can also bind synthetic plastics, suggesting that their recognition properties can be redesigned through protein engineering. Here, we used directed evolution to alter the substrate specificity of the archaeal chitin-binding protein PfChBD2 toward plastics. Two rounds of evolution, targeting surface residues and conserved aromatic residues, were screened by phage display, yielding mutants with markedly reduced chitin affinity and distinct PET-binding profiles characterized by apparent binding parameters. Structural modeling indicated that substitutions altering surface electrostatics and hydrophobicity contributed to the shift in substrate preference. When fused to a fluorescent tag, the engineered proteins bound several plastics, including PET, PS, PE, and PP, while showing minimal interaction with natural polysaccharides. The proteins also stained microplastics collected from seawater, demonstrating their potential for environmental detection. This study further demonstrates CBMs as evolutionarily adaptable scaffolds capable of recognizing synthetic polymers and highlights the potential of engineered CBM-based probes for microplastic detection, polymer analysis, and biotechnological and environmental applications.
Synthetic engineering has shown that cytokine receptors from different families can be functionally combined to a large extent. We and others have developed fully synthetic cytokine ligand-receptor systems. Here, our synthetic cytokine receptor (SyCyR) system has been used to generate highly effective non-natural cytokine receptor complexes for the interleukin (IL)-6-type signaling receptors gp130, LIFR, and the erythropoietin receptor (EpoR). Epo promotes the formation of red blood cells in the bone marrow, while IL-6 plays a role in the immune response and regulation of inflammation. Their synergistic effects in combination could contribute to effectiveness of anemia treatment in future in vivo studies. Activation of these SyCyRs is based on homo- and heterodimeric GFP-mCherry fusion proteins as synthetic cytokine ligands and VHH nanobodies (variable domain of single-domain antibodies) against GFP and mCherry fused to the transmembrane and intracellular domains of engineered cytokine receptors. Although synthetic cytokine signaling was achieved for most designs, gp130:LIFR and gp130:EpoR designs were more efficient compared to LIFR:EpoR designs. The gp130:EpoR signaling combines typical signaling signature of both pathways. Taken together, our data show that synthetic cytokine receptors for members of the IL-6 family and EpoR are generally feasible.
Poly(ethylene terephthalate) (PET) is a widely used plastic whose persistence and improper disposal pose serious environmental and health risks. In this study, three novel PET hydrolases TbPETase, AbPETase, and AfPETase were identified from Thermoanaerobacterales, Acidimicrobiales, and Actinokineospora fastidiosa, respectively. Among these, TbPETase exhibited the highest enzymatic activity and thermostability. Based on structural analysis, we performed semirational truncations targeting the intrinsically disordered N- and C-terminal regions of TbPETase, generating two improved variants ΔN36 and ΔC4. The double mutant, TbPETaseΔN36/ΔC4, demonstrated a 2.3-fold increase in overall enzymatic activity and a 2.6-fold improvement in catalytic efficiency (k cat/K m) compared to the wild-type enzyme, along with significantly enhanced thermal stability. Molecular dynamics simulations revealed that the removal of flexible terminal regions increased the overall structural rigidity of TbPETaseΔN36/ΔC4. This structural stabilization was associated with the formation of a hydrogen bond at T215 and a π–π stacking interaction at W193. In a 100 mL one-pot reaction system, the combination of TbPETaseΔN36/ΔC4 with an engineered BMHETase variant, BMHETase6M, achieved 81.2% degradation of semicrystalline PET powder at 60 °C over 60 h, yielding terephthalic acid as the major product. These findings demonstrate the potential of TbPETaseΔN36/ΔC4 as a highly efficient and industrially applicable biocatalyst for PET degradation.
Astaxanthin is a keto-carotenoid with high added value. In this study, we aimed to biosynthesize 3S,3′S-astaxanthin efficiently and sustainably from the renewable single-carbon (C1) feedstock methanol through multiplex metabolic engineering strategies in Komagataella phaffii. First, the K. phaffii cell-free terpene synthesis system was established successfully and applied to evaluate the astaxanthin synthase combinations rapidly. We then systematically engineered K. phaffii for the overproduction of 3S,3′S-astaxanthin from methanol by tuning the carotenoid synthesis module rationally, optimizing the precursor supply and carotenoid storage globally, thereby achieving a significant elevation in astaxanthin content from 3.910 mg/g to 7.513 mg/g. Thereafter, key node enzyme assembly, branch route reconstruction, and cofactor engineering were employed to further improve astaxanthin accumulation, achieving a significant increase of astaxanthin content to 11.397 mg/g. Finally, the astaxanthin production reached 4.75g/L under fed-batch fermentation, which is the highest astaxanthin level reported in an engineered microbe to date. In addition, the synthesized astaxanthin was successfully and effectively applied in shrimp farming for color enhancement and antioxidant effects. These results demonstrate the potential of K. phaffii as a promising platform for sustainable green production of value-added terpenoid compounds from organic one-carbon feedstocks and will pave the way for astaxanthin industrial production.
Abstract Traditional protein design is fundamentally constrained by known sequences and folds. To break free from these limitations, we introduce a new alternative: designing proteins directly from plain-language specifications. To achieve this, we trained MP4, a transformer-based model that maps natural language prompts to protein sequences, on a dataset of 3.2 billion points and 138k tokens. In a benchmark of 96 prompts representing a wide array of functions and contexts, MP4 excelled by simultaneously improving on three key metrics: sequence realism, predicted fold quality, and alignment to the requested function. This high performance is particularly significant as it was achieved using only text as input which is a major departure from other models. Experimental validation confirmed our computational predictions: two de novo designs were experimentally shown to be both expressible and thermostable, with high-resolution crystallography (1.30 Å and 1.77 Å) ultimately revealing one to possess a paradigm-shifting novel fold. Functionally, the designs were also active, demonstrating both ATP binding and hydrolysis in vitro. This work demonstrates the realization of natural-language intent as functional proteins that express, crystallize, and catalyze. Although the underlying approach is still in early development with incomplete coverage and controllability, MP4 delivers a profound impact: it lowers the barrier to protein design and vastly expands the space for creative exploration in molecular programming.
Abstract Although adeno-associated virus (AAV) has enjoyed enormous success as a delivery modality for gene therapy, it suffers from high prevalence of preexisting neutralizing antibodies in human populations, limiting who can receive potentially life-saving treatments. As a novel solution to this issue, we employed SpyTag-SpyCatcher molecular glue technology to facilitate packaging of AAVs inside of recombinant protein vault nanoparticles. Vaults are endogenous particles produced by mammalian cells. We therefore hypothesized that they may shield packaged molecules from neutralizing antibodies. Vaults have previously been utilized to deliver drugs and proteins into cells, but our study represents the first time anyone has packaged an entire virus inside of a vault. We showed that our vaultAAV delivery vehicle transduces cells in the presence of anti-AAV neutralizing serum. VaultAAV is positioned as a new gene therapy delivery platform with potential to overcome the neutralizing antibody problem, expanding the scope of AAV treatments.
Abstract The interaction between the receptor binding domain (RBD) of the SARS-CoV-2 spike protein (S1) and human ACE2 receptor is essential for viral entry into host epithelial cells and a key target for drug discovery. Here, we describe the evolution of a threomer, a base-modified version of threose nucleic acid (TNA), that binds to the S1 protein and inhibits its interaction with ACE2. The aptamer was isolated by in vitro selection using a DNA display strategy that linked each TNA molecule to its encoding double-stranded DNA template. Following iterative cycles of selection and amplification, lead candidates were identified by parallelized screening of individual variants in hydrogel particles. The top-performing hit exhibits low nanomolar affinity to the S1 protein and inhibits formation of the S1−ACE2 complex. Together, this work establishes threomers as an emerging platform for inhibiting protein−protein interactions in therapeutic targets.
Abstract Enfumafungin is the essential precursor for the FDA-approved antifungal agent Ibrexafungerp, yet its production is hindered by low yields, restrictive solid-state fermentation, and the genetic intractability of its producer strain, Hormonema carpetanum ATCC 74360. Here, we developed a CRISPR/Cas9-mediated platform that halved the genetic iteration cycle to less than 15 days. By rewiring the biosynthetic cluster through promoter engineering, we successfully transitioned production to scalable liquid fermentation, achieving an initial 21.6 mg/L. Then, the mevalonate pathway genes (tHMGR, ERG9, and ERG1) were overexpressed, which successfully boosted the titer to 250.1 mg/L. Following medium optimization, the engineered strain HC07 reached a record-breaking 592.4 mg/L in 250 mL shake flasks, the highest yield achieved via metabolic engineering to date. This study provides a robust microbial platform and a streamlined genetic toolbox, paving the way for the sustainable biomanufacturing of next-generation triterpenoid antifungals.
Abstract Aniline is an essential synthetic building block, with its manufacture reliant on fossil carbon feedstocks. While bioproduction could connect aniline manufacture to renewable carbon inputs, efficient metabolic pathway design remains a hurdle for aniline biosynthesis. Two alternate pathways were examined, involving the promiscuous activity of either non-oxidative phenolic acid decarboxylases or tyrosine phenol-lyases (TPLs) on an amine substrate. The decarboxylases exhibited <1% aniline yields under oxygen-scarce conditions, whereas no candidate TPLs produced aniline. Further optimization of the decarboxylase reaction is necessary to construct a viable, fully biosynthetic route towards aniline. Incidentally, we observed that (1) aniline reacts abiotically with glucose in culture medium, and (2) the existing literature on aniline biosynthesis heavily references the EC 4.1.1.24 aminobenzoate decarboxylase classification, which has not been validated since first proposed almost seventy years ago. We recommend the EC 4.1.1.24 entry be archived, supporting the EC 4.1.1.61 phenolic acid decarboxylase entry instead.
Abstract Accurate identification and classification of CRISPR-Cas systems are crucial for understanding microbial immune mechanisms and developing novel genome-editing tools. However, traditional homology-based mining methods face severe computational bottlenecks and assembly fragmentation challenges when processing massive metagenomic data. Here, we present DeepCRISPR-Typer, a comprehensive computational framework integrating a large protein language model (TEMC-Cas), a deep sequence feature extractor (CRISPR-RepTyper), and an adaptive targeted HMM profiling strategy. DeepCRISPR-Typer integrates array and Cas evidence and significantly reduces computational overhead by dynamically invoking subtype-specific HMM subsets. In metagenomic dataset evaluations, DeepCRISPR-Typer achieved a classification accuracy of 94.26% and demonstrated a significant acceleration of approximately 1 orders of magnitude compared to existing mainstream tools. This research provides a robust and scalable engine for metagenome-scale CRISPR system discovery, significantly expanding the mining toolbox for genome engineering applications.