
Genetically encoded noncanonical amino acids (ncAAs) enable site-specific installation of chemical functionalities into proteins, expanding the scope of protein engineering and bioconjugation. Here we report a photo-enhanced oxidative coupling strategy that leverages a vinyl sulfide-containing ncAA for selective protein labeling. An ortho-naphthoquinone methide (oNQM) intermediate is generated in situ from a stable precursor under mild oxidative conditions using ferricyanide, and labeling efficiency is markedly enhanced upon 365 nm irradiation. Ethyl vinyl sulfide was identified as a compact, electronically suitable dienophile that can be incorporated into proteins either through lysine modification or via genetic encoding of N6-((2-(vinylthio)ethoxy)carbonyl)-L-lysine (VtK). Under photo-enhanced oxidative conditions, robust labeling was observed at pH 6-7 with minimal background modification of wild-type proteins. Site-specific incorporation of VtK into the outer membrane protein OmpX enabled selective labeling in cell lysates and in live Escherichia coli cells, demonstrating compatibility with complex biological environments. This work establishes a genetically encodable, photo-enhanced oxidative coupling modality that complements existing ncAA-based bioconjugation strategies and expands the protein engineering toolbox.
Antibody prodrugs provide a strategy to reduce systemic toxicity by masking therapeutic antibodies until activation in specific physiological environments. Nivolumab, an anti-PD-1 checkpoint inhibitor used in cancer immunotherapy, can cause immune-related adverse events. As a first step toward a prodrug version of nivolumab, we screened an Escherichia coli affibody library using MACS and FACS, identifying affibodies that effectively mask its PD-1-binding regions. Deep sequencing revealed an unexpected enrichment of proline-rich variants, possibly mimicking a PD-1 loop. AlphaFold modeling suggested that these affibodies form interactions with nivolumab despite low alpha-helical content. Biosensor assays confirmed effective masking in a nivolumab prodrug format, with PD-1 binding restored upon proteolytic cleavage. These findings support further exploration of PD-1-mimicking affibodies as masking domains, advancing the development of more selective nivolumab prodrugs with improved safety profiles.
The integration of machine learning tools into protein engineering offers substantial promise, yet linking computational predictions to experimental performance remains challenging. Here, we applied accessible computational platforms to engineer Ideonella sakaiensis MHETase, a key enzyme in the biodeconstruction of polyethylene terephthalate (PET). Homologue identification with EnzymeMiner, followed by solubility-focused redesign using AggreProt and ProteinMPNN, enabled the finding of MHETase-like enzymes and the generation of is-MHETase variants with increase in soluble expression, as supported by predictive scoring and experimental validation. However, retaining catalytic activity proved considerably more difficult. Structural modeling and kinetic analyses revealed that ProteinMPNN introduced substitutions incompatible with MHETase function, reflecting biases toward sequence patterns common in broader esterase families rather than constraints specific to this narrow enzyme subgroup. These findings highlight both the utility and the current limitations of accessible ML tools for enzyme engineering, underscoring the continued necessity of experimental validation to guide and refine computational predictions.
Polyaspartic acid (PAA) is a biodegradable polymer with various industrial applications. To date there are only three known PAA hydrolases (from the gene PahZ) capable of degrading PAA . These enzymes are expressed in two different bacteria, Sphingomonas sp. KT-1 (PahZ1KT-1 and PahZ2KT-1) and Pedobacter sp. KP-2 (PahZ1KP-2). PahZ1KT-1 and PahZ2KT-1 form a two-component system degrading tPAA to oligoaspartic acid (OAA) and subsequently into aspartic acid. This study aims to expand the diversity of PAA hydrolases and inform efforts to improve PAA degradation. To further understand the known PahZ1 homologs, the X-ray crystal structure of PahZ1KP-2 was determined to examine its structural homology with PahZ1KT-1. Crystallographic analysis revealed PahZ1KP-2 is monomeric, contrasting with the dimeric PahZ1KT-1, yet both share a conserved serine protease catalytic triad. With the aim of expanding the PahZ1 family, four putative homologs were identified using bioinformatics and AI-based structural modeling, all of which retained the α/β hydrolase domain. Importantly, all homologs exhibited measurable PAA-degrading activity and each was classified as either monomer or dimer to further expand the diversity of PahZ1 enzymes and provide a broader toolbox for sustainable polymer degradation.
Antibodies, and in particular bi- or multispecific antibodies, are the most promising candidates for new therapy options in cancer or inflammatory diseases. So far, multiple different antibody variants containing only distinct parts of antibodies, e.g. the variable domains (Fv), and a plethora of bi- and multispecific antibody formats have been described. The vast majority of these formats have been created using classical protein engineering strategies, i.e. by combining or changing protein domains or introducing distinct mutations into the protein chains to achieve formation of the envisaged proteins. Antibodies consist of different domains that can fold independently into a three-dimensional structure. Here we have focused on the constant domain 2 of the heavy chain (CH2) that has no direct contacts to other domains in an antibody architecture. As protein complementation is a well-established methodology, we asked whether it is possible to split and reassemble the CH2 domain. Here we demonstrate that it is possible to split the CH2 domain into two parts, fuse the two CH2 domain parts to different antibody domains, and show that the split CH2 parts are able to reassemble into functional proteins. The usage of the split CH2 domain approach as a building block offers new opportunities for the design of bi- or multispecific antibody variants.
Miniprotein designs are emerging as promising antibody alternatives for therapeutic and diagnostic use. Using RFdiffusion, we designed 538 binder candidates targeting both N and C-terminal domains of the SARS-CoV-2 nucleocapsid protein, selecting 19 for recombinant production in E. coli. All were soluble and purified (8-10 mg/L yield), though 13 unexpectedly formed oligomers. CD spectroscopy confirmed proper folding and thermal refolding, with 9 miniproteins exhibiting Tm > 75°C. Binding assay revealed three miniproteins with affinities of 115 nM, 4.8 μM, and 7.32 μM. The 1.8 Å resolution crystal structure of one binder (Gpx62) matched the predicted design. However, predictive metrics (like ipTM of AlphaFold3 and computational simulations) did not align with experimental data. Together, these results reveal unintended oligomerization as a major, previously underappreciated barrier in miniprotein binder discovery, demonstrating that while RFdiffusion reliably predicts structural integrity and stability of miniprotein, its current metrics do not account for the oligomeric behavior that might critically limit binding competence.
Several technologies leverage avidin/biotin interactions for complexation, including the MAPS vaccine technology which utilizes rhizavidin, a biotin-binding protein derived from the proteobacterium Rhizobium etli, to complex antigens with biotinylated polysaccharides. Rhizavidin possesses five potential N-linked glycosylation sites which are not glycosylated in the native bacterium or when rhizavidin is recombinantly expressed in Escherichia coli. However, when expressed in eukaryotic cell systems, these sites undergo variable and non-physiologically relevant glycosylation that complicates purification. To overcome these challenges, we engineered rhizavidin with substitutions that abolish unnatural N-linked glycosylation while maintaining rhizavidin's biotin-binding functionality. As a proof of concept, this construct was genetically fused to the receptor binding domain of the spike protein from the SARS-CoV-2 virus, expressed in mammalian cells, and was successfully incorporated into MAPS technology-based vaccines. This newly engineered rhizavidin enhances the versatility of the MAPS technology, enabling the targeting of viruses and tumor-associated antigens that often require mammalian post-translational modifications.
Antibody-dependent cellular cytotoxicity is a key mechanism for antibody-based therapeutics, and current engineering strategies to enhance ADCC primarily rely on two approaches: Fc mutations to improve the antibody's intrinsic CD16a affinity, or the fusion of binding-modules targeting NK cell receptors. The former often compromises antibody stability and induces CD16a downregulation; the latter occupies sites critical for target association and limits the assembly of multi-specific therapeutics. Here, we introduce a novel Fc engineering approach wherein the CH2 domain of the Fc region is replaced with an anti-NKp46 VHH named VIF-Ig. It has been reported that NKp46 expression remains unaltered after NK cell activation across various tumor microenvironments, addressing a key limitation of CD16a-dependent strategies. The novel molecules exhibit potent ADCC, high thermal stability, and retain FcRn binding for favorable pharmacokinetic profiles. Furthermore, we demonstrate that VIF-Ig can accommodate VHHs to target various epitopes. Thus, this versatile modular platform is suitable for developing next-generation NK cell or even other cell engagers with enhanced efficacy and tunable specificity, especially for emerging multi-specific immune cell engagers.
The rapid evolution of SARS-CoV-2 presents substantial challenges to maintaining diagnostic accuracy. Single-domain antibodies (variable domains of camelid heavy-chain antibodies, VHHs), commonly referred to as nanobodies, are attractive tools for diagnostic applications due to their high specificity, stability, affinity and expression yield. In this study, three high-affinity VHHs that recognize the receptor-binding domain (RBD) of the SARS-CoV-2 BA.2 variant were isolated from a large naïve alpaca VHH phage display library. These nanobodies exhibited nanomolar affinities and high thermal stability. Notably, reformatting one VHH into a bivalent VHH-Fc fusion construct, K115.2-Fc, enhanced binding affinity 24-fold, achieving an equilibrium dissociation constant of 68.3 picomolar. Functional characterization demonstrated that K115.2-Fc cross-reacted with the Wuhan SARS-CoV-2 RBD and multiple variants, including Alpha, Beta, Gamma, Delta, Kappa, BA.1, BA.2, and BA.4/5. In indirect enzyme-linked immunosorbent assay, K115.2-Fc detected full-length spike protein with a limit of detection of 8.1 pM. Moreover, it enabled sensitive detection in western blotting and flow cytometry, highlighting its applicability across diverse diagnostic applications. Collectively, these findings identify K115.2-Fc as a promising candidate with potential for SARS-CoV-2 diagnostic utility.
Proteins and peptides underpin essential biological functions and technological applications, from targeting disease-relevant interactions to providing broad enzymatic activities. However, engineering molecules with desired properties remains difficult, owing to complex sequence-structure-function relationships and the lack of data on specific systems. Experimental selection strategies, including directed evolution, phage display, and mRNA display, address this challenge by leveraging high diversity libraries and iterative enrichment under defined selection pressures. This allows for the identification of candidates without requiring extensive prior knowledge, and can generate extensive datasets for use in machine learning. While many selection systems exist, comparisons across different selection approaches are hindered by the lack of a unifying analytical framework. Here, we developed a toolset of broadly applicable analyses for assessing selection dynamics in multi-round or multi-condition experiments, ranging from position level analysis of sequence properties to full sequence space mappings through protein language model embeddings. Performing analyses across different systems, we identify desirable traits in selection experiments including enrichment of distinct sequence patterns and correlation between enrichment and final desired functions. Notably, even under weak selection regimes with all sequences <1% frequency, functional sequences (e.g., 70nM IC50 binder to SARS-CoV-2 main protease) are still consistently enriched. We also find repeated selections of the same starting library can help differentiate selection effects of varying conditions (e.g., different delivery of metal ligand) from system noise. These findings, along with the toolset, can be used to guide experimental design, interpretation, and troubleshooting across protein and peptide discovery platforms.
T7 RNA polymerase (T7 RNAP) is a foundational enzyme for biotechnology, but its utility for many potential applications is limited by low thermal stability of 43-44°C. While stabilized variants exist, the most stable commercial version has a proprietary sequence. In this work we developed a highly stable T7 RNAP using structure-based computational design. We combined mutations from previous stabilized variants (M5, M8, V7abcd) with new mutations identified by PROSS. These mutations were filtered using data-driven heuristics to preserve function. Our final design, T7T+, contains 30 point mutations from the original T7 RNAP and demonstrates a functional stability (T50) of 54.9°C in a thermal challenge assay, which is 2.4°C higher than the most stable, published open-source variant to date. Circular dichroism spectroscopy showed an apparent melting temperature of 53.8°C. T7T+ retains 59% of wild-type activity at 37°C. 16 of the 18 tested protein designs had higher stability against thermal challenge compared with the genetic background, attesting to the high success rates of existing non deep learning computational methods for the design of stable, functional proteins. A plasmid encoding T7T+ has been deposited in AddGene and is freely available for non-commercial use.
Model organisms significantly advance research because scientific communities recognize their value, develop infrastructure around them, and utilize them to address questions relevant to multiple species. I propose adopting a similar approach at the molecular level by establishing an official model protein system. This system would acknowledge five proteins that are already informally recognized as models in protein science: GFP, lysozyme, hemoglobin/myoglobin, RNase A, and bacteriorhodopsin. Formal recognition of these proteins would create standardized reporting requirements, shared benchmarks, and reference datasets, enhancing reproducibility, comparability, and educational resources. GFP serves as a prime example of this concept: it is a single-gene, genetically portable protein with a conserved structure and a measurable phenotype that effectively connects computation and experimental research. This paper identifies the criteria for model protein designation, presents a minimal reporting checklist, and outlines initial steps for implementation at the system level.
Catechol 2,3-dioxygenases (C23DO) catalyze extradiol cleavage of catechol in aromatic hydrocarbon degradation pathways. Here, we report mutagenesis studies on two C23DO isoenzymes from Diaphorobacter sp. strain DS2, aimed at elucidating the mechanism. Docking studies with various substrates on the two isozymes identified critical active-site residues and second-sphere interactions contributing to substrate binding and catalysis. These modeling studies identified eight site-directed mutants, which were designed, produced, and kinetically characterized. Substitution of histidine-to-glutamine (His-to-Gln) (H206Q64/H200Q68) activates the previously inert substrate 2,3-dihydroxybenzoic acid (2,3-DHBA), allowing its intradiol cleavage under strong acidic conditions. The mutants retained ~20-30% of their native activity toward catechol but exhibited a mechanistic switch from Fe2+-assisted extradiol to Fe3+-mediated intradiol cleavage with 2,3-DHBA. These results highlight the catalytic adaptability of acid-base residues and demonstrate how subtle second-sphere mutation can alter dioxygenase regiospecificity, providing new insights for biocatalyst engineering and environmental bioremediation.
Edited by: Robert E. Campbell Based on the Anticalin H1GA which tightly binds Aβ40 and Aβ42 peptides - both established biomarkers of Alzheimer's disease - we describe the design of a protein-dye conjugate as analytical reagent that shows strongly elevated fluorescence upon Aβ binding. An unpaired Cys residue was introduced at seven positions within the four loop segments that shape the ligand pocket of the engineered lipocalin. Five of these mutants were purified in the monomeric state and allowed the site-specific conjugation with IANBD amide as a solvatochromic fluorophore. Three conjugates showed ligand-dependent fluorescence and one of these, derived from H1GA(D45C), exhibited sixfold higher emission at 546 nm upon complex formation with the peptide while revealing a low KD value of 1.2 ± 0.8 nM, even in the presence of 5% (w/v) albumin. This NBD-conjugated Anticalin offers a novel biosensor with potential for the detection of Aβ peptides in biochemical assays or human body fluid samples.
Engineering improved protease activity using directed evolution is challenged by uncertainty in sequence-function mapping and inefficiency in evaluating activity of candidate mutants. We implemented a generalizable yeast surface display approach that co-displays protease mutants with substrate on the same Aga2 anchor protein. Identification of enhanced activity mutants is enabled by protease cleavage of tethered substrate removing an N-terminal epitope tag, which empowers flow cytometric isolation of cells with a decrease in signal from fluorophore-linked anti-epitope antibodies. The sequence space of tobacco etch virus protease (TEVp), commonly used for specific cleavage of recombinant protein affinity tags, has previously been investigated through random mutagenesis. Leveraging our display platform, we performed high throughput screens on seven active site combinatorial libraries created via saturation mutagenesis. Beneficial mutations were incorporated into a single second-generation library, which was screened to identify individual beneficial mutations that performed optimally in a multi-mutant context. The vast majority of resultant TEVp multi-mutants improved catalytic efficiency, generally by decreasing KM. The yeast surface protease/substrate co-display system, the insights gleaned on rational library design and mutation combination strategy, and the TEVp sequence-function map will aid future protease engineering efforts.
Nanobodies offer unique advantages in biomedical and biotechnological applications due to their smaller size, ability to bind challenging epitopes, and affordable production using recombinant technology. However, challenges in large-scale production, stability, and solubility limit their widespread use. To address this, we use artificial intelligence tools to optimize the scaffold region of nanobodies. We apply our approach to four nanobodies against clinically relevant targets: the cytokine tumor necrosis factor alpha, the chemotherapeutic drug methotrexate, the pancreatic biomarker amylase, and the placental hormone chorionic gonadotropin. For all the nanobodies tested, we improve stability, production, and intracellular stability while maintaining antigen-binding affinity. Our results thus demonstrate the potential for using AI-driven protein engineering to enhance the properties of nanobodies, offering insights into the interplay between stability, solubility, and antigen binding. Given the high conservation of the scaffold, we propose some mutations that could directly transfer to other nanobodies, providing an easy-to-implement, generalizable engineering strategy.
Interleukin-17A (IL-17A) is a cytokine involved in pro-inflammatory responses and tissue regeneration, with potential therapeutic and research applications. However, its short serum half-life limits in vivo use. Here, we report the systematic design of fc-IL-17A fusion proteins for extended half-life. Through computational analysis of 25 design variants using AlphaFold, we found that IL-17A's native N-terminal unstructured region functions as a crucial natural linker that cannot be effectively replaced by artificial sequences. We therefore generated mouse and human fc-IL-17A variants using direct N-terminal fusion without additional linkers. The resulting proteins retain IL-17A's ability to stimulate IL-6 production and erythroid cell growth. Pharmacokinetic analysis confirms that the Fc fusion increases the serum half-life in mice from 1.5 to 13 hours post-subcutaneous injection. This enables tractable experimental use of IL-17A in vivo for studying its role in inflammation and tissue repair. We further perform pharmacokinetics and pharmacodynamics modeling and propose a dosing regimen with reduced frequency of injection for delivering comparable IL-17A activity. This work provides a valuable pharmacological tool for injectable delivery, enabling investigation of IL-17A's biological functions in homeostasis and disease and exploration of its therapeutic potential in tissue regeneration.
Clustered regularly interspaced short palindromic repeat interference (CRISPRi), the fusion of nuclease-inactive Cas9 with transcriptional repressor domains, is a powerful platform enabling site-specific gene knockdown across diverse biological contexts. Previously described CRISPRi systems typically utilize two distinct domain classes: (1) Krüppel-associated box domains and (2) truncations of the multifunctional protein, MeCP2. Despite widespread adoption of MeCP2 truncations for developing CRISPRi platforms, individual contributions of subdomains within MeCP2's transcriptional repression domain (TRD) toward enhancing gene knockdown remain unclear. Here, we dissect MeCP2's TRD and observe that two subdomains, the expected NcoR/SMRT interaction domain (NID) and an embedded nuclear localization signal (NLS), can separately enhance gold-standard CRISPRi platform performance beyond levels attained with the canonical MeCP2 protein truncation. Incorporating side-by-side analyses of nuclear localization and gene knockdown for over 30 constructs featuring MeCP2 subdomains or virus-derived NLS sequences, we demonstrate that appending C-terminal NLS motifs to dCas9-based transcriptional regulators, both repressors and activators, can significantly improve their effector function across several cell lines. We also observe that NLS placement greatly impacts CRISPRi repressor performance, and that modifying the subdomain configuration natively found within MeCP2 can also enhance gene suppression capabilities in certain contexts. Overall, this work demonstrates the interplay of two complimentary chimeric protein design considerations, transcriptional domain 'dissection' and NLS motif placement, for optimizing CRISPR-mediated transcriptional regulation in mammalian systems.
Tuning in vivo activity of protein therapeutics can improve their safety. In this vein, it is possible to add a ‘mask’ moiety to a protein therapeutic such that its ability to bind its target is prevented until the mask has been proteolytically removed, for instance by a tumor-associated protease. As such, new methods to isolate functional masking sequences can aid development of protein therapies. Here, we describe a yeast display-based method to discover peptide sequences that prevent binding of antibody fragments to their antigen target. Our method includes an in situ ability to screen for restoration of binding by scFvs after proteolytic mask removal, and it takes advantage of the antigenic target itself to guide mask discovery. First, we genetically linked a yeast-displayed αPSCA scFv to overlapping ‘tiles’ of its target. By selecting for reduced antigen binding via flow cytometry, we discovered two peptide masks that we confirmed to be linear epitopes of the PSCA antigen. We then expanded our method towards developing masks for three-dimensional epitopes by using a co-crystal structure of an αHer2 antibody in complex with its antigen to guide combinatorial mask design. In sum, our efforts show the feasibility of employing yeast-displayed, antigen-based libraries to find antibody masks.
Erythropoietin (EPO) suppresses apoptosis and promotes survival by signaling through EPO-R/EPO-R on hematopoietic progenitors or EPO-R/CD131 on non-hematopoietic cells. However, EPO signaling through EPO-R/CD131 is controversial and there is no solved structure of a complex. Here, we constructed a structural model of EPO-R/CD131 and designed several anti-EPO-R, anti-CD131 bispecific proteins that selectively activate EPO-R/CD131. Treatment with these fusion proteins is sufficient to activate STAT5 phosphorylation downstream of EPO-R/CD131 without engaging EPO-R/EPO-R. We demonstrated that proteins with a tandem scFv or bispecific antibody format activate EPO-R/CD131, in contrast to an equimolar mixture of the individual scFvs. Finally, we explored the effect of modifications to binding domain arrangement and linker length and found results consistent with our structural model of an EPO-R/CD131 complex. These findings highlight the utility of bispecific scaffolds in the development of cytokine receptor agonists and provide a foundation for the study of EPO-R/CD131 biology and future clinical development.