
The fetal form of the insulin receptor (IRA) is predominantly expressed in many cancers, where IR signaling promotes progression of the disease. Alternatively, the adult IR isoform (IRB) is predominantly expressed in normal tissues. Previous attempts to target IR indiscriminately via dual IR/IGF1R tyrosine kinase inhibitors frequently failed due to on-target off-tumor effects. However, an IRA-specific therapeutic has the potential to specifically target cancer while sparing normal tissues. Furthermore, an IRA-specific binder would advance research regarding this disease-associated isoform. Yet, isoform-specific IR binders have not been developed. Herein, we engineer six high-affinity, IRA-specific protein binders from a synthetic miniprotein scaffold. These miniprotein variants have IRA KDs ranging from 0.98-7.6 nM with no appreciable IRB binding at 100 nM ligand. Five of the six variants are synergistic agonists. One variant (V1) is a passive, IRA-specific binder with a KD = 6.3 nM and no IRB binding at 1 μM ligand as well as no IGF1R binding at 100 nM, which provides compelling potential as a targeted therapeutic. Additionally, we present the discovery process for these molecules, a yeast display directed evolution campaign utilizing a combination of sorting techniques. Retrospective analysis of the discovery process revealed the importance of alternating selections for target affinity and isoform specificity to generate binders with a favorable combination of these properties.
Bacteriophage M13 substrate libraries are a useful tool to discover novel protease substrates. In this study, we have mapped 7-mer protease substrates cleaved by native Escherichia coli proteases during phage outgrowth. Cleavage was monitored by the loss of a biotinylated AviTag at the N-terminus of the recombinant M13 minor coat protein, protein III. Using this method, we isolated 120 phage-displayed substrates that contain dibasic amino acid motifs (Lys-Lys, Arg-Arg, Lys-Arg, and Arg-Lys), of which the Lys-Lys motif has not been previously described using substrate phage. Since dibasic motifs are known to be substrates of the outer membrane protein proteases OmpP and OmpT, the phage library was transferred into a mutant strain deficient in both protease activities, and the selection process was repeated. Sequencing of recovered clones revealed no enrichment of the four dibasic motifs, demonstrating that all four motifs are substrates of Omptin proteases. These findings indicate that phage propagation in Omptin-expressing bacteria leads to the loss of dibasic amino acid motifs from phage-displayed combinatorial peptide, protein, and antibody libraries. This suggests that the presence of dibasic residues in proteins and peptides displayed on bacteriophage M13 should necessitate the use of OmpT- and OmpP-deficient bacterial hosts.
Protein bioconjugates have a wide range of applications, including for targeted drug delivery and in vivo diagnostics. Expanding the chemistry available to link protein to cargo remains a critical consideration for the development of novel bioconjugates for unmet needs. Currently, most FDA-approved protein bioconjugates are based on the antibody structure, which constrains the available conjugation chemistries and applications. Alternative scaffold proteins have potential as targeting molecules that advance the synthetic possibilities. Here, we report two methods of high-yield incorporation of azide-containing groups into an Fn3 domain protein scaffold to enable click chemistry conjugation. One approach incorporates noncanonical amino acid azidohomoalanine during recombinant protein expression, and the other method develops a linker system using a unique thiol residue in the Fn3 structure. These synthetic approaches enabled orthogonal, dual labeling. The resulting conjugates retain functionality in receptor binding assays, validating these methods for diverse applications in engineering scaffold proteins for bioconjugate applications.
Generative models have transformed structural bioinformatics, enabling antibody and nanobody design against protein epitopes; however, nanobody engineering for small molecule sensing remains largely experimental. In this work, we evaluate seven state-of-the-art structure predictors, AlphaFold3, Chai-1, Boltz-2x, RoseTTAFold3, Protenix, FlowDock and OmegaFold, on nanobody-small molecule complexes. Most predictors accurately reproduced nanobody and CDR geometries but struggled with ligand placement and orientation, although co-folding improved overall accuracy. Contact analysis revealed that CDR1, rather than CDR3, was predominantly involved in ligand binding. Intrinsic confidence scores correlated poorly with experimental accuracy and showed limited power in distinguishing binders from non-binders. Increasing the number of samples and seeds yielded modest gains in accuracy, whereas additional recycles did not. These findings highlight both the strengths and limitations of structure prediction methods for nanobody-small molecule complexes.
With this status report, we aim to provide a timely snapshot of the protein engineering field as a broad and rapidly advancing discipline that integrates computational, molecular biology, structure-guided, evolutionary, and synthetic approaches to create new and improved proteins with tailored structures and useful functions. The report is organized into eight thematic areas spanning core methodologies and major application domains, including enzymes, therapeutics, detection, synthetic biology, and materials. Contributions from experts across these areas highlight both the historical foundations and recent advances in their respective fields, with particular emphasis on the growing influence of machine learning and artificial intelligence-based methods. Emerging from this broad overview is a central message: protein engineering appears to be entering a golden age, defined by a rapidly accelerating pace of progress, even as significant challenges in design, screening, and real-world application remain. Looking ahead, the continued integration of computational and experimental strategies is poised to further accelerate the impact of protein engineering across an expanding range of economically and societally important sectors, from therapeutics and molecular imaging to diagnostics, plastic recycling, and industrial chemistry.
Engineering proteins for therapeutic applications is a field that has seen substantial growth in the past two decades, but challenges remain. A major deficiency in current strategies is the lack of a means to efficiently screen for binding to key epitopes to yield a desired function. To meet this challenge, we designed a tethered yeast surface display construct that leverages high local concentrations by tethering a candidate protein binder to a target protein of interest (POI) via a flexible peptide linker. These high local concentrations enable screening of epitope-specific binders based on decreased signal due to inhibitory function when the construct is assayed alongside a competitive POI binder. We demonstrate that epitope-specific screening and enrichment is possible based on fluorescent output and study key optimization parameters. We anticipate this technique will accelerate binder development by reducing the need for downstream low throughput epitope mapping of isolated proteins.
Recent years have seen remarkable clinical success with therapies that harness the patient's immune system, with checkpoint inhibitors in oncology as a prominent example. Non-invasive monitoring of immune activation in vivo has the potential to accelerate both basic immunology research and clinical drug development, offering a valuable means to track responses to emerging immunotherapies. CD69 is a rapidly induced activation marker on lymphocytes and other leukocytes, making it an attractive imaging target. For such applications, affibody molecules offer distinct advantages as radio imaging tracers due to their small size, high affinity, and rapid pharmacokinetics, resulting in excellent imaging contrast. Here, we combined directed evolution with computational design to optimise a CD69-binding affibody molecule. An alanine scan of the parental binder informed the construction of a diversification library, which was displayed on Escherichia coli and subjected to iterative MACS and FACS selections with stringent off-rate competition. The top selection hit was subsequently refined through site-directed mutagenesis, including variants suggested by a general protein language model. The resulting lead, Z1525, bound human CD69 with single-digit nanomolar affinity and showed improved thermal stability while retaining solubility and refolding capacity, consistent with suitability for radiolabelling and in vivo targeting. Most importantly, it displayed selective binding to CD69 on stimulated Jurkat cells, with negligible binding to resting cells. These results establish E. coli display with off-rate-driven selection as an efficient strategy for affibody affinity maturation and demonstrate that protein language models can effectively guide improvements in folding stability. Z1525 represents a promising radio imaging tracer candidate for monitoring immune activation in vivo and warrants further preclinical development.
Genetically encoded noncanonical amino acids (ncAAs) enable site-specific installation of chemical functionalities into proteins, expanding the scope of protein engineering and bioconjugation. Here we report a photo-enhanced oxidative coupling strategy that leverages a vinyl sulfide-containing ncAA for selective protein labeling. An ortho-naphthoquinone methide (oNQM) intermediate is generated in situ from a stable precursor under mild oxidative conditions using ferricyanide, and labeling efficiency is markedly enhanced upon 365 nm irradiation. Ethyl vinyl sulfide was identified as a compact, electronically suitable dienophile that can be incorporated into proteins either through lysine modification or via genetic encoding of N6-((2-(vinylthio)ethoxy)carbonyl)-L-lysine (VtK). Under photo-enhanced oxidative conditions, robust labeling was observed at pH 6-7 with minimal background modification of wild-type proteins. Site-specific incorporation of VtK into the outer membrane protein OmpX enabled selective labeling in cell lysates and in live Escherichia coli cells, demonstrating compatibility with complex biological environments. This work establishes a genetically encodable, photo-enhanced oxidative coupling modality that complements existing ncAA-based bioconjugation strategies and expands the protein engineering toolbox.
Antibody prodrugs provide a strategy to reduce systemic toxicity by masking therapeutic antibodies until activation in specific physiological environments. Nivolumab, an anti-PD-1 checkpoint inhibitor used in cancer immunotherapy, can cause immune-related adverse events. As a first step toward a prodrug version of nivolumab, we screened an Escherichia coli affibody library using MACS and FACS, identifying affibodies that effectively mask its PD-1-binding regions. Deep sequencing revealed an unexpected enrichment of proline-rich variants, possibly mimicking a PD-1 loop. AlphaFold modeling suggested that these affibodies form interactions with nivolumab despite low alpha-helical content. Biosensor assays confirmed effective masking in a nivolumab prodrug format, with PD-1 binding restored upon proteolytic cleavage. These findings support further exploration of PD-1-mimicking affibodies as masking domains, advancing the development of more selective nivolumab prodrugs with improved safety profiles.
The integration of machine learning tools into protein engineering offers substantial promise, yet linking computational predictions to experimental performance remains challenging. Here, we applied accessible computational platforms to engineer Ideonella sakaiensis MHETase, a key enzyme in the biodeconstruction of polyethylene terephthalate (PET). Homologue identification with EnzymeMiner, followed by solubility-focused redesign using AggreProt and ProteinMPNN, enabled the finding of MHETase-like enzymes and the generation of is-MHETase variants with increase in soluble expression, as supported by predictive scoring and experimental validation. However, retaining catalytic activity proved considerably more difficult. Structural modeling and kinetic analyses revealed that ProteinMPNN introduced substitutions incompatible with MHETase function, reflecting biases toward sequence patterns common in broader esterase families rather than constraints specific to this narrow enzyme subgroup. These findings highlight both the utility and the current limitations of accessible ML tools for enzyme engineering, underscoring the continued necessity of experimental validation to guide and refine computational predictions.
Polyaspartic acid (PAA) is a biodegradable polymer with various industrial applications. To date there are only three known PAA hydrolases (from the gene PahZ) capable of degrading PAA . These enzymes are expressed in two different bacteria, Sphingomonas sp. KT-1 (PahZ1KT-1 and PahZ2KT-1) and Pedobacter sp. KP-2 (PahZ1KP-2). PahZ1KT-1 and PahZ2KT-1 form a two-component system degrading tPAA to oligoaspartic acid (OAA) and subsequently into aspartic acid. This study aims to expand the diversity of PAA hydrolases and inform efforts to improve PAA degradation. To further understand the known PahZ1 homologs, the X-ray crystal structure of PahZ1KP-2 was determined to examine its structural homology with PahZ1KT-1. Crystallographic analysis revealed PahZ1KP-2 is monomeric, contrasting with the dimeric PahZ1KT-1, yet both share a conserved serine protease catalytic triad. With the aim of expanding the PahZ1 family, four putative homologs were identified using bioinformatics and AI-based structural modeling, all of which retained the α/β hydrolase domain. Importantly, all homologs exhibited measurable PAA-degrading activity and each was classified as either monomer or dimer to further expand the diversity of PahZ1 enzymes and provide a broader toolbox for sustainable polymer degradation.
Antibodies, and in particular bi- or multispecific antibodies, are the most promising candidates for new therapy options in cancer or inflammatory diseases. So far, multiple different antibody variants containing only distinct parts of antibodies, e.g. the variable domains (Fv), and a plethora of bi- and multispecific antibody formats have been described. The vast majority of these formats have been created using classical protein engineering strategies, i.e. by combining or changing protein domains or introducing distinct mutations into the protein chains to achieve formation of the envisaged proteins. Antibodies consist of different domains that can fold independently into a three-dimensional structure. Here we have focused on the constant domain 2 of the heavy chain (CH2) that has no direct contacts to other domains in an antibody architecture. As protein complementation is a well-established methodology, we asked whether it is possible to split and reassemble the CH2 domain. Here we demonstrate that it is possible to split the CH2 domain into two parts, fuse the two CH2 domain parts to different antibody domains, and show that the split CH2 parts are able to reassemble into functional proteins. The usage of the split CH2 domain approach as a building block offers new opportunities for the design of bi- or multispecific antibody variants.
Miniprotein designs are emerging as promising antibody alternatives for therapeutic and diagnostic use. Using RFdiffusion, we designed 538 binder candidates targeting both N and C-terminal domains of the SARS-CoV-2 nucleocapsid protein, selecting 19 for recombinant production in E. coli. All were soluble and purified (8-10 mg/L yield), though 13 unexpectedly formed oligomers. CD spectroscopy confirmed proper folding and thermal refolding, with 9 miniproteins exhibiting Tm > 75°C. Binding assay revealed three miniproteins with affinities of 115 nM, 4.8 μM, and 7.32 μM. The 1.8 Å resolution crystal structure of one binder (Gpx62) matched the predicted design. However, predictive metrics (like ipTM of AlphaFold3 and computational simulations) did not align with experimental data. Together, these results reveal unintended oligomerization as a major, previously underappreciated barrier in miniprotein binder discovery, demonstrating that while RFdiffusion reliably predicts structural integrity and stability of miniprotein, its current metrics do not account for the oligomeric behavior that might critically limit binding competence.
Several technologies leverage avidin/biotin interactions for complexation, including the MAPS vaccine technology which utilizes rhizavidin, a biotin-binding protein derived from the proteobacterium Rhizobium etli, to complex antigens with biotinylated polysaccharides. Rhizavidin possesses five potential N-linked glycosylation sites which are not glycosylated in the native bacterium or when rhizavidin is recombinantly expressed in Escherichia coli. However, when expressed in eukaryotic cell systems, these sites undergo variable and non-physiologically relevant glycosylation that complicates purification. To overcome these challenges, we engineered rhizavidin with substitutions that abolish unnatural N-linked glycosylation while maintaining rhizavidin's biotin-binding functionality. As a proof of concept, this construct was genetically fused to the receptor binding domain of the spike protein from the SARS-CoV-2 virus, expressed in mammalian cells, and was successfully incorporated into MAPS technology-based vaccines. This newly engineered rhizavidin enhances the versatility of the MAPS technology, enabling the targeting of viruses and tumor-associated antigens that often require mammalian post-translational modifications.
Antibody-dependent cellular cytotoxicity is a key mechanism for antibody-based therapeutics, and current engineering strategies to enhance ADCC primarily rely on two approaches: Fc mutations to improve the antibody's intrinsic CD16a affinity, or the fusion of binding-modules targeting NK cell receptors. The former often compromises antibody stability and induces CD16a downregulation; the latter occupies sites critical for target association and limits the assembly of multi-specific therapeutics. Here, we introduce a novel Fc engineering approach wherein the CH2 domain of the Fc region is replaced with an anti-NKp46 VHH named VIF-Ig. It has been reported that NKp46 expression remains unaltered after NK cell activation across various tumor microenvironments, addressing a key limitation of CD16a-dependent strategies. The novel molecules exhibit potent ADCC, high thermal stability, and retain FcRn binding for favorable pharmacokinetic profiles. Furthermore, we demonstrate that VIF-Ig can accommodate VHHs to target various epitopes. Thus, this versatile modular platform is suitable for developing next-generation NK cell or even other cell engagers with enhanced efficacy and tunable specificity, especially for emerging multi-specific immune cell engagers.
The rapid evolution of SARS-CoV-2 presents substantial challenges to maintaining diagnostic accuracy. Single-domain antibodies (variable domains of camelid heavy-chain antibodies, VHHs), commonly referred to as nanobodies, are attractive tools for diagnostic applications due to their high specificity, stability, affinity and expression yield. In this study, three high-affinity VHHs that recognize the receptor-binding domain (RBD) of the SARS-CoV-2 BA.2 variant were isolated from a large naïve alpaca VHH phage display library. These nanobodies exhibited nanomolar affinities and high thermal stability. Notably, reformatting one VHH into a bivalent VHH-Fc fusion construct, K115.2-Fc, enhanced binding affinity 24-fold, achieving an equilibrium dissociation constant of 68.3 picomolar. Functional characterization demonstrated that K115.2-Fc cross-reacted with the Wuhan SARS-CoV-2 RBD and multiple variants, including Alpha, Beta, Gamma, Delta, Kappa, BA.1, BA.2, and BA.4/5. In indirect enzyme-linked immunosorbent assay, K115.2-Fc detected full-length spike protein with a limit of detection of 8.1 pM. Moreover, it enabled sensitive detection in western blotting and flow cytometry, highlighting its applicability across diverse diagnostic applications. Collectively, these findings identify K115.2-Fc as a promising candidate with potential for SARS-CoV-2 diagnostic utility.
Structural dynamics play a crucial role in protein function, and tuning these dynamics through mutagenesis has emerged as a promising strategy for enhancing activity. However, identifying dynamics hotspots for protein engineering remains a labor-intensive challenge. Here, we demonstrate that NMR peak intensity analysis—a rapid, qualitative method with residue-level resolution—can identify functionally relevant dynamic regions with high precision. Using a family of red fluorescent proteins (RFPs) as a case study, we reveal that flexibility in specific regions of their structures correlates with function. Specifically, as quantum yield increases, the side of the β-barrel closest to the chromophore phenolate moiety becomes more rigid, while the opposite side, closest to the acylimine group, gains flexibility. Notably, the phenolate face corresponds to a mutational hotspot frequently targeted in directed evolution campaigns aimed at enhancing brightness, underscoring its functional significance. B-factor analysis of non-cryogenic X-ray crystal structures further supports our findings. Our results establish NMR peak intensity analysis as a promising tool for mapping functional dynamics hotspots to guide protein engineering campaigns.
Proteins and peptides underpin essential biological functions and technological applications, from targeting disease-relevant interactions to providing broad enzymatic activities. However, engineering molecules with desired properties remains difficult, owing to complex sequence-structure-function relationships and the lack of data on specific systems. Experimental selection strategies, including directed evolution, phage display, and mRNA display, address this challenge by leveraging high diversity libraries and iterative enrichment under defined selection pressures. This allows for the identification of candidates without requiring extensive prior knowledge, and can generate extensive datasets for use in machine learning. While many selection systems exist, comparisons across different selection approaches are hindered by the lack of a unifying analytical framework. Here, we developed a toolset of broadly applicable analyses for assessing selection dynamics in multi-round or multi-condition experiments, ranging from position level analysis of sequence properties to full sequence space mappings through protein language model embeddings. Performing analyses across different systems, we identify desirable traits in selection experiments including enrichment of distinct sequence patterns and correlation between enrichment and final desired functions. Notably, even under weak selection regimes with all sequences <1% frequency, functional sequences (e.g., 70nM IC50 binder to SARS-CoV-2 main protease) are still consistently enriched. We also find repeated selections of the same starting library can help differentiate selection effects of varying conditions (e.g., different delivery of metal ligand) from system noise. These findings, along with the toolset, can be used to guide experimental design, interpretation, and troubleshooting across protein and peptide discovery platforms.
T7 RNA polymerase (T7 RNAP) is a foundational enzyme for biotechnology, but its utility for many potential applications is limited by low thermal stability of 43-44°C. While stabilized variants exist, the most stable commercial version has a proprietary sequence. In this work we developed a highly stable T7 RNAP using structure-based computational design. We combined mutations from previous stabilized variants (M5, M8, V7abcd) with new mutations identified by PROSS. These mutations were filtered using data-driven heuristics to preserve function. Our final design, T7T+, contains 30 point mutations from the original T7 RNAP and demonstrates a functional stability (T50) of 54.9°C in a thermal challenge assay, which is 2.4°C higher than the most stable, published open-source variant to date. Circular dichroism spectroscopy showed an apparent melting temperature of 53.8°C. T7T+ retains 59% of wild-type activity at 37°C. 16 of the 18 tested protein designs had higher stability against thermal challenge compared with the genetic background, attesting to the high success rates of existing non deep learning computational methods for the design of stable, functional proteins. A plasmid encoding T7T+ has been deposited in AddGene and is freely available for non-commercial use.
Model organisms significantly advance research because scientific communities recognize their value, develop infrastructure around them, and utilize them to address questions relevant to multiple species. I propose adopting a similar approach at the molecular level by establishing an official model protein system. This system would acknowledge five proteins that are already informally recognized as models in protein science: GFP, lysozyme, hemoglobin/myoglobin, RNase A, and bacteriorhodopsin. Formal recognition of these proteins would create standardized reporting requirements, shared benchmarks, and reference datasets, enhancing reproducibility, comparability, and educational resources. GFP serves as a prime example of this concept: it is a single-gene, genetically portable protein with a conserved structure and a measurable phenotype that effectively connects computation and experimental research. This paper identifies the criteria for model protein designation, presents a minimal reporting checklist, and outlines initial steps for implementation at the system level.