
Directed evolution is a powerful approach for engineering new biomolecular and cellular functions [1–3]. In contrast to rational design approaches, directed evolution exploits diversity and evolution to shape the behavior of biological matter by applying the Darwinian cycle of mutation, selection, and amplification of genes and genomes. By doing so, the field of directed evolution has generated important insights into the evolutionary process [4–6] as well as useful RNAs, proteins, and systems with wide-ranging applications across biotechnology and medicine [7–11]. To mimic the evolutionary process, classical directed evolution approaches carry out cycles of ex vivo diversification on genes of interest (GOIs), transformation of the resulting gene libraries into cells, and selection of the desired function (Figure 1.1). Each iteration of this cycle is defined as a round of evolution, and as selection stringency increases over rounds, either automatically through competition or manually through changing conditions (or both), this process can lead GOIs closer and closer to the desired function. This overall process makes practical sense for a number of reasons, especially for the goal of protein engineering (i.e. GOI encodes a protein). First, ex vivo diversification is appropriate, because test tube molecular biology techniques such as DNA shuffling, site-directed saturation mutagenesis, and error-prone (ep) polymerase chain reaction (PCR) [2] are capable of generating exceptionally high and precise levels of sequence diversity for any GOI. Second, transforming diversified libraries of the GOI into cells is appropriate, because each GOI variant needs to be translated into a protein in order to express its function, and cells, especially model microbes, are naturally robust hosts for protein expression. Third, carrying out selection inside cells is appropriate, because (i) cells automatically maintain the genotype–phenotype connection between the GOI and expressed protein that is necessary for amplification of desired variants,
Genetic engineering of enzymes has played a significant role in multiple pharmaceutical synthetic processes. The need for selective, robust catalysts that can operate under process chemistry conditions has driven the need for mutagenesis approaches, which have successfully delivered a variety of process-optimized biocatalysts. In this chapter we give an overview of different directed evolution methods that have been applied to pharmaceutical intermediates during the last decade. The exciting progress that has been made since the landmark sitagliptin study in 2010 demonstrates the role that engineered biocatalysts can play in the manufacture of active pharmaceutical ingredients (APIs).
Artificial metalloenzymes are an emerging class of functional biomacromolecules that both provide useful information for native metalloenzymes and have the potential to catalyze important transformations in industrial biocatalysis. Metalloenzymes can be engineered by two complementary methods: rational design and directed evolution. Rational design uses predetermined knowledge of biochemical systems to design a novel protein, often with the aid of computation, while directed evolution exploits the genetic nature of proteins to select for a desired attribute, such as reactivity or stability. Both methods have resulted in the development of metalloenzymes that catalyze abiological reactions, use abiological cofactors, and produce commercially significant compounds. This chapter seeks to provide a survey of the state-of-the-art in the field of metalloenzyme engineering.
In vivo biosensors enable protein functional development, by associating genetic changes to distinguishable traits, in high throughput directed evolution methodologies. With advances in molecular tools for biosensor design and discovery, the scope of tunable proteins and characteristics has expanded considerably in recent years. As the capacity to modulate independent and metabolic pathway proteins increases, directed protein evolution techniques are becoming more relevant in promising sectors such as value-added chemicals, nutraceuticals, and therapeutics biosynthesis. In vivo biosensors can deliver rapid evaluation of protein intracellular responses, including their behaviors in the context of an entire proteome. This chapter presents an overview on the types of in vivo biosensors commonly used in directed protein evolution and some key considerations for biosensor design and selection. Recent novel and inventive applications of biosensors toward single and multivariate directed protein evolution will also be discussed.
Cytochrome P450 monooxygenases are a large class of enzymes, which have evolved to carry out a myriad of oxidation reactions across all kingdoms of life. Reflecting the functional diversity of naturally evolved P450s, these enzymes have proven to be remarkably pliable to protein engineering for improving or altering their substrate profile, catalytic properties, selectivity, and stability. In this chapter, major contributions made in this area are discussed with an emphasis on the different protein engineering strategies, including directed evolution, rational design, or semi-rational approaches, applied for the development of engineered P450 catalysts for biotechnological applications. This survey illustrates the diverse and growing range of applications of engineered P450s, which encompasses the synthesis of chiral building blocks, the preparation of drug metabolites, and the late-stage functionalization of natural products, highlighting their potential toward enabling the synthesis and discovery of bioactive molecules and toward developing sustainable processes for manufacturing of pharmaceuticals and other high-value compounds.
Proteins are composed of 20 proteinogenic amino acids and are central in many biological processes. While traditional mutagenesis is restricted to the 20 proteinogenic amino acids, unnatural amino acid s ( UAA s) can be incorporated into protein through chemical synthesis, chemical mutagenesis, using auxotrophic strains, genetic codon expansion, and other methods, which greatly expands the toolbox for protein engineering. UAAs have been used to increase protein stability, study mechanism of action, tune enzyme activity or selectivity, design novel protein functions, and even engineer synthetic life. Together with other protein engineering methods, UAA incorporation will play an important role in protein engineering, greatly expand what the already very powerful molecules are capable of.
Directed evolution of proteins is a powerful approach to introduce novel functions in proteins through the manipulation of corresponding nucleic acids. Similar to natural selection, directed evolution is based on the fact that genetic diversity leads to phenotypic diversity. It takes advantage of random mutagenesis to generate a broad set of protein variants, of protein production in the host cells, and of the recovery and amplification of clones based on a link between genotype and phenotype. The library design, evolutionary landscape, phenotypic target, and screening system are all among the critical components of a successful directed evolution. Cell-surface display techniques provide a robust platform for the directed evolution of protein variants using high-throughput in vitro screening methods, such as fluorescent-activated cell sorting (FACS). The cell-displayed platforms are capable of carrying thousands and millions of protein mutants, which allow for a rapid selection of the protein variants with desired function, such as improved binding affinity, selectivity, stability, or enzyme activity. Directed evolution based on cell-display platforms is dependent upon the link between the genotype and phenotype of the displayed proteins on the cell surface, in conjunction with the encoding DNA inside the cell. This link between genotype and phenotype allows for the screening and selection of a library of protein mutants based on a desired phenotype, as well as the extraction and isolation of plasmid DNA from isolated single cells. Various techniques for the generation of a library of protein variants, including random mutagenesis and gene shuffling, were used to create a diverse pool of protein mutants to be screened using high-throughput screening methods, such as FACS or plate-reading assays. In this chapter, cell-display techniques, selection methods and strategies, and the modification of cell-surface display systems are described. Some of the recent advancements in revolutionizing directed evolution based on cell-displayed techniques are also discussed. This is a very brief overview of the techniques of and advancements in the directed evolution of proteins using cell-display platforms. Unfortunately, due to limited space, it would be impossible to cover the whole spectrum of the research done in this area.
Protein-based imaging probes provide the power to detect the location and activity of molecular events. Using protein engineering and directed evolution, genetically encoded protein modules have been developed to image live cells. In this chapter, we first introduce the protein engineering and development of various fluorescent proteins and their biosensor derivatives. We then discuss the engineering of binding proteins, i.e. antibody and non-antibody molecular scaffolds, which can be conjugated with various detection modules to image molecules in cells and target specific antigens in vivo . These developed fluorescent proteins and protein motifs are further introduced in applications related to molecular and cellular imaging of signaling events. Lastly, we discuss the applications of these tools for diagnostics and therapeutics via high throughput drug screening. In summary, protein engineering has driven the development of numerous molecular imaging probes for a wide range of applications, spanning from basic biology research to disease diagnosis and treatment.
Directed evolution has matured in academia and industry as a versatile algorithm to redesign enzymes to match demands in biotechnological applications (as documented by the Nobel Prize in chemistry in 2018). Based on the obtained knowledge, computational methods (e.g. FRESCO, FoldX, CNA, PROSS, ProSAR) emerged to be predictive methods to especially improve properties that could be localized within a protein (e.g. thermostability, selectivity, catalytic efficiency, and activity). The main limitation to efficiently explore and benefit from nature's potential in generating better enzymes is the size of the protein sequence space; experimentalists have to admit that they will never be able to experimentally sample through the whole sequence space. A combination of experimental and computational methods proved to be time efficient in redesigning enzymes to meet the application demands (e.g. in chemical and pharmaceuticals synthesis). In this chapter, we highlighted protein engineering strategies that combine directed evolution and computational analysis to efficiently reengineer enzymes and that partly contribute to a molecular understanding of structure function relationship, which can be transferred from enzymes to another. In this respect, an emphasis will be given to the "KnowVolution" (knowledge gaining directed evolution) strategy, which is generally applicable, minimizes experimental efforts, generates a molecular understanding on each positions/amino acid substitution, and was successfully applied to a broad range of and enzymes and properties.
Introduction A protein’s sequence of amino acids encodes its function. This “function” could refer to a protein’s natural biological function, or it could also be any other property including binding affinity toward a particular ligand, thermodynamic stability, or catalytic activity. A detailed understanding of how these functions are encoded would allow us to more accurately reconstruct the tree of life and possibly predict future evolutionary events, diagnose genetic diseases before they manifest symptoms, and design new proteins with useful properties. We know that a protein sequence folds into a three-dimensional structure, and this structure positions specific chemical groups to perform a function; however, we’re missing the quantitative details of this sequence-structure-function mapping. This mapping is extraordinarily complex because it involves thousands of molecular interactions that are dynamically coupled across multiple length and time scales. Computational methods can be used to model the mapping from sequence to structure to function. Tools such as molecular dynamics simulations or Rosetta use atomic representations of protein structures and physics-based energy functions to model structures and functions (1–3). While these models are based on well-founded physical principles, they often fail to capture a protein’s overall global behavior and properties. There are numerous challenges associated with physics-based models including consideration of conformational dynamics, the requirement to make energy function approximations for the sake of computational efficiency, and the fact that, for many complex properties such as enzyme catalysis, the molecular basis is simply unknown (4). In systems composed of thousands of atoms, the propagation of small errors quickly overwhelms any predictive accuracy. Despite tremendous breakthroughs and research progress over the last century, we still lack the key details to reliably predict, simulate, and design protein function. In this chapter, we present the emerging field of data-driven protein engineering. Instead of physically modeling the relationships between protein sequence, structure, and function, data-driven methods use ideas from statistics and machine learning to infer these complex relationships from data. This top-down modeling approach implicitly captures the numerous and possibly unknown factors that shape the mapping from sequence to function. Statistical models have been used to understand the molecular basis of protein function and provide exceptional predictive accuracy for protein design.
Antibodies have unique specificity, tolerability, and long half-life, which make them ideal therapeutic agents. Over the last 35 years, ever-increasing understanding of the multiple immunological roles that antibodies play has led to creative and diverse targeting modalities for the treatment of human diseases. Approved antibody therapies are now available for cancers, autoimmune disorders, and infectious diseases in formats ranging from small antibody subunits to patient-derived T cells armed with antibody-based targeting molecules. Physiological understanding of specific disease states is required to generate effective antibody-based therapies, and this must be combined with biophysical considerations for the development and manufacture of these complex proteins. Here, we review the latest progress in the field, focusing on antibody-based therapy formats, discovery methods, functional optimization, and production considerations.
In the quest to produce by directed evolution chemo-, stereo-, and regioselective enzymes as catalysts for the sustainable production of chiral compounds needed in modern society, focused saturation mutagenesis at sites surrounding the binding pocket has emerged as a particularly effective strategy. This chapter focuses on the genesis and recent progress of this concept, which has been dubbed combinatorial active-site saturation test (CAST), a procedure that can be performed iteratively (iterative saturation mutagenesis, ISM). The newest developments aimed at increasing efficacy further include guidelines on how to choose hotspots for CAST and how to design reduced amino acid alphabets semirationally, multiparameter evolution aided by mutability landscapes, utilization of the CRISP-Cas9 system, and solid-phase chemical syntheses of saturation mutagenesis libraries on Si-chips for eliminating amino acid bias. The chapter includes a table of recent CAST/ISM papers and highlights in detail three case studies: Limonene epoxide hydrolase in desymmetrization of a prochiral epoxide, alcohol dehydrogenase (TbSADH)-catalyzed enantioselective transformation of difficult-to-reduce ketones, and P450-BM3 as a biocatalyst in whole-cell cascade reactions of cyclohexane with production of ( R,R )-, ( S,S )-, and ( R,S )-1,2-dihydroxycyclohexane.
The adoptive transfer of T cells expressing chimeric antigen receptors (CARs) that target tumor antigens is a highly promising immunotherapy with untapped potential to treat a vast array of cancer types. However, despite remarkable success against B-cell malignancies, clinical outcomes of CAR-T cell therapy against other cancer types have been underwhelming, warranting ongoing research efforts to engineer CARs with optimal antitumor function. Owing to their modularity, CARs can easily be tuned with different target specificities and signaling properties, making CAR engineering highly amenable to design–build–test cycles. Here, we present an overview of each CAR component and their known contributions to CAR function, synthesizing a set of design principles to guide rational engineering efforts for optimal therapeutic efficacy.
Trabajo presentado en el 3rd Multistep Enzyme Catalyzed Processes Congress, celebrado en Madrid (Espana) del 07 al 10 de abril de 2014.
Protein research is critical in fundamental biochemistry, drug development, and biocatalyst applications. High-throughput discovery, characterization, and engineering of proteins are highly desirable, considering the overwhelming availability of genomic information that awaits functional annotation, and the enormous structure–function space that can be explored using protein engineering strategies. While one can now create a large number of engineered proteins using a variety of microbial expression systems, the ability to chemically characterize them in a high throughput manner has lagged behind. Capable of label-free analysis that utilizes native substrates and ligands, mass spectrometry (MS) is well suited for high-throughput assays to characterize complex mixtures of engineered proteins. Here, we review recent advances that improve the throughput and information content of MS-based protein assays and describe notable applications that highlight these capabilities.
The ability to stabilize enzymes and other proteins has wide-ranging applications. Most protocols for enhancing enzyme stability require multiple rounds of high-throughput screening of mutant libraries and provide only modest improvements of stability. Here, we describe a computational library design protocol that can increase enzyme stability by 20-35 °C with little experimental screening, typically fewer than 200 variants. This protocol, termed FRESCO, scans the entire protein structure to identify stabilizing disulfide bonds and point mutations, explores their effect by molecular dynamics simulations, and provides mutant libraries with variants that have a good chance (>10%) to exhibit enhanced stability. After experimental verification, the most effective mutations are combined to produce highly robust enzymes.
Because present-day proteins are composed of 20 kinds of amino acids, the number of possible amino-acid sequences in a 100-residue protein is 20100 (approximately 10130), which is larger than the total number of atoms in the universe (~1080). The number of proteins that may have existed in nature throughout the history of life on the Earth has been estimated to be less than 1050 molecules (Mandecki, 1998) or 1043 molecules (Dryden et al., 2008). Thus, the vast sequence space available remains to be explored further, and the sequence space that remains unexplored provides an opportunity to create valuable proteins with novel structures and functions for biomedical and environmental applications. Evolutionary protein engineering or directed protein evolution has been used to create artificial proteins with novel functions (Bloom et al., 2005; Hoogenboom, 2005; Leemhuis et al., 2005; Romero & Arnold, 2009) by repeated mutation, selection and amplification, mimicking Darwinian evolution in the laboratory (Figure 1).
Shu-Qun Liu1,2 Xing-Lai Ji1,2, Yan Tao1, De-Yong Tan3, Ke-Qin Zhang1 and Yun-Xin Fu1,4 1Laboratory for Conservation and Utilization of Bio-Resources & Key Laboratory for Southwest Biodiversity, Yunnan University, Kunming, 2 Sino-Dutch Biomedial and Information Engineering School, Northeastern University, Shenyang, 3School of Life Sciences, Yunnan University, Kunming, 4Human Genetics Center, School of Public Health, The University of Texas Health Science Center, Houston, Texas, 1,2,3P. R. China 4USA
IntroductionMembrane proteins, constituting ~30% of proteins encoded by whole genomes (Krogh et al., 2001), are heavily implicated in all fundamental cellular processes and, therefore, represent up to 60% of targets for all currently marketed drugs (Overington et al., 2006).Nevertheless, in spite of their significance, only few tens of spatial structures of membrane proteins have been obtained so far, while design of new types of drugs targeting membrane proteins requires precise structural information about this class of objects.Hydrophobic -helices represent a dominant structural motif found in membrane-spanning domains of proteins, excluding membrane -barrels.So, a membrane part a large variety of membrane proteins is formed by -helical bundle (polytopic proteins) or just by single -helix (bitopic proteins) (Fig. 1).Besides structural switching, oligomerization of helical membrane proteins forms the basis for various functions in the living cell including reception of extracellular signals, signal transduction, ion transfer, catalysis, energy conversion and so on (Ubarretxena-Belandia & Engelman, 2001).The mechanisms, by which helical membrane proteins fold into native structures and functionally oligomerize, are beginning to be understood from a confluence of structural and biochemical studies.Folding determinants of a membrane protein can be partially understood by dissecting its structure into pairs of interacting transmembrane (TM) helices, which, together with the connecting loops and extramembrane domains, comprise the overall structure.Obviously, the fold of helical membrane proteins along with their biological activity is largely determined by proper interactions of membrane-embedded helices.Either destroying or enhancing such helix-helix interactions can result in many diseases (developmental, oncogenic, neurodegenerative, immune, cardiovascular, and so on) related to dysfunction of different tissues in the human body.Activity regulation of bitopic proteins that have only single-spanning TM domain is mostly associated with their lateral dimerization in cell membranes.Bitopic proteins are a broad class of biologically significant membrane proteins including the majority of receptor protein kinases, immune receptors and apoptotic proteins, which are involved in development regulation and homeostasis of multicellular organisms.Homo-and heterodimerization of bitopic proteins was earlier thought to involve mostly their extracellular and cytoplasmic domains, but recent studies have been making it increasingly www.intechopen.comProtein Engineering 2 clear that the single-spanning TM domains are also critical for their dimerization and modulation of biological function.Upon bitopic protein activation, ligand-dependent or not, significant intramolecular conformational transitions result in rearrangement of the receptor domains and following receptor dimerization or switching from one dimerization state to another, e.g.ligand-dependent transition from preformed inactive dimeric state into active dimer of ErbB receptor tyrosine kinase (Schlessinger, 2000;Moriki et al., 2001;Fleishman et al., 2002;Mendrola et al., 2002).The so-called "rotation-coupled" and "flexible rotation" activation mechanisms (Moriki et al., 2001;Fleishman et al., 2002;Mendrola et al., 2002), which were initially proposed for receptor tyrosine kinases and imply active involvement of TM domains in dimerization and activation of the receptors via proper TM helix-helix packing and rearranging, are possibly widespread among bitopic proteins.However, if biological functions are carried out using only one homo-or heterodimeric state of bitopic protein TM domains, the TM helix-helix interaction can be strong, as in the case of permeabilization of the outer mitochondrial membrane by proapoptotic protein BNip3 in the course of hypoxia-acidosis induced cell death.Furthermore, amino acid polymorphisms and mutations in the TM domain of bitopic proteins have been implicated in numerous human pathological states, including many types of cancers, Alzheimer's disease, tissue dysplasias and abnormalities (Li & Hristova, 2006;Selkoe, 2001).It was shown that the mutations affect both the behavior of the isolated TM domains in model lipid bilayers, and the behavior of the full length receptors in the plasma membrane.Most probably, the effects are exerted via yet unknown mutation-induced changes in dimeric structure of the TM domains.Importantly, it was found that isolated TM domains revealed ability not only to homo-and heterodimerize in membrane-like environment, but also to specifically inhibit biological activity of bitopic proteins in cell membrane (Li & Hristova, 2006;Bennasroune et al., 2004;Rath et al., 2007).So, membrane-spanning segments of bitopic proteins represent a novel class of pharmacologically important targets, whose activity can be modulated by natural or specially designed molecules.Among the most perspective candidates for these purposes are artificial hydrophobic helical peptides, the so-called peptide "interceptors" (Bennasroune et al., 2004) or "computer helical antimembrane proteins" (CHAMPs) (Caputo et al., 2008), which are capable of specifically recognizing the target wild-type TM segments of bitopic proteins and interfering with their lateral association in cell membrane.Therefore, understanding the factors that drive packing of -helices in membranes has attracted considerable interest of researchers from both scientific and medical communities.Nevertheless, in spite of their significance, only few spatial structures of the homo-and heterodimeric single-span TM domains have been obtained so far, notwithstanding that design of new types of drugs targeting bitopic proteins requires precise structural information about this class of objects.At the present stage of development of the structural biology methods, obtaining highresolution structure of a full-length bitopic protein is a scientific challenge.Issues with crystallization of membrane proteins are inherent to X-ray techniques, whereas NMR cannot effectively handle large protein-lipid complexes.The crystallographic methods, which recently allowed obtaining high-resolution structure of such multi-span TM receptors as Gprotein coupled receptors (Cherezov et al., 2007), cannot be directly translated to multipledomains flexible receptors like receptor kinases and immune system receptors.Therefore, www.intechopen.
The utilization of ethanol produced from plant biomass (so-called “bioethanol”), which is derived from the fixation of atmospheric CO2, as an industrial carbon source and car fuel is one of the most important research issues for the realization of a sustainable global environment. Currently, bioethanol is produced mainly from agricultural crop biomass, which, however, compete with food and animal feed. Alternatively, “lignocellulosic biomass”, such as woods and agricultural residues, is an attractive feedstock, and consists of cellulose, hemicellulose, and lignin. In agricultural crops, corn for example, starch, consisting of cellulose (hexose polymers), accounts for 70% of the mass, and its biological fermentation is easy. On the other hand, in lignocellulosic biomass, hemicellulose, accounting for 30% of the mass, is composed of pentoses as well as hexoses such as xylose (and L-arabinose), which cannot be fermented by natural microorganisms.