A "chemical linearization" approach was applied to synthetic peptide macrocycles to enable their de novo sequencing from mixtures using nanoliquid chromatography-tandem mass spectrometry (nLC-MS/MS). This approach─previously applied to individual macrocycles but not to mixtures─involves cleavage of the peptide backbone at a defined position to give a product capable of generating sequence-determining fragment ions. Here, we first established the compatibility of "chemical linearization" by Edman degradation with a prominent macrocycle scaffold based on bis-Cys peptides cross-linked with the m-xylene linker, which are of major significance in therapeutics discovery. Then, using macrocycle libraries of known sequence composition, the ability to recover accurate de novo assignments to linearized products was critically tested using performance metrics unique to mixtures. Significantly, we show that linearized macrocycles can be sequenced with lower recall compared to linear peptides but with similar accuracy, which establishes the potential of using "chemical linearization" with synthetic libraries and selection procedures that yield compound mixtures. Sodiated precursor ions were identified as a significant source of high-scoring but inaccurate assignments, with potential implications for improving automated de novo sequencing more generally.
Establishing structure-activity relationships is crucial to understand and optimize the activity of peptide-based inhibitors of protein-protein interactions. Single alanine substitutions provide limited information on the residues that tolerate simultaneous modifications with retention of biological activity. To guide optimization of peptide binders, we use combinatorial peptide libraries of over 4,000 variants-in which each position is varied with either the wild-type residue or alanine-with a label-free affinity selection platform to study protein-ligand interactions. Applying this platform to a peptide binder to the oncogenic protein MDM2, several multi-alanine-substituted analogs with picomolar binding affinity were discovered. We reveal a non-additive substitution pattern in the selected sequences. The alanine substitution tolerances for peptide ligands of the 12ca5 antibody and 14-3-3 regulatory protein are also characterized, demonstrating the general applicability of this new platform. We envision that binary combinatorial alanine scanning will be a powerful tool for investigating structure-activity relationships.
An attractive approach to target intracellular macromolecular interfaces and to model putative drug interactions is to design small high-affinity proteins. Variable domains of the immunoglobulin heavy chain (VH domains) are ideal miniproteins, but their development has been restricted by poor intracellular stability and expression. Here we show that an autonomous and disufhide-free VH domain is suitable for intracellular studies and use it to construct a high-diversity phage display library. Using this library and affinity maturation techniques we identify VH domains with picomolar affinity against eIF4E, a protein commonly hyper-activated in cancer. We demonstrate that these molecules interact with eIF4E at the eIF4G binding site via a distinct structural pose. Intracellular overexpression of these miniproteins reduce cellular proliferation and expression of malignancy-related proteins in cancer cell lines. The linkage of high-diversity in vitro libraries with an intracellularly expressible miniprotein scaffold will facilitate the discovery of VH domains suitable for intracellular applications.
Establishing structure–activity relationships iscrucial to understand and optimize the activity of peptide-based inhibitors of protein–protein interactions. Single alanine mutagenesis provides limited information toward thisgoal. To guide multiple simultaneous peptide modifications with retention ofbiological activity, we used synthetic combinatorial alanine-scanninglibraries—in which each position was varied with either the wild type residueor alanine—with an affinity selection platform to study the mutational toleranceof protein–ligand interactions. Applying this platform to a peptide binder tothe oncogenic protein MDM2, several multi-alanine-substituted analogs thatretained low nanomolar affinity were discovered, including a 13-mer binder withseven alanine substitutions at non-hotspot positions. These binders served as templatesfor further modifications, generating cysteine-substituted, perfluoroaryl-stapledpeptides with sub-nanomolar affinity and ten-fold improved proteolyticstability. The alanine substitution tolerances for peptide ligands of the 12ca5antibody and 14-3-3 regulatory protein were also reported, demonstrating thegeneral applicability of this new platform. We envision that deep combinatorialalanine scanning will be a powerful tool for structure–activity optimization ofpotential peptide therapeutics.
Nature has three biopolymers: oligonucleotides, polypeptides, and oligosaccharides. Each biopolymer has independent functions, but when needed, they form mixed assemblies for higher-order purposes, as in the case of ribosomal protein synthesis. Rather than forming large complexes to coordinate the role of different biopolymers, we dovetail protein amino acids and nucleobases into a single low molecular weight precision polyamide polymer. We established efficient chemical synthesis and de novo sequencing procedures and prepared combinatorial libraries with up to 100 million biohybrid molecules. This biohybrid material has a higher bulk affinity to oligonucleotides than peptides composed exclusively of canonical amino acids. Using affinity selection mass spectrometry, we discovered variants with a high affinity for pre-microRNA hairpins. Our platform points toward the development of high throughput discovery of sequence defined polymers with designer properties, such as oligonucleotide binding.
High-diversity genetically-encoded combinatorial libraries (10 8 −10 13 members) are a rich source of peptide-based binding molecules, identified by affinity selection. Synthetic libraries can access broader chemical space, but typically examine only ~ 10 6 compounds by screening. Here we show that in-solution affinity selection can be interfaced with nano-liquid chromatography-tandem mass spectrometry peptide sequencing to identify binders from fully randomized synthetic libraries of 10 8 members—a 100-fold gain in diversity over standard practice. To validate this approach, we show that binders to a monoclonal antibody are identified in proportion to library diversity, as diversity is increased from 10 6 –10 8 . These results are then applied to the discovery of p53-like binders to MDM2, and to a family of 3–19 nM-affinity, α/β-peptide-based binders to 14-3-3. An X-ray structure of one of these binders in complex with 14-3-3σ is determined, illustrating the role of β-amino acids in facilitating a key binding contact.
Flow‐based approaches to solid phase peptide synthesis (SPPS) have been pursued since the method's early days, with anticipated gains in speed, reaction monitoring, and ease of automation. Here, we discuss how these advantages are being realized by synthesis at elevated temperature, facilitated by a 'preheat/activation' loop. This important modification both accelerates peptide synthesis—providing a wealth of new data from in‐line monitoring—and in conjunction with an optimized protocol, extends the length of peptides routinely accessible by stepwise synthesis. Streamlined synthesis of longer peptides will address a major challenge in chemical protein synthesis: preparation of peptide segments for use in chemical ligation. The context of recent results in flow‐based SPPS is discussed, highlighting remaining challenges and future opportunities.
Ribosomes produce most proteins of living cells in seconds. Here we report highly efficient chemistry matched with an automated fast-flow instrument for the direct manufacturing of peptide chains up to 164 amino acids over 328 consecutive reactions. The machine is rapid - the peptide chain elongation is complete in hours. We demonstrate the utility of this approach by the chemical synthesis of nine different protein chains that represent enzymes, structural units, and regulatory factors. After purification and folding, the synthetic materials display biophysical and enzymatic properties comparable to the biologically expressed proteins. High-fidelity automated flow chemistry is an alternative for producing single-domain proteins without the ribosome.
In the version of this article originally published, the peptide sequences of compounds 90, 92 and 93 in Fig. 5b and Supplementary Table 7 contained several errors. In Fig. 5b, position 6 of compound 90 should be Tyr instead of Phe. In both Fig. 5b and Supplementary Table 7, position 9 of compounds 92 and 93 should be Gln instead of Glu. Additionally, the surname of co-author Anupam Bandyopadhyay was incorrectly spelled as Bandyopdhyay. The errors have been corrected in the HTML and PDF versions of the paper and in the Supplementary Information PDF.
A 29-residue peptide (MP01), identified by in vitro selection for reactivity with a small molecule perfluoroaromatic, was modified and characterized using experimental and computational techniques, with the goal of understanding the molecular basis of its reactivity. These studies identified a six-amino acid point mutant (MP01-Gen4) that exhibited a reaction rate constant of 25.8 ± 1.8 M-1 s-1 at pH 7.4 and room temperature, approximately 2 orders of magnitude greater than that of its progenitor sequence and 3 orders of magnitude greater than background cysteine reactivity. MP01-Gen4 appeared to be conformationally dynamic and exhibited several properties reminiscent of larger protein molecules, including denaturant-sensitive structure and reactivity. We believe the majority of the reaction rate enhancement can be attributed to interaction of MP01-Gen4 with the perfluoroaromatic probe, which was found to stabilize a helical conformation of both MP01-Gen4 and nonreactive Cys-to-Ser or Cys-to-Ala variants. These findings demonstrate the ability of dynamic peptides to access proteinlike reaction mechanisms and the potential of perfluoroaromatic functionality to stabilize small peptide folds.
We report a site-selective cysteine-cyclooctyne conjugation reaction between a seven-residue peptide tag (DBCO-tag, Leu-Cys-Tyr-Pro-Trp-Val-Tyr) at the N or C terminus of a peptide or protein and various aza-dibenzocyclooctyne (DBCO) reagents. Compared to a cysteine peptide control, the DBCO-tag increases the rate of the thiol-yne reaction 220-fold, thereby enabling selective conjugation of DBCO-tag to DBCO-linked fluorescent probes, affinity tags, and cytotoxic drug molecules. Fusion of DBCO-tag with the protein of interest enables regioselective cysteine modification on proteins that contain multiple endogenous cysteines; these examples include green fluorescent protein and the antibody trastuzumab. This study demonstrates that short peptide tags can aid in accelerating bond-forming reactions that are often slow to non-existent in water.
The facile rearrangement of "S-acyl isopeptides" to native peptide bonds via S,N-acyl shift is central to the success of native chemical ligation, the widely used approach for protein total synthesis. Proximity-driven amide bond formation via acyl transfer reactions in other contexts has proven generally less effective. Here, we show that under neutral aqueous conditions, "O-acyl isopeptides" derived from hydroxy-asparagine [aspartic acid-β-hydroxamic acid; Asp(β-HA)] rearrange to form native peptide bonds via an O,N-acyl shift. This process constitutes a rare example of an O,N-acyl shift that proceeds rapidly across a medium-size ring (t1/2 ∼ 15 min), and takes place in water with minimal interference from hydrolysis. In contrast to serine/threonine or tyrosine, which form O-acyl isopeptides only by the use of highly activated acyl donors and appropriate protecting groups in organic solvent, Asp(β-HA) is sufficiently reactive to form O-acyl isopeptides by treatment with an unprotected peptide-αthioester, at low mM concentration, in water. These findings were applied to an acyl transfer-based chemical ligation strategy, in which an unprotected N-terminal Asp(β-HA)-peptide and peptide-αthioester react under aqueous conditions to give a ligation product ultimately linked by a native peptide bond.
Significance Combinatorial protein libraries—prepared via molecular biology-based approaches—are invaluable tools for protein engineering. The inclusion of noncanonical amino acids in such libraries is of considerable interest. However, at present no approach competes with chemical synthesis in terms of the variety and number of noncanonical amino acids that can be simultaneously incorporated into a protein molecule. Here, we describe selection from synthetic libraries as a strategy for protein engineering. The approach enables identification of small (∼30 aa), functional protein variants comprising a virtually unlimited variety of noncanonical amino acids. Increasing the throughput of synthetic library screening, which was achieved through this effort, is anticipated to improve the utility of synthetic libraries for identifying polypeptide-based ligands with de novo function.
A methodology to achieve high-throughput de novo sequencing of synthetic peptide mixtures is reported. The approach leverages shotgun nanoliquid chromatography coupled with tandem mass spectrometry-based de novo sequencing of library mixtures (up to 2000 peptides) as well as automated data analysis protocols to filter away incorrect assignments, noise, and synthetic side-products. For increasing the confidence in the sequencing results, mass spectrometry-friendly library designs were developed that enabled unambiguous decoding of up to 600 peptide sequences per hour while maintaining greater than 85% sequence identification rates in most cases. The reliability of the reported decoding strategy was additionally confirmed by matching fragmentation spectra for select authentic peptides identified from library sequencing samples. The methods reported here are directly applicable to screening techniques that yield mixtures of active compounds, including particle sorting of one-bead one-compound libraries and affinity enrichment of synthetic library mixtures performed in solution.
The burial of hydrophobic side chains in a protein core generally is thought to be the major ingredient for stable, cooperative folding. Here, we show that, for the snow flea antifreeze protein (sfAFP), stability and cooperativity can occur without a hydrophobic core, and without α-helices or β-sheets. sfAFP has low sequence complexity with 46% glycine and an interior filled only with backbone H-bonds between six polyproline 2 (PP2) helices. However, the protein folds in a kinetically two-state manner and is moderately stable at room temperature. We believe that a major part of the stability arises from the unusual match between residue-level PP2 dihedral angle bias in the unfolded state and PP2 helical structure in the native state. Additional stabilizing factors that compensate for the dearth of hydrophobic burial include shorter and stronger H-bonds, and increased entropy in the folded state. These results extend our understanding of the origins of cooperativity and stability in protein folding, including the balance between solvent and polypeptide chain entropies.
Under suitable conditions, trifluoromethanesulfonic acid performs comparably to hydrogen fluoride for the on-resin global deprotection of peptides prepared by Boc chemistry solid phase peptide synthesis (SPPS). Obviation of hydrogen fluoride in Boc chemistry SPPS enables the straightforward synthesis of peptide-αthioesters for use in native chemical ligation.
Efficient total synthesis of insulin is important to enable the application of medicinal chemistry to the optimization of the properties of this important protein molecule. Recently we described "ester insulin"--a novel form of insulin in which the function of the 35 residue C-peptide of proinsulin is replaced by a single covalent bond--as a key intermediate for the efficient total synthesis of insulin. Here we describe a fully convergent synthetic route to the ester insulin molecule from three unprotected peptide segments of approximately equal size. The synthetic ester insulin polypeptide chain folded much more rapidly than proinsulin, and at physiological pH. Both the D-protein and L-protein enantiomers of monomeric DKP ester insulin (i.e., [Asp(B10), Lys(B28), Pro(B29)]ester insulin) were prepared by total chemical synthesis. The atomic structure of the synthetic ester insulin molecule was determined by racemic protein X-ray crystallography to a resolution of 1.6 Å. Diffraction quality crystals were readily obtained from the racemic mixture of {D-DKP ester insulin + L-DKP ester insulin}, whereas crystals were not obtained from the L-ester insulin alone even after extensive trials. Both the D-protein and L-protein enantiomers of monomeric DKP ester insulin were assayed for receptor binding and in diabetic rats, before and after conversion by saponification to the corresponding DKP insulin enantiomers. L-DKP ester insulin bound weakly to the insulin receptor, while synthetic L-DKP insulin derived from the L-DKP ester insulin intermediate was fully active in binding to the insulin receptor. The D- and L-DKP ester insulins and D-DKP insulin were inactive in lowering blood glucose in diabetic rats, while synthetic L-DKP insulin was fully active in this biological assay. The structural basis of the lack of biological activity of ester insulin is discussed.
Original synthetic and structure determination methods were used to make a protein molecule with an unprecedented linear-loop polypeptide chain topology, and to characterize its X-ray structure.