Native ion channels play key roles in biological systems, and engineered versions are widely used as chemogenetic tools and in sensing devices1,2. Protein design has been harnessed to generate pore-containing transmembrane proteins, but the design of selectivity filters with precise arrangements of amino acid side chains specific for a target ion, a crucial feature of native ion channels3, has been constrained by the lack of methods for placing the metal-coordinating residues with atomic-level precision. Here we describe a bottom-up RFdiffusion-based approach to construct Ca2+ channels from defined selectivity filter residue geometries, and use this approach to design symmetric oligomeric channels with Ca2+ selectivity filters having different coordination numbers and different geometries at the entrance of a wider pore buttressed by multiple transmembrane helices. The designed channel proteins assemble into homogeneous pore-containing particles and, for both tetrameric and hexameric ion-coordinating configurations, patch-clamp experiments show that the designed channels have higher conductances for Ca2+ than for Na+ and other divalent ions (Sr2+ and Mg2+) that are eliminated after mutation of selectivity filter residues. Cryogenic electron microscopy indicates that the design method has high accuracy: the structure of the hexameric Ca2+ channel is nearly identical to that of the design model. Our bottom-up design approach now enables the testing of hypotheses relating filter geometry to ion selectivity by direct construction, and provides a roadmap for creating selective ion channels for a wide range of applications.
Recent advances in computational methods have led to considerable progress in the design of self-assembling protein nanoparticles. However, nearly all nanoparticles designed to date exhibit strict point group symmetry, with each subunit occupying an identical, symmetrically related environment. This limits the structural diversity that can be achieved and precludes anisotropic functionalization. Here, we describe a general computational strategy for designing multi-component bifaceted protein nanomaterials with two distinctly addressable sides. The method centers on docking pseudosymmetric heterooligomeric building blocks in architectures with dihedral symmetry and designing an asymmetric protein-protein interface between them. We used this approach to obtain an initial 30-subunit assembly with pseudo-D5 symmetry, and then generated an additional 15 variants in which we controllably altered the size and morphology of the bifaceted nanoparticles by designing de novo extensions to one of the subunits. Functionalization of the two distinct faces of the nanoparticles with de novo protein minibinders enabled specific colocalization of two populations of polystyrene microparticles coated with target protein receptors. The ability to accurately design anisotropic protein nanomaterials with precisely tunable structures and functions could be broadly useful in applications that require colocalizing two or more distinct target moieties.
AbstractWhile direct cell transplantation holds great promise in treating many debilitating diseases, poor cell survival and engraftment following injection have limited effective clinical translation. Though injectable biomaterials offer protection against membrane‐damaging extensional flow and supply a supportive 3D environment in vivo that ultimately improves cell retention and therapeutic costs, most are created from synthetic or naturally harvested polymers that are immunogenic and/or chemically ill‐defined. This work presents a shear‐thinning and self‐healing telechelic recombinant protein‐based hydrogel designed around XTEN – a well‐expressible, non‐immunogenic, and intrinsically disordered polypeptide previously evolved as a genetically encoded alternative to PEGylation to “eXTENd” the in vivo half‐life of fused protein therapeutics. By flanking XTEN with self‐associating coil domains derived from cartilage oligomeric matrix protein, single‐component physically crosslinked hydrogels exhibiting rapid shear thinning and self‐healing through homopentameric coiled‐coil bundling are formed. Individual and combined point mutations that variably stabilize coil association enables a straightforward method to genetically program material viscoelasticity and biodegradability. Finally, these materials protect and sustain viability of encapsulated human fibroblasts, hepatocytes, embryonic kidney (HEK), and embryonic stem‐cell‐derived cardiomyocytes (hESC‐CMs) through culture, injection, and transcutaneous implantation in mice. These injectable XTEN‐based hydrogels show promise for both in vitro cell culture and in vivo cell transplantation applications.
We describe a modular bond-centric approach to protein nanomaterial design inspired by the rich diversity of chemical structures that can be generated from the small number of atomic valencies and bonding interactions. We design protein building blocks with regular coordination geometries and bonding interactions that enable the assembly of a wide variety of closed and opened nanomaterials using simple geometrical principles. Experimental characterization confirms successful formation of more than twenty multi-component polyhedral protein cages, 2D arrays, and 3D protein lattices, with a high (10-50 %) success rate and electron microscopy data closely matching the corresponding design models. Because of the modularity, individual building blocks can assemble with different partners to generate distinct regular assemblies, resulting in an economy of parts and enabling the construction of reconfigurable systems.
Four, eight or twenty C3 symmetric protein trimers can be arranged with tetrahedral, octahedral or icosahedral point group symmetry to generate closed cage-like structures1,2. Viruses access more complex higher triangulation number icosahedral architectures by breaking perfect point group symmetry3-9, but nature appears not to have explored similar symmetry breaking for tetrahedral or octahedral symmetries. Here we describe a general design strategy for building higher triangulation number architectures starting from regular polyhedra through pseudosymmetrization of trimeric building blocks. Electron microscopy confirms the structures of T = 4 cages with 48 (tetrahedral), 96 (octahedral) and 240 (icosahedral) subunits, each with 4 distinct chains and 6 different protein-protein interfaces, and diameters of 33 nm, 43 nm and 75 nm, respectively. Higher triangulation number viruses possess very sophisticated functionalities; our general route to higher triangulation number nanocages should similarly enable a next generation of multiple antigen-displaying vaccine candidates10,11 and targeted delivery vehicles12,13.
Design models for the associated publication
A wooden house frame consists of many different lumber pieces, but because of the regularity of these building blocks, the structure can be designed using straightforward geometrical principles. The design of multicomponent protein assemblies, in comparison, has been much more complex, largely owing to the irregular shapes of protein structures1. Here we describe extendable linear, curved and angled protein building blocks, as well as inter-block interactions, that conform to specified geometric standards; assemblies designed using these blocks inherit their extendability and regular interaction surfaces, enabling them to be expanded or contracted by varying the number of modules, and reinforced with secondary struts. Using X-ray crystallography and electron microscopy, we validate nanomaterial designs ranging from simple polygonal and circular oligomers that can be concentrically nested, up to large polyhedral nanocages and unbounded straight 'train track' assemblies with reconfigurable sizes and geometries that can be readily blueprinted. Because of the complexity of protein structures and sequence-structure relationships, it has not previously been possible to build up large protein assemblies by deliberate placement of protein backbones onto a blank three-dimensional canvas; the simplicity and geometric regularity of our design platform now enables construction of protein nanomaterials according to 'back of an envelope' architectural blueprints.
Molecular systems with coincident cyclic and superhelical symmetry axes have considerable advantages for materials design as they can be readily lengthened or shortened by changing the length of the constituent monomers. Among proteins, alpha-helical coiled coils have such symmetric, extendable architectures, but are limited by the relatively fixed geometry and flexibility of the helical protomers. Here we describe a systematic approach to generating modular and rigid repeat protein oligomers with coincident C 2 to C 8 and superhelical symmetry axes that can be readily extended by repeat propagation. From these building blocks, we demonstrate that a wide range of unbounded fibres can be systematically designed by introducing hydrophilic surface patches that force staggering of the monomers; the geometry of such fibres can be precisely tuned by varying the number of repeat units in the monomer and the placement of the hydrophilic patches.
Four, eight or twenty C3 symmetric protein trimers can be arranged with tetrahedral (T-sym), octahedral (O-sym) or icosahedral (I-sym) point group symmetry to generate closed cage-like structures 1,2 . Generating more complex closed structures requires breaking perfect point group symmetry. Viruses do this in the icosahedral case using quasi-symmetry or pseudo-symmetry to access higher triangulation number architectures 3–9 , but nature appears not to have explored higher triangulation number tetrahedral or octahedral symmetries. Here, we describe a general design strategy for building T = 4 architectures starting from simpler T = 1 structures through pseudo-symmetrization of trimeric building blocks. Electron microscopy confirms the structures of T = 4 cages with 48 (T-sym), 96 (O-sym), and 240 (I-sym) subunits, each with four distinct chains and six different protein-protein interfaces, and diameters of 33nm, 43nm, and 75nm, respectively. Higher triangulation number viruses possess very sophisticated functionalities; our general route to higher triangulation number nanocages should similarly enable a next generation of multiple antigen displaying vaccine candidates 10,11 and targeted delivery vehicles 12,13 .
De novo protein design methods can create proteins with folds not yet seen in nature. These methods largely focus on optimizing the compatibility between the designed sequence and the intended conformation, without explicit consideration of protein folding pathways. Deeply knotted proteins, whose topologies may introduce substantial barriers to folding, thus represent an interesting test case for protein design. Here we report our attempts to design proteins with trefoil (3(1)) and pentafoil (5(1)) knotted topologies. We extended previously described algorithms for tandem repeat protein design in order to construct deeply knotted backbones and matching designed repeat sequences (N = 3 repeats for the trefoil and N = 5 for the pentafoil). We confirmed the intended conformation for the trefoil design by X ray crystallography, and we report here on this protein's structure, stability, and folding behaviour. The pentafoil design misfolded into an asymmetric structure (despite a 5-fold symmetric sequence); two of the four repeat-repeat units matched the designed backbone while the other two diverged to form local contacts, leading to a trefoil rather than pentafoil knotted topology. Our results also provide insights into the folding of knotted proteins.
Engineered proteins with precisely defined shapes can scaffold functional protein domains in 3D space to fine-tune their functions, such as the regulation of cellular signaling by ligand positioning or the design of self-assembling protein materials with specific forms. Methods for simply and efficiently generating the protein backbones to initiate these design processes remain limited. In this work, we develop a lightweight neural network to guide helix fragment assembly along a guideline using a GAN architecture and show that this approach can rapidly generate viable samples while being computationally inexpensive. Key to our approach is the transformation of the input structural data used for training into a parametric representation of helices to reduce the generator network size, which in turn facilitates rapid backpropagation to find specific helical arrangements during generation. This approach provides a method to quickly generate helical protein scaffolds.
Deep learning generative approaches provide an opportunity to broadly explore protein structure space beyond the sequences and structures of natural proteins. Here, we use deep network hallucination to generate a wide range of symmetric protein homo-oligomers given only a specification of the number of protomers and the protomer length. Crystal structures of seven designs are very similar to the computational models (median root mean square deviation: 0.6 angstroms), as are three cryo-electron microscopy structures of giant 10-nanometer rings with up to 1550 residues and C33 symmetry; all differ considerably from previously solved structures. Our results highlight the rich diversity of new protein structures that can be generated using deep learning and pave the way for the design of increasingly complex components for nanomachines and biomaterials.
Dynamic dimerization is a common regulatory interaction between biological molecules, underpinning many signaling functions. Because of its ubiquity, many biological engineering efforts have focused on building dimerizing proteins, such as the SYNZIPs and de novo Designed HeteroDimers (DHDs). Using the DHDs as a model system, we show that low-affinity protein interactions can be competitively displaced by a high-affinity "dominant negative" heterodimer. We demonstrate the utility of this signaling motif by using competitive displacement to implement negative feedback in a synthetic circuit. Competitive displacement could be extended to other heterodimer systems to expand the functionality of protein circuits and enable new biotechnology applications.
Pullulanases are glycoside hydrolase family 13 (GH13) enzymes that target alpha 1,6 glucosidic linkages within starch and aid in the degradation of the alpha 1,4- and alpha 1,6- linked glucans pullulan, glycogen and amylopectin. The human gut bacterium Ruminococcus bromii synthesizes two extracellular pullulanases, Amy10 and Amy12, that are incorporated into the multiprotein amylosome complex that enables the digestion of granular resistant starch from the diet. Here we provide a comparative biochemical analysis of these pullulanases and the x-ray crystal structures of the wild type and the nucleophile mutant D392A of Amy12 complexed with maltoheptaose and 63 alpha-D glucosyl-maltotriose. While Amy10 displays higher catalytic efficiency on pullulan and cleaves only alpha 1,6 linkages, Amy12 has some activity on alpha 1,4 linkages suggesting that these enzymes are not redundant within the amylosome. Our structures of Amy12 include a mucin-binding protein (MucBP) domain that follows the Cdomain of the GH13 fold, an atypical feature of these enzymes. The wild type Amy12 structure with maltoheptaose captured two oligosaccharides in the active site arranged as expected following catalysis of an alpha 1,6 branch point in amylopectin. The nucleophile mutant D392A complexed with maltoheptaose or 63-alpha-D glucosylmaltotriose captured beta-glucose at the reducing end in the -1 subsite, facilitated by the truncation of the active site aspartate and stabilized by stacking with Y279. The core interface between the co-crystallized ligands and Amy12 occurs within the -2 through + 1 subsites, which may allow for flexible recognition of alpha 1,6 linkages within a variety of starch structures.
The design of modular protein logic for regulating protein function at the posttranscriptional level is a challenge for synthetic biology. Here, we describe the design of two-input AND, OR, NAND, NOR, XNOR, and NOT gates built from de novo-designed proteins. These gates regulate the association of arbitrary protein units ranging from split enzymes to transcriptional machinery in vitro, in yeast and in primary human T cells, where they control the expression of the TIM3 gene related to T cell exhaustion. Designed binding interaction cooperativity, confirmed by native mass spectrometry, makes the gates largely insensitive to stoichiometric imbalances in the inputs, and the modularity of the approach enables ready extension to three-input OR, AND, and disjunctive normal form gates. The modularity and cooperativity of the control elements, coupled with the ability to de novo design an essentially unlimited number of protein components, should enable the design of sophisticated posttranslational control logic over a wide range of biological functions.
Single-cell RNA sequencing (scRNA-seq) has become an essential tool for characterizing gene expression in eukaryotes, but current methods are incompatible with bacteria. Here, we introduce microSPLiT (microbial split-pool ligation transcriptomics), a high-throughput scRNA-seq method for Gram-negative and Gram-positive bacteria that can resolve heterogeneous transcriptional states. We applied microSPLiT to >25,000 Bacillus subtilis cells sampled at different growth stages, creating an atlas of changes in metabolism and lifestyle. We retrieved detailed gene expression profiles associated with known, but rare, states such as competence and prophage induction and also identified unexpected gene expression states, including the heterogeneous activation of a niche metabolic pathway in a subpopulation of cells. MicroSPLiT paves the way to high-throughput analysis of gene expression in bacterial communities that are otherwise not amenable to single-cell analysis, such as natural microbiota.
SummaryRuminococcus bromii is a dominant member of the human colonic microbiota that plays a ‘keystone’ role in degrading dietary resistant starch. Recent evidence from one strain has uncovered a unique cell surface ‘amylosome’ complex that organizes starch‐degrading enzymes. New genome analysis presented here reveals further features of this complex and shows remarkable conservation of amylosome components between human colonic strains from three different continents and a R. bromii strain from the rumen of Australian cattle. These R. bromii strains encode a narrow spectrum of carbohydrate active enzymes (CAZymes) that reflect extreme specialization in starch utilization. Starch hydrolysis products are taken up mainly as oligosaccharides, with only one strain able to grow on glucose. The human strains, but not the rumen strain, also possess transporters that allow growth on galactose and fructose. R. bromii strains possess a full complement of sporulation and spore germination genes and we demonstrate the ability to form spores that survive exposure to air. Spore formation is likely to be a critical factor in the ecology of this nutritionally highly specialized bacterium, which was previously regarded as ‘non‐sporing’, helping to explain its widespread occurrence in the gut microbiota through the ability to transmit between hosts.