The design of proteins with tailored functions is of immense interest to biotechnology, medicine, and the chemical industry. While protein design is rapidly evolving with the use of AI techniques, the design of complex enzymes remains a challenge. Here, we present the use of two large language models (LLMs), ZymCTRL and ProtGPT2, for the generation of de novo enzymes that catalyze the triosephosphate isomerase (TIM) reaction. Natural TIM enzymes are obligatory oligomers that catalyze a multi-step isomerization reaction near the diffusion limit. This makes TIM an ideal target to assess the generative ability of protein language models. Newly generated sequences were filtered to obtain a set of twelve candidates from each approach for experimental validation. Multiple constructs from both language models exhibit the intended function in vivo through their ability to complement a TIM-deficient E. coli strain. In-depth characterization of the best-behaving artificial enzyme reveals behavior and catalytic efficiency close to its natural counterparts. These findings support the use of conditional and fine-tuned unconditional LLMs for the generation of complex enzymes. ### Competing Interest Statement The authors have declared no competing interest.
The ability to design customized proteins to perform specific tasks is of great interest. We are particularly interested in the design of sensitive and specific small molecule ligand‐binding proteins for biotechnological or biomedical applications. Computational methods can narrow down the immense combinatorial space to find the best solution and thus provide starting points for experimental procedures. However, success rates strongly depend on accurate modeling and energetic evaluation. Not only intra‐ but also intermolecular interactions have to be considered. To address this problem, we developed PocketOptimizer, a modular computational protein design pipeline, that predicts mutations in the binding pockets of proteins to increase affinity for a specific ligand. Its modularity enables users to compare different combinations of force fields, rotamer libraries, and scoring functions. Here, we present a much‐improved version––PocketOptimizer 2.0. We implemented a cleaner user interface, an extended architecture with more supported tools, such as force fields and scoring functions, a backbone‐dependent rotamer library, as well as different improvements in the underlying algorithms. Version 2.0 was tested against a benchmark of design cases and assessed in comparison to the first version. Our results show how newly implemented features such as the new rotamer library can lead to improved prediction accuracy. Therefore, we believe that PocketOptimizer 2.0, with its many new and improved functionalities, provides a robust and versatile environment for the design of small molecule‐binding pockets in proteins. It is widely applicable and extendible due to its modular framework. PocketOptimizer 2.0 can be downloaded at https://github.com/Hoecker-Lab/pocketoptimizer .
Protein design aims to build novel proteins customized for specific purposes, thereby holding the potential to tackle many environmental and biomedical problems. Recent progress in Transformer-based architectures has enabled the implementation of language models capable of generating text with human-like capabilities. Here, motivated by this success, we describe ProtGPT2, a language model trained on the protein space that generates de novo protein sequences following the principles of natural ones. The generated proteins display natural amino acid propensities, while disorder predictions indicate that 88% of ProtGPT2-generated proteins are globular, in line with natural sequences. Sensitive sequence searches in protein databases show that ProtGPT2 sequences are distantly related to natural ones, and similarity networks further demonstrate that ProtGPT2 is sampling unexplored regions of protein space. AlphaFold prediction of ProtGPT2-sequences yields well-folded non-idealized structures with embodiments and large loops and reveals topologies not captured in current structure databases. ProtGPT2 generates sequences in a matter of seconds and is freely available.
The experimental characterization and computational prediction of protein structures has become increasingly rapid and precise. However, the analysis of protein structures often requires researchers to use several software packages or web servers, which complicates matters. To provide long-established structural analyses in a modern, easy-to-use interface, we implemented ProteinTools, a web server toolkit for protein structure analysis. ProteinTools gathers four applications so far, namely the identification of hydrophobic clusters, hydrogen bond networks, salt bridges, and contact maps. In all cases, the input data is a PDB identifier or an uploaded structure, whereas the output is an interactive dynamic web interface. Thanks to the modular nature of ProteinTools, the addition of new applications will become an easy task. Given the current need to have these tools in a single, fast, and interpretable interface, we believe that ProteinTools will become an essential toolkit for the wider protein research community. The web server is available at https://proteintools.uni-bayreuth.de.
Modern proteins have been shown to share evolutionary relationships via subdomain-sized fragments. The assembly of such fragments through duplication and recombination events led to the complex structures and functions we observe today. We previously implemented a pipeline that identified more than 1,000 of these fragments that are shared by different protein folds and developed a web interface to analyze and search for them. This resource named Fuzzle helps structural and evolutionary biologists to identify and analyze conserved parts of a protein but it also provides protein engineers with building blocks for example to design proteins by fragment combination. Here, we describe a new version of this web resource that was extended to include ligand information. This addition is a significant asset to the database since now protein fragments that bind specific ligands can be identified and analyzed. Often the mode of ligand binding is conserved in proteins thereby supporting a common evolutionary origin. The same can now be explored for subdomain-sized fragments within this database. This ligand binding information can also be used in protein engineering to graft binding pockets into other protein scaffolds or to transfer functional sites via recombination of a specific fragment. Fuzzle 2.0 is freely available at https://fuzzle.uni-bayreuth.de/2.0.
Natural evolution has generated an impressively diverse protein universe via duplication and recombination from a set of protein fragments that served as building blocks. The application of these concepts to the design of new proteins using subdomain-sized fragments from different folds has proven to be experimentally successful. To better understand how evolution has shaped our protein universe, we performed an all-against-all comparison of protein domains representing all naturally existing folds and identified conserved homologous protein fragments. Overall, we found more than 1000 protein fragments of various lengths among different folds through similarity network analysis. These fragments are present in very different protein environments and represent versatile building blocks for protein design. These data are available in our web server called F(old P)uzzle (fuzzle.uni-bayreuth.de), which allows to individually filter the dataset and create customized networks for folds of interest. We believe that our results serve as an invaluable resource for structural and evolutionary biologists and as raw material for the design of custom-made proteins.
Experiences as the basis for value creation and competitive positioning are increasingly placed at the centre of marketing activities to create an emotional customer-brand relationship. Especially in the travel and tourism market, the demand for brand experiences becomes apparent and is reflected in a wide range of services ranging from transport and accommodation to entertainment and relaxation. The cruise ship industry as the fastest growing segment in the travel industry provides a holistic experiential package designed to meet the travellers' expectations for pleasure and satisfaction. Against this backdrop, the aim of this paper is to empirically investigate antecedents of consumer perception and related consumption behaviour with special focus on brand experiences and sustainability orientation. The results reveal that consumers have an increasing demand for personal and authentic experiences combined with a rising concern regarding ethical and environmental values. As a consequence, addressing brand experience and sustainability orientation is a promising way to create successful differentiation strategies in the travel and tourism industry.
Due to increasing recognition of the benefits provided by mangrove ecosystems, protection policies have emerged under both wetland and forestry programs. However, little consistency remains among these programs and inadequate coordination exists among sectors of government. With approximately 123 countries containing mangroves, the need for global management of these ecosystems is crucial to sustain the industries (i.e., fisheries, timber, and tourism) and coastal communities that mangroves support and protect. To determine the most effective form of mangrove management, this review examines management guidelines, particularly those associated with Integrated Coastal Zone Management (ICZM). Five case studies were reviewed to further explore the fundamentals of mangrove management. The management methodologies of two developed nations as well as three developing nations were assessed to encompass comprehensive influences on mangrove management, such as socioeconomics, politics, and land-use regulations. Based on this review, successful mangrove management will require a blend of forestry, wetland, and ICZM programs in addition to the cooperation of all levels of government. Legally binding policies, particularly at the international level, will be essential to successful mangrove management, which must include the preservation of existing mangrove habitat and restoration of damaged mangroves.
Developmental programs have the fidelity to form neural circuits with the same structure and function among individuals of the same species. It is less well understood, however, to what extent entire neural circuits of different individuals are similar. Previously, we reported the neuronal connectome of the visual eye circuit from the head of a Platynereis dumerilii larva (Randel et al., 2014). We now report a full-body serial section transmission electron microscopy (ssTEM) dataset of another larva of the same age, for which we describe the connectome of the visual eyes and the larval eyespots. Anatomical comparisons and quantitative analyses of the two circuits reveal a high inter-individual stereotypy of the cell complement, neuronal projections, and synaptic connectivity, including the left-right asymmetry in the connectivity of some neurons. Our work shows the extent to which the eye circuitry in Platynereis larvae is hard-wired.
Finding evolutionary links between protein superfamilies has proven challenging. Advanced bioinformatics tools now identify relationships across two superfolds as well as a hybrid family whose structure displays characteristics of both.
The rehabilitation of patients should not only be limited to the first phases during intense hospital care but also support and therapy should be guaranteed in later stages, especially during daily life activities if the patient's state requires this. However, aid should only be given to the patient if needed and as much as it is required. To allow this, automatic self-initiated movement support and patient-cooperative control strategies have to be developed and integrated into assistive systems. In this work, we first give an overview of different kinds of neuromuscular diseases, review different forms of therapy, and explain possible fields of rehabilitation and benefits of robotic aided rehabilitation. Next, the mechanical design and control scheme of an upper limb orthosis for rehabilitation are presented. Two control models for the orthosis are explained which compute the triggering function and the level of assistance provided by the device. As input to the model fused sensor data from the orthosis and physiology data in terms of electromyography (EMG) signals are used.
In this work an upper limb active orthosis for assistive rehabilitation is presented. The design and torque control scheme of the orthosis that take into account important aspects of human rehabilitation, are described. Furthermore, first results of successful muscle activity detection and processing for the operation of the orthosis in two movement directions are presented. The proposed system is the first step towards an adaptive support of patients with respect to the strength of their muscle activity. To allow an adaptive support, different methods for EMG analysis have to be applied which allow to correlate muscle activity strength with the recorded signal and thus enable to adapt the support of the orthosis to the needs of the patient and state of
The CCR4-NOT complex plays a crucial role in post-transcriptional mRNA regulation in eukaryotes. This complex catalyzes the removal of mRNA poly(A) tails, thereby repressing translation and committing an mRNA to degradation. The conserved core of the complex is assembled by the interaction of at least two modules: the NOT module, which minimally consists of NOT1, NOT2 and NOT3, and a catalytic module comprising two deadenylases, CCR4 and POP2/CAF1. Additional complex subunits include CAF40 and two newly identified human subunits, NOT10 and C2orf29. The role of the NOT10 and C2orf29 subunits and how they are integrated into the complex are unknown. Here, we show that the Drosophila melanogaster NOT10 and C2orf29 orthologs form a complex that interacts with the N-terminal domain of NOT1 through C2orf29. These interactions are conserved in human cells, indicating that NOT10 and C2orf29 define a conserved module of the CCR4-NOT complex. We further investigated the assembly of the D. melanogaster CCR4-NOT complex, and demonstrate that the conserved armadillo repeat domain of CAF40 interacts with a region of NOT1, comprising a domain of unknown function, DUF3819. Using tethering assays, we show that each subunit of the CCR4-NOT complex causes translational repression of an unadenylated mRNA reporter and deadenylation and degradation of a polyadenylated reporter. Therefore, the recruitment of a single subunit of the complex to an mRNA target induces the assembly of the complete CCR4-NOT complex, resulting in a similar regulatory outcome.
Research and/or Engineering Questions/Objective The vehicle of the future will support its driver by advising him regarding potential hazards. Essential prerequisite therefore is the sensor based perception of the traffic situation. For the recognition of traffic related objects, camera based sensors, deepness cameras, vehicle sensors as well as radar and lidar sensors are used. For the future development of ADAS the fusion of multiple sensor data to a consistent environmental picture will play a key role. The evaluation approach of real world driving tests will no longer be sufficient due to the complexity of the system interactions. New simulation methods are needed to evaluate ADAS by using virtual test driving with realistic vehicle behavior and complex traffic environment. Methodology Therefore it is important to integrate camera based components in a “closed loop”-simulation platform to be able to test sensor data fusion technologies under realistic conditions. To test new driver assistance systems in a simulation environment today animation data is filmed, subsequently this data is used to test an image processing algorithm or a fusion algorithm. But this method cannot be applied if wide-angle cameras such as cameras with fisheye lenses will be used. Within a research frame work for autonomous driving functions a new simulation technology was developed to integrate virtual cameras beside the well know environment sensor in the vehicle dynamic simulation CarMaker. For this purpose the real-time animation was extended with a sophisticated virtual camera model so called “VideoDataStream” to generate simultaneous video data (also PMD for 3D images). The camera positions as well as the camera properties could be applied individually. Additionally it is possible to freely define the type of the camera lens (e.g. fisheye) with lens settings like opening angle and the typical lens failures (e.g. distortion and vignetting). With this new technology it is possible that e.g. camera and radar data can be provided time and place synchronal for the fusion algorithm which should be tested! Results The video data could be used for evaluating image processing and sensor data fusion in Model-/Software-/Hardware-in-the-Loop applications within virtual test driving conditions. Here the created method and examples of image based perception of the vehicle environment as well as sensor data fusion algorithms shall be presented. Among others this covers first of all the recognition of traffic lanes, traffic signs and other traffic partners as well as the fusion of the single information up to a comprehensive environment picture. A further field of application will be the conjunction with navigation systems and digital maps, by which the virtual vehicle supports the navigation system with related GPS position and gets back the “MPP—Most Probable Path” with the “electronic horizon”, which is a type of predictive sensor, with all related preview information in front of the vehicle which are defined in the ADASIS protocol. Conclusion By using the introduced method the capability and efficiency of function development and testing in the area of Advanced Driver Assistant Systems will significantly be improved. Due to a powerful simulation environment a broad range of validation tests can be shifted into simulation because also complex test scenarios can be replicated and the tests are reproducible. The simulation data can be provided time and place synchronal, which is absolutely important, e.g. for a fusion algorithm which should be tested.
This work introduces the architecture of a novel brain-arm haptic interface usable to improve the operation of complex robotic systems, or to deliver a fine rehabilitation therapy to the human upper limb. The proposed control scheme combines different approaches from the areas of robotics, neuroscience and human-machine interaction in order to overcome the limitations of each single field. Via the adaptive Brain Reading Interface (aBRI) user movements are anticipated by classification of surface electroencephalographic data in a millisecond range. This information is afterwards integrated into the control strategy of a wearable exoskeleton in order to finely modulate its impedance and therefore to comply with the motion preparation of the user. Results showing the efficacy of the proposed control approach are presented for the single joint case.
The LINE-1 (L1) retrotransposon emerges as a major source of human interindividual genetic variation, with important implications for evolution and disease. L1 retrotransposition is poorly understood at the molecular level, and the mechanistic details and evolutionary origin of the L1-encoded L1ORF1 protein (L1ORF1p) are particularly obscure. Here three crystal structures of trimeric L1ORF1p and NMR solution structures of individual domains reveal a sophisticated and highly structured, yet remarkably flexible, RNA-packaging protein. It trimerizes via an N-terminal, ion-containing coiled coil that serves as scaffold for the flexible attachment of the central RRM and the C-terminal CTD domains. The structures explain the specificity for single-stranded RNA substrates, and a mutational analysis indicates that the precise control of domain flexibility is critical for retrotransposition. Although the evolutionary origin of L1ORF1p remains unclear, our data reveal previously undetected structural and functional parallels to viral proteins.
In this paper an innovative low pressure servo-valve is presented. The device was designed with the main aim to be easily integrable into complex hydraulic/pneumatic actuation systems, and to operate at relatively low pressure (< 50.10(5)Pa). Characteristics like compactness, lightweight, high bandwidth, and autonomous sensory capability, where considered during the design process in order to achieve a device that fulfills the basic requirements for a wearable robotic system. Preliminary results about the prototype performances are presented here, in particular its dynamic behavior was measured for different working conditions, and a non-linear model identified using a recursive Hammerstein-Wiener parameter adaptation algorithm.
To the Editor: Applications of rapidly advancing sequencing technologies exacerbate the need to interpret individual sequence variants. Sequencing of phenotyped clinical subjects will soon become a method of choice in studies of the genetic causes of Mendelian and complex diseases. New exon capture techniques will direct sequencing efforts towards the most informative and easily interpretable protein-coding fraction of the genome. Thus, the demand for computational predictions of the impact of protein sequence variants will continue to grow. Here we present a new method and the corresponding software tool, PolyPhen-2 (http://genetics.bwh.harvard.edu/pph2/), which is different from the early tool PolyPhen1 in the set of predictive features, alignment pipeline, and the method of classification (Fig. 1a). PolyPhen-2 uses eight sequence-based and three structure-based predictive features (Supplementary Table 1) which were selected automatically by an iterative greedy algorithm (Supplementary Methods). Majority of these features involve comparison of a property of the wild-type (ancestral, normal) allele and the corresponding property of the mutant (derived, disease-causing) allele, which together define an amino acid replacement. Most informative features characterize how well the two human alleles fit into the pattern of amino acid replacements within the multiple sequence alignment of homologous proteins, how distant the protein harboring the first deviation from the human wild-type allele is from the human protein, and whether the mutant allele originated at a hypermutable site2. The alignment pipeline selects the set of homologous sequences for the analysis using a clustering algorithm and then constructs and refines their multiple alignment (Supplementary Fig. 1). The functional significance of an allele replacement is predicted from its individual features (Supplementary Figs. 2–4) by Naive Bayes classifier (Supplementary Methods). Figure 1 PolyPhen-2 pipeline and prediction accuracy. (a) Overview of the algorithm. (b) Receiver operating characteristic (ROC) curves for predictions made by PolyPhen-2 using five-fold cross-validation on HumDiv (red) and HumVar3 (light green). UniRef100 (solid ... We used two pairs of datasets to train and test PolyPhen-2. We compiled the first pair, HumDiv, from all 3,155 damaging alleles with known effects on the molecular function causing human Mendelian diseases, present in the UniProt database, together with 6,321 differences between human proteins and their closely related mammalian homologs, assumed to be non-damaging (Supplementary Methods). The second pair, HumVar3, consists of all the 13,032 human disease-causing mutations from UniProt, together with 8,946 human nsSNPs without annotated involvement in disease, which were treated as non-damaging. We found that PolyPhen-2 performance, as presented by its receiver operating characteristic curves, was consistently superior compared to PolyPhen (Fig. 1b) and it also compared favorably with the three other popular prediction tools4–6 (Fig. 1c). For a false positive rate of 20%, PolyPhen-2 achieves the rate of true positive predictions of 92% and 73% on HumDiv and HumVar, respectively (Supplementary Table 2). One reason for a lower accuracy of predictions on HumVar is that nsSNPs assumed to be non-damaging in HumVar contain a sizable fraction of mildly deleterious alleles. In contrast, most of amino acid replacements assumed non-damaging in HumDiv must be close to selective neutrality. Because alleles that are even mildly but unconditionally deleterious cannot be fixed in the evolving lineage, no method based on comparative sequence analysis is ideal for discriminating between drastically and mildly deleterious mutations, which are assigned to the opposite categories in HumVar. Another reason is that HumDiv uses an extra criterion to avoid possible erroneous annotations of damaging mutations. For a mutation, PolyPhen-2 calculates Naive Bayes posterior probability that this mutation is damaging and reports estimates of false positive (the chance that the mutation is classified as damaging when it is in fact non-damaging) and true positive (the chance that the mutation is classified as damaging when it is indeed damaging) rates. A mutation is also appraised qualitatively, as benign, possibly damaging, or probably damaging (Supplementary Methods). The user can choose between HumDiv- and HumVar-trained PolyPhen-2. Diagnostics of Mendelian diseases requires distinguishing mutations with drastic effects from all the remaining human variation, including abundant mildly deleterious alleles. Thus, HumVar-trained PolyPhen-2 should be used for this task. In contrast, HumDiv-trained PolyPhen-2 should be used for evaluating rare alleles at loci potentially involved in complex phenotypes, dense mapping of regions identified by genome-wide association studies, and analysis of natural selection from sequence data, where even mildly deleterious alleles must be treated as damaging.
Roded Sharan合作论文数Tel-Aviv University;School of Computer Science2