Protein kinase C θ (PKCθ) is expressed in T lymphocytes, in which it plays a crucial role in cell activation, proliferation, and differentiation. Previously, we discovered that the Thr335-Pro motif within the PKCθ-V3 regulatory domain serves as a critical site for PKCθ activation and demonstrated that phosphorylation of the Thr335-Pro motif creates a binding site for the peptidyl-prolyl cis-trans isomerase (PPIase), Pin1. Herein, we elaborate on the functional consequences of the Pin1-PKCθ association and identify Pin1 as a regulator of PKCθ catalytic activity. Pin1 downregulated the activity of PKCθ in vitro in PMA-stimulated human Jurkat T cells and C57BL/6J mouse spleen- and thymus-derived T lymphocytes, an effect that was reversed by juglone, a Pin1-activity inhibitor. Utilizing Pin1 isomerase-deficient mutants, we demonstrate that the functionality of the Pin1 catalytic domain is essential for its regulation of PKCθ, suggesting that Pin1 mediates its inhibitory effect on PKCθ via cis-trans isomerization. In silico docking analysis supported the role of critical residues within the Pin1-PPIase domain that enables the cis-trans interconversion of the PKCθ phospho-Thr335-Pro motif. Retroviral knockdown of Pin1 in Jurkat T cells led to an elevation in PKCθ kinase activity. Furthermore, stimulation of [Lckcre × Pin1lox] F1 mice-derived Pin1-deficient T cells augmented the phosphorylation of SPAK kinase, a bona fide PKCθ downstream substrate. Together, our results support a role for the Pin1 isomerase as a regulator of PKCθ activity and highlight the potential contribution of Pin1 to the fine-tuning of the T cell activation dynamics.
Commercial bumblebee colonies are routinely used for crop pollination in greenhouses and are increasingly introduced into orchards as well. Bumblebee spillover to natural habitats near the orchards may interfere with local wild bees and impact the pollination of non-crop plants. Concurrently, foraging in natural habitats may diversify the bumblebees’ diets and improve colony development. To evaluate these potential effects, we placed commercial Bombus terrestris colonies in blooming Rosaceae orchards, 25–125 m away from the margins. We recorded the colonies’ mass gain, population sizes, composition of stored pollen, and temperature regulation. We monitored bee activity, and seed sets of the non-crop plant Eruca sativa, along transects in a semi-natural shrubland up to 100 m away from the orchards, with managed bumblebees either present or absent. Rosaceae pollen comprised 1/3 of the colonies’ pollen stores at all distances from the orchard margins. Colonies placed closest to the margins showed prolonged development, produced fewer reproductive individuals, and had poorer thermoregulation than colonies closer to the orchards’ center. Possibly, abiotic stressors inhibited the bumblebees’ development near orchard borders. Wild bees were as active during the colonies’ deployment as after their removal. E. sativa’s seed sets decreased after bumblebee removal, but similar declines also occurred near a control orchard without managed bumblebees. Altogether, we found no short-term spillover effects of managed bumblebees on nearby plant-bee communities during the orchards’ two-week flowering. The colonies’ prompt removal after blooming can reduce longer-term ecological risks associated with managed bumblebees.
Abstract About 40% of human mRNAs contain upstream open reading frames (uORFs) in their 5’-untranslated regions. Some of these uORF sequences thought to attenuate scanning ribosomes or lead to mRNA degradation were recently shown to be translated, although the function of most encoded peptides remains unknown. Our recent study was the first to reveal a uORF-encoded peptide exhibiting kinase inhibitory activity (Jayaram DR, et al. PNAS (2021)). This uORF (uORF2), upstream of a protein kinase C (PKC) family member (PKCeta) main ORF, encodes for a peptide (uPEP2) containing the typical PKC pseudosubstrate (PS) motif present in all PKCs that auto-inhibits their kinase activity. We show direct binding and selective inhibition of the catalytic activity of novel PKCs by uPEP2 (but not of classical or atypical PKCs). Functionally, treatment of breast cancer cells with uPEP2 diminished cell survival and their migration and synergized with chemotherapy by interfering with the response to DNA damage. Furthermore, in a xenograft of MDA-MB-231 triple-negative breast cancer in mice models, uPEP2 suppressed tumor progression, invasion, and metastasis. The endogenous deletion of uORF2 by CRISPR/Cas9 in the non-tumorigenic MCF-7 cells showed that it impedes cell proliferation, acting as a growth suppressor. Its role in tumor formation and metastasis is presently studied in breast cancer models with metastasis to the lungs. Together, we point to a new layer of protein regulation by a uORF-encoded kinase inhibitor, which could present a forerunner for other uORFs, possessing kinase inhibitory function. Notably, we have recently identified an additional uORF with a PS motif upstream of another PKC. As kinase inhibitors, uORF-encoded peptides may play a role in signaling and stress, acting in cis and/or in trans, forming diverse regulatory networks, thereby opening new views into eukaryotic protein regulation and cancer biology. Citation Format: Divya Ram Jayaram, Sigal Frost, Chanan Argov, Vijayasteltar Belsamma Liju, Nikhil Ponnoor Anto, Amitha Muraleedharan, Assaf Ben-Ari, Rose Sinay, Ilan Smoly, Ofra Novoplansky, Noah Isakov, Debra Toiber, Chen Keasar, Moshe Elkabets, Esti Yeger-Lotem, Etta Livneh. Beyond the Canonical: uORF-encoded Micro Peptides as Novel Kinase Inhibitors in Cancer Therapy [abstract]. In: Proceedings of Frontiers in Cancer Science; 2023 Nov 6-8; Singapore. Philadelphia (PA): AACR; Cancer Res 2024;84(8_Suppl):Abstract nr P42.
Insects are highly abundant and diverse, and play major roles in ecosystem functions. Monitoring of insect populations is key to their sustainable management. However, the labor and expertise needed to identify insects, and the challenges of archiving the wealth of data collected in monitoring programs, often limit these efforts. We describe a pipeline to reduce the barriers associated with curating and mining big data of insect biodiversity. The pipeline, STARdbi, includes capturing flying insects with sticky traps, scanning the traps, storing the trap-images in a public database with a web-based interface, and applying machine learning models to extract information from the images. To illustrate the insights that can be gained from STARdbi, we describe two case studies. One of them involves monitoring of circadian activity patterns of grain pests and of their natural enemies, and the other compares insect abundance, biomass and size distributions between agricultural and semi-natural habitats. We invite the community of insect ecologists to contribute to the STARdbi database, and to use its image analysis tools to address diverse ecological and evolutionary questions.
Protein kinase C-θ (PKCθ) is a member of the novel PKC subfamily known for its selective and predominant expression in T lymphocytes where it regulates essential functions required for T cell activation and proliferation. Our previous studies provided a mechanistic explanation for the recruitment of PKCθ to the center of the immunological synapse (IS) by demonstrating that a proline-rich (PR) motif within the V3 region in the regulatory domain of PKCθ is necessary and sufficient for PKCθ IS localization and function. Herein, we highlight the importance of Thr335-Pro residue in the PR motif, the phosphorylation of which is key in the activation of PKCθ and its subsequent IS localization. We demonstrate that the phospho-Thr335-Pro motif serves as a putative binding site for the peptidyl-prolyl cis-trans isomerase (PPIase), Pin1, an enzyme that specifically recognizes peptide bonds at phospho-Ser/Thr-Pro motifs. Binding assays revealed that mutagenesis of PKCθ-Thr335-to-Ala abolished the ability of PKCθ to interact with Pin1, while Thr335 replacement by a Glu phosphomimetic, restored PKCθ binding to Pin1, suggesting that Pin1-PKCθ association is contingent upon the phosphorylation of the PKCθ-Thr335-Pro motif. Similarly, the Pin1 mutant, R17A, failed to associate with PKCθ, suggesting that the integrity of the Pin1 N-terminal WW domain is a requisite for Pin1-PKCθ interaction. In silico docking studies underpinned the role of critical residues in the Pin1-WW domain and the PKCθ phospho-Thr335-Pro motif, to form a stable interaction between Pin1 and PKCθ. Furthermore, TCR crosslinking in human Jurkat T cells and C57BL/6J mouse-derived splenic T cells promoted a rapid and transient formation of Pin1-PKCθ complexes, which followed a T cell activation-dependent temporal kinetic, suggesting a role for Pin1 in PKCθ-dependent early activation events in TCR-triggered T cells. PPIases that belong to other subfamilies, i.e., cyclophilin A or FK506-binding protein, failed to associate with PKCθ, indicating the specificity of the Pin1-PKCθ association. Fluorescent cell staining and imaging analyses demonstrated that TCR/CD3 triggering promotes the colocalization of PKCθ and Pin1 at the cell membrane. Furthermore, interaction of influenza hemagglutinin peptide (HA307-319)-specific T cells with antigen-fed antigen presenting cells (APCs) led to colocalization of PKCθ and Pin1 at the center of the IS. Together, we point to an uncovered function for the Thr335-Pro motif within the PKCθ-V3 regulatory domain to serve as a priming site for its activation upon phosphorylation and highlight its tenability to serve as a regulatory site for the Pin1 cis-trans isomerase.
The inverse protein folding problem, also known as protein sequence design, seeks to predict an amino acid sequence that folds into a specific structure and performs a specific function. Recent advancements in machine learning techniques have been successful in generating functional sequences, outperforming previous energy function-based methods. However, these machine learning methods are limited in their interoperability and robustness, especially when designing proteins that must function under non-ambient conditions, such as high temperature, extreme pH, or in various ionic solvents. To address this issue, we propose a new Physics-Informed Neural Networks (PINNs)-based protein sequence design approach. Our approach combines all-atom molecular dynamics simulations, a PINNs MD surrogate model, and a relaxation of binary programming to solve the protein design task while optimizing both energy and the structural stability of proteins. We demonstrate the effectiveness of our design framework in designing proteins that can function under non-ambient conditions.
Computationally generated models of protein structures bridge the gap between the practically negligible price tag of sequencing and the high cost of experimental structure determination. By providing a low-cost (and often free) partial alternative to experimentally determined structures, these models help biologists design and interpret their experiments. Obviously, the more accurate the models the more useful they are. However, methods for protein structure prediction generate many structural models of various qualities, necessitating means for the estimation of their accuracy. In this work we present MESHI_consensus, a new method for the estimation of model accuracy. The method uses a tree-based regressor and a set of structural, target-based, and consensus-based features. The new method achieved high performance in the EMA (Estimation of Model Accuracy) track of the recent CASP14 community-wide experiment ( https://predictioncenter.org/casp14/index.cgi ). The tertiary structure prediction track of that experiment revealed an unprecedented leap in prediction performance by a single prediction group/method, namely AlphaFold2. This achievement would inevitably have a profound impact on the field of protein structure prediction, including the accuracy estimation sub-task. We conclude this manuscript with some speculations regarding the future role of accuracy estimation in a new era of accurate protein structure prediction.
Recent advancements in machine learning techniques for protein structure prediction motivate better results in its inverse problem–protein design. In this work we introduce a new graph mimetic neural network, MimNet, and show that it is possible to build a reversible architecture that solves the structure and design problems in tandem, allowing to improve protein backbone design when the structure is better estimated. We use the ProteinNet data set and show that the state of the art results in protein design can be met and even improved, given recent architectures for protein folding.
Approximately 40% of human messenger RNAs (mRNAs) contain upstream open reading frames (uORFs) in their 5' untranslated regions. Some of these uORF sequences, thought to attenuate scanning ribosomes or lead to mRNA degradation, were recently shown to be translated, although the function of the encoded peptides remains unknown. Here, we show a uORF-encoded peptide that exhibits kinase inhibitory functions. This uORF, upstream of the protein kinase C-eta (PKC-η) main ORF, encodes a peptide (uPEP2) containing the typical PKC pseudosubstrate motif present in all PKCs that autoinhibits their kinase activity. We show that uPEP2 directly binds to and selectively inhibits the catalytic activity of novel PKCs but not of classical or atypical PKCs. The endogenous deletion of uORF2 or its overexpression in MCF-7 cells revealed that the endogenously translated uPEP2 reduces the protein levels of PKC-η and other novel PKCs and restricts cell proliferation. Functionally, treatment of breast cancer cells with uPEP2 diminished cell survival and their migration and synergized with chemotherapy by interfering with the response to DNA damage. Furthermore, in a xenograft of MDA-MB-231 breast cancer tumor in mice models, uPEP2 suppressed tumor progression, invasion, and metastasis. Tumor histology showed reduced proliferation, enhanced cell death, and lower protein expression levels of novel PKCs along with diminished phosphorylation of PKC substrates. Hence, our study demonstrates that uORFs may encode biologically active peptides beyond their role as translation regulators of their downstream ORFs. Together, we point to a unique function of a uORF-encoded peptide as a kinase inhibitor, pertinent to cancer therapy.
Deep mutational scanning provides unprecedented wealth of quantitative data regarding the functional outcome of mutations in proteins. A single experiment may measure properties (eg, structural stability) of numerous protein variants. Leveraging the experimental data to gain insights about unexplored regions of the mutational landscape is a major computational challenge. Such insights may facilitate further experimental work and accelerate the development of novel protein variants with beneficial therapeutic or industrially relevant properties. Here we present a novel, machine learning approach for the prediction of functional mutation outcome in the context of deep mutational screens. Using sequence (one-hot) features of variants with known properties, as well as structural features derived from models thereof, we train predictive statistical models to estimate the unknown properties of other variants. The utility of the new computational scheme is demonstrated using five sets of mutational scanning data, denoted "targets": (a) protease specificity of APPI (amyloid precursor protein inhibitor) variants; (b-d) three stability related properties of IGBPG (immunoglobulin G-binding β1 domain of streptococcal protein G) variants; and (e) fluorescence of GFP (green fluorescent protein) variants. Performance is measured by the overall correlation of the predicted and observed properties, and enrichment-the ability to predict the most potent variants and presumably guide further experiments. Despite the diversity of the targets the statistical models can generalize variant examples thereof and predict the properties of test variants with both single and multiple mutations.
Ecology documents and interprets the abundance and distribution of organisms. Ecoinformatics addresses this challenge by analyzing databases of observational data. Ecoinformatics of insects has high scientific and applied importance, as insects are abundant, speciose, and involved in many ecosystem functions. They also crucially impact human well-being, and human activities dramatically affect insect demography and phenology. Hazards, such as pollinator declines, outbreaks of agricultural pests and the spread insect-borne diseases, raise an urgent need to develop ecoinformatics strategies for their study. Yet, insect databases are mostly focused on a small number of pest species, as data acquisition is labor-intensive and requires taxonomical expertise. Thus, despite decades of research, we have only a qualitative notion regarding fundamental questions of insect ecology, and only limited knowledge about the spatio-temporal distribution of insects. We describe a novel high throughput cost-effective approach for monitoring flying insects as an enabling step toward “big data” entomology. The proposed approach combines “high tech” deep learning with “low tech” sticky traps that sample flying insects in diverse locations. As a proof of concept we considered three recent insect invaders of Israel’s forest ecosystem: two hemipteran pests of eucalypts and a parasitoid wasp that attacks one of them. We developed software, based on deep learning, to identify the three species in images of sticky traps from Eucalyptus forests. These image processing tasks are quite difficult as the insects are small (<5 mm) and stick to the traps in random poses. The resulting deep learning model discriminated the three focal organisms from one another, as well as from other elements such as leaves and other insects, with high precision. We used the model to compare the abundances of these species among six sites, and validated the results by manually counting insects on the traps. Having demonstrated the power of the proposed approach, we started a more ambitious study that monitors these insects at larger spatial and temporal scales. We aim at building an ecoinformatics repository for trap images and generating data-driven models of the populations’ dynamics and morphological traits.
MOTIVATION:The Protein Data Bank (PDB), the ultimate source for data in structural biology, is inherently imbalanced. To alleviate biases, virtually all structural biology studies use nonredundant (NR) subsets of the PDB, which include only a fraction of the available data. An alternative approach, dubbed redundancy-weighting (RW), down-weights redundant entries rather than discarding them. This approach may be particularly helpful for machine-learning (ML) methods that use the PDB as their source for data. Methods for secondary structure prediction (SSP) have greatly improved over the years with recent studies achieving above 70% accuracy for eight-class (DSSP) prediction. As these methods typically incorporate ML techniques, training on RW datasets might improve accuracy, as well as pave the way toward larger and more informative secondary structure classes.RESULTS:This study compares the SSP performances of deep-learning models trained on either RW or NR datasets. We show that training on RW sets consistently results in better prediction of 3- (HCE), 8- (DSSP) and 13-class (STR2) secondary structures.AVAILABILITY AND IMPLEMENTATION:The ML models, the datasets used for their derivation and testing, and a stand-alone SSP program for DSSP and STR2 predictions, are freely available under LGPL license in http://meshi1.cs.bgu.ac.il/rw.SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
The Protein Data Bank (PDB), the ultimate source for data in structural biology, is inherently imbalanced. To alleviate biases, virtually all structural biology studies use non-redundant subsets of the PDB, which include only a fraction of the available data. An alternative approach, dubbed redundancy-weighting, down-weights redundant entries rather than discarding them. This approach may be particularly helpful for Machine Learning (ML) methods that use the PDB as their source for data. Current state-of-art methods for Secondary Structure Prediction of proteins (SSP) use non-redundant datasets to predict either 3-letter or 8-letter secondary structure annotations. The current study challenges both choices: the dataset and alphabet size. Non-redundant datasets are presumably unbiased, but are also inherently small, which limits the machine learning performance. On the other hand, the utility of both 3- and 8-letter alphabets is limited by the aggregation of parallel, anti-parallel, and mixed beta-sheets in a single class. Each of these subclasses imposes different structural constraints, which makes the distinction between them desirable. In this study we show improvement in prediction accuracy by training on a redundancy-weighted dataset. Further, we show the information content is improved by extending the alphabet to consider beta subclasses while hardly effecting SSP accuracy.
Motivation Methods for protein structure prediction (PSP) generate multiple alternative structural models (aka decoys). Thus, supervised learning methods for the evaluation and ranking of these models are crucial elements of PSP. Supervised learning involves optimization of loss functions, but their influence on performance is typically overlooked. Here we put the loss functions in the spotlight, and study their effect on prediction performance. Results Here we report the performances of three variants of MESHI-score, a supervised learning method for the estimation of model accuracy (EMA). Each variant was trained with a different loss function and showed better performance in different aspects of the EMA problem. Most importantly, better discrimination between models of the same target, is gained by target centered loss functions. Availability All data is available at http://meshi1.cs.bgu.ac.il/SidiAndKeasar2018Data_download/ . The MESHI-package (version 9.412) is available at https://github.com/meshiprot/meshi/releases ). Contact chen.keasar@gmail.com
Proteins are the major building blocks of life, and actuators of almost all chemical and biophysical events in living organisms. Their native structures in turn enable their biological functions which have a fundamental role in drug design. This motivates predicting the structure of a protein from its sequence of amino acids, a fundamental problem in computational biology. In this work, we demonstrate state-of-the-art protein structure prediction (PSP) results using embeddings and deep learning models for prediction of backbone atom distance matrices and torsion angles. We recover 3D coordinates of backbone atoms and reconstruct full atom protein by optimization. We create a new gold standard dataset of proteins which is comprehensive and easy to use. Our dataset consists of amino acid sequences, Q8 secondary structures, position specific scoring matrices, multiple sequence alignment co-evolutionary features, backbone atom distance matrices, torsion angles, and 3D coordinates. We evaluate the quality of our structure prediction by RMSD on the latest Critical Assessment of Techniques for Protein Structure Prediction (CASP) test data and demonstrate competitive results with the winning teams and AlphaFold in CASP13 and supersede the results of the winning teams in CASP12. We make our data, models, and code publicly available.
The function of a protein is determined by its structure, which creates a need for efficient methods of protein structure determination to advance scientific and medical research. Because current experimental structure determination methods carry a high price tag, computational predictions are highly desirable. Given a protein sequence, computational methods produce numerous 3D structures known as decoys. Selection of the best quality decoys is both challenging and essential as the end users can handle only a few ones. Therefore, scoring functions are central to decoy selection. They combine measurable features into a single number indicator of decoy quality. Unfortunately, current scoring functions do not consistently select the best decoys. Machine learning techniques offer great potential to improve decoy scoring. This paper presents two machine-learning based scoring functions to predict the quality of proteins structures, i.e., the similarity between the predicted structure and the experimental one without knowing the latter. We use different metrics to compare these scoring functions against three state-of-the-art scores. This is a first attempt at comparing different scoring functions using the same non-redundant dataset for training and testing and the same features. The results show that adding informative features may be more significant than the method used.
We tackle the problem of protein secondary structure prediction using a common task framework. This lead to the introduction of multiple ideas for neural architectures based on state of the art building blocks, used in this task for the first time. We take a principled machine learning approach, which provides genuine, unbiased performance measures, correcting longstanding errors in the application domain. We focus on the Q8 resolution of secondary structure, an active area for continuously improving methods. We use an ensemble of strong predictors to achieve accuracy of 70.7% (on the CB513 test set using the CB6133filtered training set). These results are statistically indistinguishable from those of the top existing predictors. In the spirit of reproducible research we make our data, models and code available, aiming to set a gold standard for purity of training and testing sets. Such good practices lower entry barriers to this domain and facilitate reproducible, extendable research.
Every two years groups worldwide participate in the Critical Assessment of Protein Structure Prediction (CASP) experiment to blindly test the strengths and weaknesses of their computational methods. CASP has significantly advanced the field but many hurdles still remain, which may require new ideas and collaborations. In 2012 a web-based effort called WeFold, was initiated to promote collaboration within the CASP community and attract researchers from other fields to contribute new ideas to CASP. Members of the WeFold coopetition (cooperation and competition) participated in CASP as individual teams, but also shared components of their methods to create hybrid pipelines and actively contributed to this effort. We assert that the scale and diversity of integrative prediction pipelines could not have been achieved by any individual lab or even by any collaboration among a few partners. The models contributed by the participating groups and generated by the pipelines are publicly available at the WeFold website providing a wealth of data that remains to be tapped. Here, we analyze the results of the 2014 and 2016 pipelines showing improvements according to the CASP assessment as well as areas that require further adjustments and research.
Methods to reliably estimate the quality of 3D models of proteins are essential drivers for the wide adoption and serious acceptance of protein structure predictions by life scientists. In this article, the most successful groups in CASP12 describe their latest methods for estimates of model accuracy (EMA). We show that pure single model accuracy estimation methods have shown clear progress since CASP11; the 3 top methods (MESHI, ProQ3, SVMQA) all perform better than the top method of CASP11 (ProQ2). Although the pure single model accuracy estimation methods outperform quasi-single (ModFOLD6 variations) and consensus methods (Pcons, ModFOLDclust2, Pcomb-domain, and Wallner) in model selection, they are still not as good as those methods in absolute model quality estimation and predictions of local quality. Finally, we show that when using contact-based model quality measures (CAD, lDDT) the single model quality methods perform relatively better.