P110α is a member of the phosphoinositide 3-kinase (PI3K) enzyme family that functions downstream of RAS. RAS proteins contribute to the activation of p110α by interacting directly with its RAS binding domain (RBD), resulting in the promotion of many cellular functions such as cell growth, proliferation and survival. Previous work from our lab has highlighted the importance of the p110α/RAS interaction in tumour initiation and growth. Here we report the discovery and characterisation of a cyclic peptide inhibitor (cyclo-CRVLIR) that interacts with the p110α-RBD and blocks its interaction with KRAS. cyclo-CRVLIR was discovered by screening a “split-intein cyclisation of peptides and proteins” (SICLOPPS) cyclic peptide library. The primary cyclic peptide hit from the screen initially showed a weak affinity for the p110α-RBD (Kd about 360 µM). However, two rounds of amino acid substitution led to cyclo-CRVLIR, with an improved affinity for p110α-RBD in the low µM (Kd 3 µM). We show that cyclo-CRVLIR binds selectively to the p110α-RBD but not to KRAS or the structurally-related RAF-RBD. Further, using biophysical, biochemical and cellular assays, we show that cyclo-CRVLIR effectively blocks the p110α/KRAS interaction in a dose dependent manner and reduces phospho-AKT levels in several oncogenic KRAS cell lines.
Cancers, such as squamous cell carcinoma, frequently invade as multicellular units. However, these invading units can be organized in a variety of ways, ranging from thin discontinuous strands to thick ‘pushing’ collectives. Here we employ an integrated experimental and computational approach to identify the factors that determine the mode of collective cancer cell invasion. We find that matrix proteolysis is linked to the formation of wide strands, but has little effect on the maximum extent of invasion. Cell-cell junctions also favour wide strands, but our analysis also reveals a requirement for cell-cell junctions for efficient invasion in response to uniform directional cues. Unexpectedly, the ability to generate wide invasive strands is coupled to the ability to grow effectively when surrounded by ECM in 3D assays. Combinatorial perturbation of both matrix proteolysis and cell-cell adhesion demonstrates that the most aggressive cancer behaviour, both in terms of invasion and growth, is achieved at high levels of cell-cell adhesion and high levels of proteolysis. Contrary to expectation, cells with canonical mesenchymal traits – no cell-cell junctions and high proteolysis – exhibit reduced growth and lymph node metastasis. Thus, we conclude that the ability of squamous cell carcinoma cells to invade effectively is also linked to their ability to generate space for proliferation in confined contexts. These data provide an explanation for the apparent advantage of retaining cell-cell junctions in SCC.
During meiosis, programmed DNA double-strand breaks (DSBs) are repaired by homologous recombination. DMC1, a conserved recombinase, plays a central role in this process. DMC1 promotes DNA strand exchange between homologous chromosomes, thus creating the physical linkage between them. Its function is regulated not only by several accessory proteins but also by bivalent ions. Here, we show that whereas calcium ions in the presence of ATP cause a conformational change within DMC1, stimulating its DNA binding and D-loop formation, they inhibit the extension of the invading strand within the D-loop. Based on structural studies, we have generated mutants of two highly conserved amino acids - E162 and D317 - in human DMC1, which are deficient in calcium regulation. In vivo studies of their yeast homologues further showed that they exhibit severe defects in meiosis, thus emphasizing the importance of calcium ions in the regulation of DMC1 function and meiotic recombination.
We present the results for CAPRI Round 50, the fourth joint CASP‐CAPRI protein assembly prediction challenge. The Round comprised a total of twelve targets, including six dimers, three trimers, and three higher‐order oligomers. Four of these were easy targets, for which good structural templates were available either for the full assembly, or for the main interfaces (of the higher‐order oligomers). Eight were difficult targets for which only distantly related templates were found for the individual subunits. Twenty‐five CAPRI groups including eight automatic servers submitted ~1250 models per target. Twenty groups including six servers participated in the CAPRI scoring challenge submitted ~190 models per target. The accuracy of the predicted models was evaluated using the classical CAPRI criteria. The prediction performance was measured by a weighted scoring scheme that takes into account the number of models of acceptable quality or higher submitted by each group as part of their five top‐ranking models. Compared to the previous CASP‐CAPRI challenge, top performing groups submitted such models for a larger fraction (70–75%) of the targets in this Round, but fewer of these models were of high accuracy. Scorer groups achieved stronger performance with more groups submitting correct models for 70–80% of the targets or achieving high accuracy predictions. Servers performed less well in general, except for the MDOCKPP and LZERD servers, who performed on par with human groups. In addition to these results, major advances in methodology are discussed, providing an informative overview of where the prediction of protein assemblies currently stands.
Vγ9Vδ2 T cells respond in a TCR-dependent fashion to both microbial and host-derived pyrophosphate compounds (phosphoantigens, or P-Ag). Butyrophilin-3A1 (BTN3A1), a protein structurally related to the B7 family of costimulatory molecules, is necessary but insufficient for this process. We performed radiation hybrid screens to uncover direct TCR ligands and cofactors that potentiate BTN3A1’s P-Ag sensing function. These experiments identified butyrophilin-2A1 (BTN2A1) as essential to Vγ9Vδ2 T cell recognition. BTN2A1 synergised with BTN3A1 in sensitizing P-Ag-exposed cells for Vγ9Vδ2 TCR-mediated responses. Surface plasmon resonance experiments established Vγ9Vδ2 TCRs used germline-encoded Vγ9 regions to directly bind the BTN2A1 CFG-IgV domain surface. Notably, somatically recombined CDR3 loops implicated in P-Ag recognition were uninvolved. Immunoprecipitations demonstrated close cell-surface BTN2A1-BTN3A1 association independent of P-Ag stimulation. Thus, BTN2A1 is a BTN3A1-linked co-factor critical to Vγ9Vδ2 TCR recognition. Furthermore, these results suggest a composite-ligand model of P-Ag sensing wherein the Vγ9Vδ2 TCR directly interacts with both BTN2A1 and an additional ligand recognized in a CDR3-dependent manner.
Many of the biological functions of the cell are driven by protein-protein interactions. However, determining which proteins interact and exactly how they do so to enable their functions, remain major research questions. Functional interactions are dependent on a number of complicated factors; therefore, modeling the three-dimensional structure of protein-protein complexes is still considered a complex endeavor. Nevertheless, the rewards for modeling protein interactions to atomic level detail are substantial, and there are numerous examples of how models can provide useful information for drug design, protein engineering, systems biology, and understanding of the immune system. Here, we provide practical guidelines for docking proteins using the web-server, SwarmDock, a flexible protein-protein docking method. Moreover, we provide an overview of the factors that need to be considered when deciding whether docking is likely to be successful.
The formation of specific protein-protein interactions is often a key to a protein's function. During complex formation, each protein component will undergo a change in the conformational state, for some these changes are relatively small and reside primarily at the sidechain level; however, others may display notable backbone adjustments. One of the classic problems in the protein-docking field is to be able to a priori predict the extent of such conformational changes. In this work, we investigated three protocols to find the most suitable input structure conformations for cross-docking, including a robust sampling approach in normal mode space. Counterintuitively, knowledge of the theoretically best combination of normal modes for unbound-bound transitions does not always lead to the best results. We used a novel spatial partitioning library, Aether Engine (see Supplementary Materials), to efficiently search the conformational states of 56 receptor/ligand pairs, including a recent CAPRI target, in a systematic manner and selected diverse conformations as input to our automated docking server, SwarmDock, a server that allows moderate conformational adjustments during the docking process. In essence, here we present a dynamic cross-docking protocol, which when benchmarked against the simpler approach of just docking the unbound components shows a 10% uplift in the quality of the top docking pose.
We present the results for CAPRI Round 46, the third joint CASP-CAPRI protein assembly prediction challenge. The Round comprised a total of 20 targets including 14 homo-oligomers and 6 heterocomplexes. Eight of the homo-oligomer targets and one heterodimer comprised proteins that could be readily modeled using templates from the Protein Data Bank, often available for the full assembly. The remaining 11 targets comprised 5 homodimers, 3 heterodimers, and two higher-order assemblies. These were more difficult to model, as their prediction mainly involved "ab-initio" docking of subunit models derived from distantly related templates. A total of ~30 CAPRI groups, including 9 automatic servers, submitted on average ~2000 models per target. About 17 groups participated in the CAPRI scoring rounds, offered for most targets, submitting ~170 models per target. The prediction performance, measured by the fraction of models of acceptable quality or higher submitted across all predictors groups, was very good to excellent for the nine easy targets. Poorer performance was achieved by predictors for the 11 difficult targets, with medium and high quality models submitted for only 3 of these targets. A similar performance "gap" was displayed by scorer groups, highlighting yet again the unmet challenge of modeling the conformational changes of the protein components that occur upon binding or that must be accounted for in template-based modeling. Our analysis also indicates that residues in binding interfaces were less well predicted in this set of targets than in previous Rounds, providing useful insights for directions of future improvements.
The atomic structures of protein complexes can provide useful information for drug design, protein engineering, systems biology, and understanding pathology. Obtaining this information experimentally can be challenging. However, if the structures of the subunits are known, then it is often possible to model the complex computationally. This chapter provide practical guidelines for docking proteins using the SwarmDock flexible protein-protein docking method, providing an overview of the factors that need to be considered when deciding whether docking is likely to be successful, the preparation of structural input, generation of docked poses, analysis and ranking of docked poses, and the validation of models using external data.
T lymphocytes expressing γδ T cell antigen receptors (TCRs) comprise evolutionarily conserved cells with paradoxical features. On the one hand, clonally expanded γδ T cells with unique specificities typify adaptive immunity. Conversely, large compartments of γδTCR+ intraepithelial lymphocytes (γδ IELs) exhibit limited TCR diversity and effect rapid, innate-like tissue surveillance. The development of several γδ IEL compartments depends on epithelial expression of genes encoding butyrophilin-like (Btnl (mouse) or BTNL (human)) members of the B7 superfamily of T cell co-stimulators. Here we found that responsiveness to Btnl or BTNL proteins was mediated by germline-encoded motifs within the cognate TCR variable γ-chains (Vγ chains) of mouse and human γδ IELs. This was in contrast to diverse antigen recognition by clonally restricted complementarity-determining regions CDR1-CDR3 of the same γδTCRs. Hence, the γδTCR intrinsically combines innate immunity and adaptive immunity by using spatially distinct regions to discriminate non-clonal agonist-selecting elements from clone-specific ligands. The broader implications for antigen-receptor biology are considered.
ABSTRACTReliable identification of near‐native poses of docked protein–protein complexes is still an unsolved problem. The intrinsic heterogeneity of protein–protein interactions is challenging for traditional biophysical or knowledge based potentials and the identification of many false positive binding sites is not unusual. Often, ranking protocols are based on initial clustering of docked poses followed by the application of an energy function to rank each cluster according to its lowest energy member. Here, we present an approach of cluster ranking based not only on one molecular descriptor (e.g., an energy function) but also employing a large number of descriptors that are integrated in a machine learning model, whereby, an extremely randomized tree classifier based on 109 molecular descriptors is trained. The protocol is based on first locally enriching clusters with additional poses, the clusters are then characterized using features describing the distribution of molecular descriptors within the cluster, which are combined into a pairwise cluster comparison model to discriminate near‐native from incorrect clusters. The results show that our approach is able to identify clusters containing near‐native protein–protein complexes. In addition, we present an analysis of the descriptors with respect to their power to discriminate near native from incorrect clusters and how data transformations and recursive feature elimination can improve the ranking performance. Proteins 2017; 85:528–543. © 2016 Wiley Periodicals, Inc.
We present optimisations applied to a bespoke bio-physical molecular dynamics simulation designed to investigate chromosome condensation. Our primary focus is on domain-specific algorithmic improvements to determining short-range interaction forces between particles, as certain qualities of the simulation render traditional methods less effective. We implement tuned versions of the code for both traditional CPU architectures and the modern many-core architecture found in the Intel Xeon Phi coprocessor and compare their effectiveness. We achieve speed-ups starting at a factor of 10 over the original code, facilitating more detailed and larger-scale experiments.
We present an updated and integrated version of our widely used protein–protein docking and binding affinity benchmarks. The benchmarks consist of non-redundant, high-quality structures of protein–protein complexes along with the unbound structures of their components. Fifty-five new complexes were added to the docking benchmark, 35 of which have experimentally measured binding affinities. These updated docking and affinity benchmarks now contain 230 and 179 entries, respectively. In particular, the number of antibody–antigen complexes has increased significantly, by 67% and 74% in the docking and affinity benchmarks, respectively. We tested previously developed docking and affinity prediction algorithms on the new cases. Considering only the top 10 docking predictions per benchmark case, a prediction accuracy of 38% is achieved on all 55 cases and up to 50% for the 32 rigid-body cases only. Predicted affinity scores are found to correlate with experimental binding energies up to r = 0.52 overall and r = 0.72 for the rigid complexes.
Mitotic chromosomes were one of the first cell biological structures to be described, yet their molecular architecture remains poorly understood. We have devised a simple biophysical model of a 300 kb-long nucleosome chain, the size of a budding yeast chromosome, constrained by interactions between binding sites of the chromosomal condensin complex, a key component of interphase and mitotic chromosomes. Comparisons of computational and experimental (4C) interaction maps, and other biophysical features, allow us to predict a mode of condensin action. Stochastic condensin-mediated pairwise interactions along the nucleosome chain generate native-like chromosome features and recapitulate chromosome compaction and individualization during mitotic condensation. Higher order interactions between condensin binding sites explain the data less well. Our results suggest that basic assumptions about chromatin behavior go a long way to explain chromosome architecture and are able to generate a molecular model of what the inside of a chromosome is likely to look like.
ABSTRACTWithin the crowded, seemingly chaotic environment of the cell, proteins are still able to find their binding partners. This is achieved via an ensemble of trajectories, which funnel them towards their functional binding sites, the binding funnel. Here, we characterize funnel‐like energy structures on the global energy landscape using time‐homogeneous finite state Markov chain models. These models are based on the idea that transitions can occur between structurally similar docking solutions, with transition probabilities determined by their difference in binding energy. Funnel‐like energy structures are those containing solutions with very high equilibrium populations. Although these are found surrounding both near‐native and false positive binding sites, we show that the removal of nonfunnel‐like energy structures, by filtering away solutions with low maximum equilibrium population, can significantly improve the ranking of docked poses. Proteins 2013; 81:2143–2149. © 2013 Wiley Periodicals, Inc.
Protein-protein interactions are central to almost all biological functions, and the atomic details of such interactions can yield insights into the mechanisms that underlie these functions. We present a web server that wraps and extends the SwarmDock flexible protein-protein docking algorithm. After uploading PDB files of the binding partners, the server generates low energy conformations and returns a ranked list of clustered docking poses and their corresponding structures. The user can perform full global docking, or focus on particular residues that are implicated in binding. The server is validated in the CAPRI blind docking experiment, against the most current docking benchmark, and against the ClusPro docking server, the highest performing server currently available.Availability: The server is freely available and can be accessed at: http://bmm.cancerresearchuk.org/%7ESwarmDock/.Contact: Paul.Bates@cancer.org.ukSupplementary information: Supplementary data are available at Bioinformatics online.
Abduction and induction are two forms of reasoning that have been widely used in machine learning. The combination of abduction and induction has recently been explored from a number of angles, one of which is the area of systems biology. The research reported in this article is being conducted as part of the MetaLog project, which aims to build causal models of the actions of toxins from empirical data in the form of nuclear magnetic resonance (NMR) data, together with information on networks of known metabolic reactions from the Kyoto Encyclopedia of Genes and Genomes (KEGG). The NMR spectra provide information concerning the flux of metabolite concentrations before, during, and after administration of a toxin
In previous CAPRI rounds (3–5) we showed that using MD‐generated ensembles, as inputs for a rigid‐body docking algorithm, increased our success rate, especially for targets exhibiting substantial amounts of induced fit. In recent rounds (6–11), our cross‐docking was followed by a short MD‐based local refinement for the subset of solutions with the lowest interaction energies after minimization. The above approach showed promising results for target 20, where we were able to recover 30% of native contacts for one of our submitted models. Further tests, performed a posteriori, revealed that cross‐docking approach produces more near‐native (NN) solutions but only for targets with large conformational changes upon binding. However, at the time of the blind docking experiment, these improved solutions were not chosen for the subsequent refinement, as their interaction energies after minimization ranked poorly compared with other solutions. This indicates deficiencies in the present scoring schemes that are based on interaction energies of minimized structures. Refinement MD simulations substantially increase the fraction of native contacts for NN docked solutions, but generally worsen interface and ligand RMSD. Further analysis shows that although MD simulations are able to improve sidechain packing across the interface, which results in an increased fraction of native contacts, they are not capable of improving interface and ligand backbone RMSD for NN structures beyond 1.5 and 3.5 Å, respectively, even if explicit solvent is used. Proteins 2007. © 2007 Wiley‐Liss, Inc.
In this paper we use a logic-based representation and a combination of Abduction and Induction to model inhibition in metabolic networks. In general, the integration of abduction and induction is required when the following two conditions hold. Firstly, the given background knowledge is incomplete. Secondly, the problem must require the learning of general rules in the circumstance in which the hypothesis language is disjoint from the observation language. Both these conditions hold in the application considered in this paper. Inhibition is very important from the therapeutic point of view since many substances designed to be used as drugs can have an inhibitory effect on other enzymes. Any system able to predict the inhibitory effect of substances on the metabolic network would therefore be very useful in assessing the potential harmful side-effects of drugs. In modelling the phenomenon of inhibition in metabolic networks, background knowledge is used which describes the network topology and functional classes of inhibitors and enzymes. This background knowledge, which represents the present state of understanding, is incomplete. In order to overcome this incompleteness hypotheses are considered which consist of a mixture of specific inhibitions of enzymes (ground facts) together with general (non-ground) rules which predict classes of enzymes likely to be inhibited by the toxin. The foreground examples are derived from in vivo experiments involving NMR analysis of time-varying metabolite concentrations in rat urine following injections of toxins. The model's performance is evaluated on training and test sets randomly generated from a real metabolic network. It is shown that even in the case where the hypotheses are restricted to be ground, the predictive accuracy increases with the number of training examples and in all cases exceeds the default (majority class). Experimental results also suggest that when sufficient training data is provided, non-ground hypotheses show a better predictive accuracy than ground hypotheses. The model is also evaluated in terms of the biological insight that it provides.