Many protein-protein interactions (PPIs) are mediated by the binding of short linear motifs (SLiMs) to peptide recognition domains (PRDs). Here, we describe PrePPI-SLiM, a proteome-scale computational pipeline that leverages data from the Eukaryotic Linear Motif (ELM) database to predict whether two proteins will form a peptide-mediated complex. The ELM database defines classes of protein-peptide interactions with SLiMs represented by sequence motifs and PRDs represented by Pfam domains. PrePPI-SLiM systematically evaluates all pairwise combinations of proteins within a proteome and identifies PRD-SLiM pairs that occur in the same ELM class. This evidence together with disorder prediction and sequence conservation of the motif are integrated in a naïve Bayes framework to assign a likelihood for complex formation. To obtain potential PDB templates for atomistic models of PrePPI-SLiM interactions, we associate individual PPI predictions with homologous PDB complexes involving the same PRD Pfam domain and SLIM, and obtain PDB templates for 92% of our high-confidence predictions. Moreover, studies with AF3Complex suggest that prior knowledge of the interacting PRD and SLiM, as provided here, is a critical starting point for creating a 3D model of the specific sequences of the PRD and SLiM query proteins. Finally, we demonstrate that clustering of the high-confidence PrePPI-SLiM interactome yields functionally coherent PPI networks that reveal mechanistic insights into cellular processes. The PrePPI webserver provides convenient access to high-confidence PrePPI-SLiM predictions, PDB templates for modeling, and functional networks.
PrePPI is a structure-based pipeline that predicts protein-protein interactions (PPIs) between two structured domains and between structured domains and short linear motifs (SLiMs) on a proteome-wide scale. Since the 2023 Computational Resource Issue of JMB, the PrePPI website has been significantly expanded and redesigned. The resource now includes interactomes for human, yeast, and E. coli proteomes with 3D models for high-confidence domain-level complexes and PDB templates for most of the SLiM-mediated predicted interactions. A key new addition is derived from the clustering of the PrePPI interactomes based entirely on the structure-based likelihood of an interaction. Remarkably these clusters exhibit functional coherence and provide an unprecedented proteome-wide depiction of the subnetworks of PPIs that underlie biological phenomena. The new website - https://honigcomplab.c2b2.columbia.edu/PrePPI - provides convenient access to these clusters, to structural models for each pairwise complex, and to function annotations for individual proteins, enabling multiple modes of biological discovery.
A significant portion of disease-causing mutations occur at protein-protein interfaces however, the number of structurally resolved multi-protein complexes is extremely small. Here we present a computational pipeline, PIONEER2.0, that integrates 3D structural similarity with geometric deep learning to accurately predict protein binding partner-specific interfacial residues for all experimentally observed human binary protein-protein interactions. We estimate that AlphaFold3 fails to produce high-quality structural models for about half of the human interactome; for these challenging cases, PIONEER2.0 significantly outperforms AlphaFold3 in predicting their interface residues, making PIONEER2.0 an excellent alternative and complementary tool in real-world applications. We further systematically validated PIONEER2.0 predictions experimentally by generating 1,866 mutations and testing their impact on 5,010 mutation-interaction pairs, confirming PIONEER-predicted interfaces are comparable in accuracy as experimentally determined interfaces using PDB co-complex structures. We then used PIONEER2.0 to create a comprehensive multiscale structurally informed human interactome encompassing all 352,124 experimentally determined binary human protein interactions in the literature. We find that PIONEER2.0-predicted interfaces are instrumental in prioritizing disease-associated mutations and thus provide insight into their underlying molecular mechanisms. Overall, our PIONEER2.0 framework offers researchers a valuable tool at an unprecedented scale for studying disease etiology and advancing personalized medicine.
Aberrant signaling pathway activity is a hallmark of tumorigenesis and progression, which has guided targeted inhibitor design for over 30 years. Yet, adaptive resistance mechanisms, induced by rapid, context-specific signaling network rewiring, continue to challenge therapeutic efficacy. By leveraging progress in proteomic technologies and network-based methodologies, over the past decade, we developed VESPA-an algorithm designed to elucidate mechanisms of cell response and adaptation to drug perturbations-and used it to analyze 7-point phosphoproteomic time series from colorectal cancer cells treated with clinically-relevant inhibitors and control media. Interrogation of tumor-specific enzyme/substrate interactions accurately inferred kinase and phosphatase activity, based on their inferred substrate phosphorylation state, effectively accounting for signal cross-talk and sparse phosphoproteome coverage. The analysis elucidated time-dependent signaling pathway response to each drug perturbation and, more importantly, cell adaptive response and rewiring that was experimentally confirmed by CRISPRko assays, suggesting broad applicability to cancer and other diseases.
We introduce ZEPPI (Z-score Evaluation of Protein-Protein Interfaces), a framework to evaluate structural models of a complex based on sequence coevolution and conservation involving residues in protein-protein interfaces. The ZEPPI score is calculated by comparing metrics for an interface to those obtained from randomly chosen residues. Since contacting residues are defined by the structural model, this obviates the need to account for indirect interactions. Further, although ZEPPI relies on species-paired multiple sequence alignments, its focus on interfacial residues allows it to leverage quite shallow alignments. ZEPPI can be implemented on a proteome-wide scale and is applied here to millions of structural models of dimeric complexes in the Escherichia coli and human interactomes found in the PrePPI database. PrePPI's scoring function is based primarily on the evaluation of protein-protein interfaces, and ZEPPI adds a new feature to this analysis through the incorporation of evolutionary information. ZEPPI performance is evaluated through applications to experimentally determined complexes and to decoys from the CASP-CAPRI experiment. As we discuss, the standard CAPRI scores used to evaluate docking models are based on model quality and not on the ability to give yes/no answers as to whether two proteins interact. ZEPPI is able to detect weak signals from PPI models that the CAPRI scores define as incorrect and, similarly, to identify potential PPIs defined as low confidence by the current PrePPI scoring function. A number of examples that illustrate how the combination of PrePPI and ZEPPI can yield functional hypotheses are provided.
Supplementary File 1. ZEPPI results for 1A: Bacterial PDB heterodimer complexes1B: Human PDB heterodimer complexes 1C: Bacterial PDB homodimer complexes1D: Human PDB homodimer complexes
Supplementary File 2. ZEPPI results for2A: PrePPI-AF predictions for E. coli PPIs, at FPR ≤0.0012B: PrePPI-AF predictions for human PPIs, at FPR ≤0.001
We describe the Predicting Protein-Compound Interactions (PrePCI) database which comprises over 5 billion predicted interactions between 6.8 million chemical compounds and 19,797 human proteins. PrePCI relies on a proteome-wide database of structural models based on both traditional modeling techniques and the AlphaFold Protein Structure Database. Sequence- and structural similarity-based metrics are established between template proteins, T, in the Protein Data Bank that bind compounds, C, and query proteins in the model database, Q. When the metrics exceed threshold values, it is assumed that C also binds to Q with a likelihood ratio (LR) derived from machine learning. If the relationship is based on structural similarity, the LR is based on a scoring function that measures the extent to which C is compatible with the binding site of Q as described in the LT-scanner algorithm. For every predicted complex derived in this way, chemical similarity based on the Tanimoto coefficient identifies other small molecules that may bind to Q. An overall LR for the binding of C to Q is obtained from Naive Bayesian statistics. The PrePCI database can be queried by entering a UniProt ID or gene name for a protein to obtain a list of compounds predicted to bind to it along with associated LRs. Alternatively, entering an identifier for the compound outputs a list of proteins it is predicted to bind. Specific applications of the database to lead discovery, elucidation of drug mechanism of action, and biological function annotation are described.
Predicting protein-protein interactions (PPI) is a challenging problem of central importance in fundamental biology. With the increasing number of available PPI prediction methods and databases, an effective evaluation model would be extremely valuable. Here we introduce ZIPPI (Z-score for Information about Protein-Protein Interfaces), which evaluates structural models of a complex based on sequence co-evolution and conservation involving residues that are in contact in the interface. The interface Z-score (ZIPPI score) is calculated by comparing metrics for interface contacts to metrics obtained from randomly chosen surface residues. Since contacting residues are defined by the structural model, this obviates the need of accounting for indirect interactions with methods such as Direct Coupling Analysis. Although ZIPPI relies on species-paired multiple sequence alignments, its focus on contacting interfacial residues and the avoidance of direct coupling methods makes it computationally efficient. The performance of ZIPPI is evaluated through applications to experimentally determined complexes from the Protein Data Bank (PDB) and to decoys from the Critical Assessment of PRedicted Interactions (CAPRI) experiment. We demonstrate how ZIPPI can be implemented on a genome-wide scale by calculating scores for millions of structural models of protein-protein interactions in the E. coli interactome as predicted by PrePPI. Many PrePPI predictions filtered by ZIPPI score are novel. In all, this proteome-scale method shows promising feasibility for applications to the full human protein interactome, which is not yet accessible to deep learning methods.
We present an updated version of the Predicting Protein-Protein Interactions (PrePPI) webserver which predicts PPIs on a proteome-wide scale. PrePPI combines structural and non-structural clues within a Bayesian framework to compute a likelihood ratio (LR) for essentially every possible pair of proteins in a proteome; the current database is for the human interactome. The structural modeling (SM) clue is derived from templatebased modeling and its application on a proteome-wide scale is enabled by a unique scoring function used to evaluate a putative complex. The updated version of PrePPI leverages AlphaFold structures that are parsed into individual domains. As has been demonstrated in earlier applications, PrePPI performs extremely well as measured by receiver operating characteristic curves derived from testing on E. coli and human protein-protein interaction (PPI) databases. A PrePPI database of ~1.3 million human PPIs can be queried with a webserver application that comprises multiple functionalities for examining query proteins, template complexes, 3D models for predicted complexes, and related features ( https://honiglab.c2b2.columbia.edu/PrePPI ). PrePPI is a state-of- the-art resource that offers an unprecedented structure-informed view of the human interactome.Graphic Abstract:
Systems biology is a data-heavy field that focuses on systems-wide depictions of biological phenomena necessarily sacrificing a detailed characterization of individual components. As an example, genome-wide protein interaction networks are widely used in systems biology and continuously extended and refined as new sources of evidence become available. Despite the vast amount of information about individual protein structures and protein complexes that has accumulated in the past 50 years in the Protein Data Bank, the data, computational tools, and language of structural biology are not an integral part of systems biology. However, increasing effort has been devoted to this integration, and the related literature is reviewed here. Relationships between proteins that are detected via structural similarity offer a rich source of information not available from sequence similarity, and homology modeling can be used to leverage Protein Data Bank structures to produce 3D models for a significant fraction of many proteomes. A number of structure-informed genomic and cross-species (i.e., virus-host) interactomes will be described, and the unique information they provide will be illustrated with a number of examples. Tissue- and tumor-specific interactomes have also been developed through computational strategies that exploit patient information and through genetic interactions available from increasingly sensitive screens. Strategies to integrate structural information with these alternate data sources will be described. Finally, efforts to link protein structure space with chemical compound space offer novel sources of information in drug design, off-target identification, and the identification of targets for compounds found to be effective in phenotypic screens.
Physical sciences are often overlooked in the field of cancer research. The Physical Sciences in Oncology Initiative was launched to integrate physics, mathematics, chemistry, and engineering with cancer research and clinical oncology through education, outreach, and collaboration. Here, we provide a framework for education and outreach in emerging transdisciplinary fields.
Tumor-specific elucidation of physical and functional oncoprotein interactions could improve tumorigenic mechanism characterization and therapeutic response prediction. Current interaction models and pathways, however, lack context specificity and are not oncoprotein specific. We introduce SigMaps as context-specific networks, comprising modulators, effectors and cognate binding-partners of a specific oncoprotein. SigMaps are reconstructed de novo by integrating diverse evidence sources—including protein structure, gene expression and mutational profiles—via the OncoSig machine learning framework. We first generated a KRAS-specific SigMap for lung adenocarcinoma, which recapitulated published KRAS biology, identified novel synthetic lethal proteins that were experimentally validated in three-dimensional spheroid models and established uncharacterized crosstalk with RAB/RHO. To show that OncoSig is generalizable, we first inferred SigMaps for the ten most mutated human oncoproteins and then for the full repertoire of 715 proteins in the COSMIC Cancer Gene Census. Taken together, these SigMaps show that the cell’s regulatory and signaling architecture is highly tissue specific.
Abstract Breast cancer is one of the most common cancers among women worldwide. Recent studies suggest that cancer cell homeostatic stable phenotypes may be implemented, controlled and maintained by Master Regulators (MR) proteins. In essence, all MR proteins are Transcription Factors, which frequently are intrinsically (IDP) or partially disordered proteins, depending on the extent of disordered regions (IDRs). IDPs/IDRs are known to be promiscuous binders that commonly function as central hubs in protein networks. Furthermore, they contain Molecular Recognition Features (MoRFs) and/or Short Linear Motifs (SLiMs), regions that occupy a unique structural and functional niche in which function is a direct consequence of intrinsic disorder. Gathering information on MR protein interactome has significant implications for understanding the molecular basis of the mechanisms responsible for the stability of tumor cell states. In this work, we focused on a subset of 128 MR proteins predicted from RNA-Seq profiles to be responsible for the Breast Invasive Carcinoma (BRCA) tumor homeostatic control across five distinct cancer subtypes and analyzed their disorder prevalence, functional importance, and potential roles in protein association networks. The disorder profile annotation was performed using the MobiDB-lite disorder meta-predictor. The proteins were classified according to the percentage of intrinsically disordered residues (PIDR) as highly ordered (PIDR < 10%), moderately disordered (10% ≤ PIDR < 30%), and highly disordered (PIDR ≥ 30%). We also evaluated target proteins using the CH-CDF plot, which is a combination of two binary disorder predictor methods: the charge-hydropathy plot (CH), that considers charge and hydrophobicity of amino acids, and the cumulative distribution function (CDF) plot, which is based on the distribution of per-residue disorder predictor scores. Disorder functional annotation was performed by mapping MoRFs and SLiMs using experimentally curated databases. The PrePPI database was used to predict interactivity of the BRCA MR proteins by constructing five cancer-specific protein-protein interaction (PPI) networks. Statistical significance of connectivity for the subtype interactomes was measured by the network-based distance method. Annotation of abundant/ housekeeping genes and network topology (hubs and bottlenecks) was introduced to correlate disorder, gene essentiality, and promiscuous binding activity. The result of this work is a comprehensive, functionally annotated inventory of the intrinsically disordered status of BRCA MR proteins. Our work revealed that the master regulator candidates are enriched in moderate (44%) and high (28%) intrinsic disorder and provide the groundwork for experimental studies that may uncover novel pathways associated with the aberrant protein activity that determine stable tumor states and predict IDPs as novel drug targets. Citation Format: Julia V. Castro, Evan O. Paull, Jose I. Garzon, Diana Murray, Andrea Califano, Barry Honig. Uncovering the disorder of breast invasive carcinoma homeostasis proteins [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2018; 2018 Apr 14-18; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2018;78(13 Suppl):Abstract nr 285.
The dysregulated signaling responsible for tumor cell state implementation and maintenance has been shown to be mediated by the concerted action of master regulator (MR) proteins. Highly connected MR protein modules serve as “tumor checkpoints” by translating upstream genetic alterations into aberrant protein activity that drives pathophysiological cellular phenotypes. Direct pathways for information flow among MR proteins, and between individual MR proteins and their upstream modulators and downstream effectors, are largely unknown. The Prediction of Protein-Protein Interactions (PrePPI) database contains 1.4M predictions for ~85% of the human proteome; of these, ~300K are predicted to be physical, i.e. direct one-to-one, interactions.We apply network analysis concepts, as implemented in the R package igraph, to the PrePPI interactome to predict structural mechanisms underlying tumorigenic signal transduction. Shortest path and random walk with restart algorithms elucidate 1) physical interactions between MR proteins, 2) physical interactions that connect MR proteins that do not directly bind to each other, 3) upstream signaling proteins with recurrent and patient-specific mutations that physically interact with MR proteins, and 4) downstream cofactors and transcription factors that may be involved in effecting a dysregulated phenotype.Our analysis predicts that physical interactions fully connect BACH2, BCL6, IRF8, and SPIB, established germinal center MR proteins. 1) BACH2 and BCL6 are predicted to interact directly through their BTB/POZ domains; BTB domains are known to mediate heterodimeric protein-protein interactions. 2) Additional proteins, including IBTK, ANFY1, PAX7, ZEB1, GABPB1/B2, connect SPIB to BACH2 and to BCL6, in some cases through ankyrin repeats, which are structural protein-protein interaction motifs. Several of the PrePPI-predicted proteins, e.g. IBTK, ZEB1, GABP, have previously been implicated in B-cell biology. The possible implications of our analyses for a wide range of MR proteins modules will be presented.The computational prediction of physical protein-protein interactions within the regulatory architecture of a tumor cell is a powerful approach to discovering novel tumorigenic signaling mechanisms, formulating specific experimentally testable hypotheses of function, including compound mechanisms of action, and prioritizing novel drug targets. Citation Format: Diana Murray, Kamrun N. Begum, Andrea Califano, Barry Honig. Network analysis of the human protein-protein interactome: Tumorigenic signaling mechanisms [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2018; 2018 Apr 14-18; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2018;78(13 Suppl):Abstract nr 3300.
The largely incomplete and tissue-independent nature of cancer pathways represents a key limitation to the ability to elucidate mechanistic determinants of cancer phenotypes and to predict adaptive response to targeted therapy. To address these challenges, we propose replacing canonical cancer pathways with a more accurate, comprehensive, and context-specific architecture – dubbed a Protein-Centric molecular interaction Map (PC-Map) – representing modulators, effectors, and cognate binding-partners of any oncoprotein of interest. To reconstruct these complex molecular architectures de novo , we introduce a novel OncoSig algorithm. Validation of a lung adenocarcinoma specific (LUAD) KRAS-centric PC-Map recapitulated known KRAS biology and, more critically, identified a novel repertoire of proteins eliciting synthetic lethality in KRASG12D LUAD organoid cultures. Showing the generalizable nature of the algorithm, we elucidated PC-Maps for ten recurrently mutated oncoproteins, including KRAS, in distinct tumor contexts. This revealed a highly context-specific nature of cancer’s regulatory and signaling architectures to an unprecedented degree of resolution.
We report a template-based method, LT-scanner, which scans the human proteome using protein structural alignment to identify proteins that are likely to bind ligands that are present in experimentally determined complexes. A scoring function that rapidly accounts for binding site similarities between the template and the proteins being scanned is a crucial feature of the method. The overall approach is first tested based on its ability to predict the residues on the surface of a protein that are likely to bind small-molecule ligands. The algorithm that we present, LBias, is shown to compare very favorably to existing algorithms for binding site residue prediction. LT-scanner's performance is evaluated based on its ability to identify known targets of Food and Drug Administration (FDA)-approved drugs and it too proves to be highly effective. The specificity of the scoring function that we use is demonstrated by the ability of LT-scanner to identify the known targets of FDA-approved kinase inhibitors based on templates involving other kinases. Combining sequence with structural information further improves LT-scanner performance. The approach we describe is extendable to the more general problem of identifying binding partners of known ligands even if they do not appear in a structurally determined complex, although this will require the integration of methods that combine protein structure and chemical compound databases.
Abstract The Cancer Systems Therapeutics (CaST) Outreach helps foster community building throughout the NCI Cancer Systems Biology Consortium (CSBC). The overall mission of the CaST Outreach Core is to advance progress in cancer systems biology by tapping into and making connections across research talent at all stages of educational and professional attainment. A unique component is our partnership with motivated undergraduate, post-baccalaureate, and graduate students from Brooklyn College of the City University of New York (CUNY). These students comprise the first cohort of the CaST Scholars Program and are immersed in the cutting-edge biological and biomedical science underlying Cancer Systems Biology. In subsequent years, CaST Outreach will partner with students from other CUNY Senior Colleges while continuing interactions with our Brooklyn College Scholars. This integrative approach will create a network of CaST Scholars throughout the CUNY system. Because Systems Biology is highly interdisciplinary, participants in the program can come from many backgrounds, including biology, computer science, physics, chemistry, genomics, engineering, and other fields. The program has no fixed curriculum and each of nine Brooklyn College students has developed an individualized CSBC educational plan tailored to their interests. The CaST Scholars Program opportunities include 1) tuition-free attendance at a master’s level Systems Biology course taught by a CaST investigator and offered through Columbia University’s School of Professional Studies; 2) supported attendance at the AACR annual meeting; 3) monthly informal discussion meetings with CaST investigators; 4) weekly formal tutorials in the design and applications of CaST computational algorithms; 5) research internships with CaST investigators; 6) participation in a wide-range of Center activities; 7) a scholarship to support participation; and 8) ongoing mentorship from CaST investigators. This poster provides summaries of the collaborative, synergistic experiences and achievements of each Scholar and of the group as a whole. The CaST Scholars Program offers a unique opportunity for early-stage scientists to learn how Systems Biology is transforming cancer research and precision medicine at Columbia University Medical Center. Regardless of their chosen professional path, CaST Scholars will obtain a solid foundation in the methods and potential of Cancer Systems Biology. Citation Format: Rahimah Ahmad, Tasnim Azad, Carlos Barreto, Kamrun Begum, Mahlaqa Butt, Daniel Gruffat, Aulon Jerliu, Jonathan Kwiat, Mikaela Murph, Katherine A. Rivera Gómez, Barry Honig, Andrea Califano, Shaneen Singh, Diana Murray. Columbia University’s Center for Cancer Systems Therapeutics (CaST) 2017 Scholars Program: A synergistic partnership with students from Brooklyn College, CUNY [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2017; 2017 Apr 1-5; Washington, DC. Philadelphia (PA): AACR; Cancer Res 2017;77(13 Suppl):Abstract nr 2609. doi:10.1158/1538-7445.AM2017-2609
Abstract Elucidation of the full complement of members of oncogenic signaling pathways is critical for understanding dysregulated function in cancers. Our knowledge of experimentally verified human gene-regulatory networks and protein-protein interactions remains incomplete, thus limiting their use for predicting novel members of these critical pathways. We present a method that uses a machine learning approach to predict novel members of these aberrantly activated or inactivated pathways. We integrate predictions of protein-protein interactions, gene regulatory networks, post-translational modifications and co-segregation of mutational status and protein activity. This approach produces novel predictions for these pathways. For the Kras and EGFR pathway, these predictions are highly enriched for genes known to by synthetic lethal for these two pathways. Knockdown of top candidates derived from this approach decreased growth in a Kras-activated context, indicating their importance in Kras regulated pathways. Citation Format: Joshua Broyde, David Simpson, Diana Murray, Alexander Lachmann, Federico M. Giorgi, Barry Honig, Alejandro E. Sweet-Cordero, Andrea Califano. Computational detection of oncogene-centric pathway members [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2017; 2017 Apr 1-5; Washington, DC. Philadelphia (PA): AACR; Cancer Res 2017;77(13 Suppl):Abstract nr LB-289. doi:10.1158/1538-7445.AM2017-LB-289