Protein cis-regulatory elements (CREs) are regions that modulate the activity of a protein through intramolecular interactions. Kinases, pivotal enzymes in numerous biological processes, often undergo regulatory control via inhibitory interactions in cis. This study delves into the mechanisms of cis regulation in kinases mediated by CREs, employing a combined structural and sequence analysis. To accomplish this, we curated an extensive dataset of kinases featuring annotated CREs, organized into homolog families through multiple sequence alignments. Key molecular attributes, including disorder and secondary structure content, active and ATP-binding sites, post-translational modifications, and disease-associated mutations, were systematically mapped onto all sequences. Additionally, we explored the potential for conformational changes between active and inactive states. Finally, we explored the presence of these kinases within membraneless organelles and elucidated their functional roles therein. CREs display a continuum of structures, ranging from short disordered stretches to fully folded domains. The adaptability demonstrated by CREs in achieving the common goal of kinase inhibition spans from direct autoinhibitory interaction with the active site within the kinase domain, to CREs binding to an alternative site, inducing allosteric regulation revealing distinct types of inhibitory mechanisms, which we exemplify by archetypical representative systems. While this study provides a systematic approach to comprehend kinase CREs, further experimental investigations are imperative to unravel the complexity within distinct kinase families. The insights gleaned from this research lay the foundation for future studies aiming to decipher the molecular basis of kinase dysregulation, and explore potential therapeutic interventions.
Short Linear Motifs (SLiMs) are the smallest structural and functional components of modular eukaryotic proteins. They are also the most abundant, especially when considering post-translational modifications. As well as being found throughout the cell as part of regulatory processes, SLiMs are extensively mimicked by intracellular pathogens. At the heart of the Eukaryotic Linear Motif (ELM) Resource is a representative (not comprehensive) database. The ELM entries are created by a growing community of skilled annotators and provide an introduction to linear motif functionality for biomedical researchers. The 2024 ELM update includes 346 novel motif instances in areas ranging from innate immunity to both protein and RNA degradation systems. In total, 39 classes of newly annotated motifs have been added, and another 17 existing entries have been updated in the database. The 2024 ELM release now includes 356 motif classes incorporating 4283 individual motif instances manually curated from 4274 scientific publications and including >700 links to experimentally determined 3D structures. In a recent development, the InterPro protein module resource now also includes ELM data. ELM is available at: http://elm.eu.org.
ABSTRACT The modulation of actin polymerization is a common theme among microbial pathogens. Even though microorganisms show a wide repertoire of strategies to subvert the activity of actin, most of them converge in the ones that activate nucleating factors, such as the Arp2/3 complex. Brucella spp. are intracellular pathogens capable of establishing chronic infections in their hosts. The ability to subvert the host cell response is dependent on the capacity of the bacterium to attach, invade, avoid degradation in the phagocytic compartment, replicate in an endoplasmic reticulum-derived compartment and egress. Even though a significant number of mechanisms deployed by Brucella in these different phases have been identified and characterized, none of them have been described to target actin as a cellular component. In this manuscript, we describe the identification of a novel virulence factor (NpeA) that promotes niche formation. NpeA harbors a short linear motif (SLiM) present within an amphipathic alpha helix that has been described to bind the GTPase-binding domain (GBD) of N-WASP and stabilizes the autoinhibited state. Our results show that NpeA is secreted in a Type IV secretion system-dependent manner and that deletion of the gene diminishes the intracellular replication capacity of the bacterium. In vitro and ex vivo experiments demonstrate that NpeA binds N-WASP and that the short linear motif is required for the biological activity of the protein.IMPORTANCEThe modulation of actin-binding effectors that regulate the activity of this fundamental cellular protein is a common theme among bacterial pathogens. The neural Wiskott–Aldrich syndrome protein (N-WASP) is a protein that several pathogens target to hijack actin dynamics. The highly adapted intracellular bacterium Brucella has evolved a wide repertoire of virulence factors that modulate many activities of the host cell to establish successful intracellular replication niches, but, to date, no effector proteins have been implicated in the modulation of actin dynamics. We present here the identification of a virulence factor that harbors a short linear motif (SLiM) present within an amphipathic alpha helix that has been described to bind the GTPase-binding domain (GBD) of N-WASP stabilizing its autoinhibited state. We demonstrate that this protein is a Type IV secretion effector that targets N-WASP-promoting intracellular survival and niche formation.
Structural resolution of protein interactions enables mechanistic and functional studies as well as interpretation of disease variants. However, structural data is still missing for most protein interactions because we lack computational and experimental tools at scale. This is particularly true for interactions mediated by short linear motifs occurring in disordered regions of proteins. We find that AlphaFold-Multimer predicts with high sensitivity but limited specificity structures of domain-motif interactions when using small protein fragments as input. Sensitivity decreased substantially when using long protein fragments or full length proteins. We delineated a protein fragmentation strategy particularly suited for the prediction of domain-motif interfaces and applied it to interactions between human proteins associated with neurodevelopmental disorders. This enabled the prediction of highly confident and likely disease-related novel interfaces, which we further experimentally corroborated for FBXO23-STX1B, STX1B-VAMP2, ESRRG-PSMC5, PEX3-PEX19, PEX3-PEX16, and SNRPB-GIGYF1 providing novel molecular insights for diverse biological processes. Our work highlights exciting perspectives, but also reveals clear limitations and the need for future developments to maximize the power of Alphafold-Multimer for interface predictions.
The pathogenic, tropical Leishmania flagellates belong to an early-branching eukaryotic lineage (Kinetoplastida) with several unique features. Unfortunately, they are poorly understood from a molecular biology perspective, making development of mechanistically novel and selective drugs difficult. Here, we explore three functionally critical targeting short linear motif systems as well as their receptors in depth, using a combination of structural modeling, evolutionary sequence divergence and deep learning. Secretory signal peptides, endoplasmic reticulum (ER) retention motifs (KDEL motifs), and autophagy signals (motifs interacting with ATG8 family members) are ancient and essential components of cellular life. Although expected to be conserved amongst the kinetoplastids, we observe that all three systems show a varying degree of divergence from their better studied equivalents in animals, plants, or fungi. We not only describe their behaviour, but also build models that allow the prediction of localization and potential functions for several uncharacterized Leishmania proteins. The unusually Ala/Val-rich secretory signal peptides, endoplasmic reticulum resident proteins ending in Asp-Leu-COOH and atypical ATG8-like proteins are all unique molecular features of kinetoplastid parasites. Several of their critical protein-protein interactions could serve as targets of selective antimicrobial agents against Leishmaniasis due to their systematic divergence from the host.
DisProt is the primary repository of Intrinsically Disordered Proteins (IDPs). This database is manually curated and the annotations there have strong experimental support. Currently, DisProt contains a relatively small number of proteins highlighting the importance of transferring annotations regarding verified disorder state and corresponding functions to homologous proteins in other species. In such a way, providing them with highly valuable information to better understand their biological roles. While the principles and practicalities of homology transfer are well-established for globular proteins, these are largely lacking for disordered proteins. We used DisProt to evaluate the transferability of the annotation terms to orthologous proteins. For each protein, we looked for their orthologs, with the assumption that they will have a similar function. Then, for each protein and their orthologs, we made multiple sequence alignments (MSAs). Disordered sequences are fast evolving and can be hard to align, therefore, we implemented alignment quality control steps ensuring robust alignments before mapping the annotations. We have designed a pipeline to obtain good-quality MSAs and to transfer annotations from any protein to their orthologs. Applying the pipeline to DisProt proteins, from the 1731 entries with 5623 annotations, we can reach 97,555 orthologs and transfer a total of 301,190 terms by homology. We also provide a web server for consulting the results of DisProt proteins and execute the pipeline for any other protein. The server Homology Transfer IDP (HoTIDP) is accessible at http://hotidp.leloir.org.ar.
The SH2-binding phosphotyrosine class of short linear motifs (SLiMs) are key conditional regulatory elements, particularly in signaling protein complexes beneath the cell's plasma membrane. In addition to transmitting cellular signaling information, they can also play roles in cellular hijack by invasive pathogens. Researchers can take advantage of bioinformatics tools and resources to predict the motifs at conserved phosphotyrosine residues in regions of intrinsically disordered protein. A candidate SH2-binding motif can be established and assigned to one or more of the SH2 domain subgroups. It is, however, not so straightforward to predict which SH2 domains are capable of binding the given candidate. This is largely due to the cooperative nature of the binding amino acids which enables poorer binding residues to be tolerated when the other residues are optimal. High-throughput peptide arrays are powerful tools used to derive SH2 domain-binding specificity, but they are unable to capture these cooperative effects and also suffer from other shortcomings. Tissue and cell type expression can help to restrict the list of available interactors: for example, some well-studied SH2 domain proteins are only present in the immune cell lineages. In this article, we provide a table of motif patterns and four bioinformatics strategies that introduce a range of tools that can be used in motif hunting in cellular and pathogen proteins. Experimental followup is essential to determine which SH2 domain/motif-containing proteins are the actual functional partners.
ABSTRACT The pathogenic tropical flagellates Leishmania belong to an early-branching eukaryotic lineage (Kinetoplastida) with several unique features. Here, we explore three ancient protein targeting linear motif systems and their receptors and demonstrate how they resemble or differ from other eukaryotic organisms, including their hosts. Secretory signal peptides, endoplasmic reticulum (ER) retention motifs (KDEL motifs), and autophagy signals (motifs interacting with ATG8 family members) are essential components of cellular life. Although expected to be conserved, we observe that all three systems show a varying degree of divergence from the eukaryotic version observed in animals, plants, or fungi. We not only describe their behavior but also build predictive models that allow the prediction of localization or function for several proteins in Leishmania species for the first time. Several of these critical protein-protein interactions could serve as targets of selective antimicrobial agents against Leishmaniasis due to their divergence from the host.
Leishmaniasis is a detrimental disease causing serious changes in quality of life and some forms lead to death. The disease is spread by the parasite Leishmania transmitted by sandfly vectors and their primary hosts are vertebrates including humans. The pathogen penetrates host cells and secretes proteins (the secretome) to repurpose cells for pathogen growth and to alter cell signaling via host-pathogen Protein-Protein Interactions (PPIs). Here we present LeishMANIAdb, a database specifically designed to investigate how Leishmania virulence factors may interfere with host proteins. Since the secretomes of different Leishmania species are only partially characterized, we collected various experimental evidence and used computational predictions to identify Leishmania secreted proteins to generate a user-friendly unified web resource allowing users to access all information available on experimental and predicted secretomes. In addition, we manually annotated host-pathogen interactions of 211 proteins, and the localization/function of 3764 transmembrane (TM) proteins of different Leishmania species. We also enriched all proteins with automatic structural and functional predictions that can provide new insights in the molecular mechanisms of infection. Our database, available at https://leishmaniadb.ttk.hu may provide novel insights into Leishmania host-pathogen interactions and help to identify new therapeutic targets for this neglected disease.
DNA mismatch repair (MMR) is essential for correction of DNA replication errors. Germline mutations of the human MMR gene MLH1 are the major cause of Lynch syndrome, a heritable cancer predisposition. In the MLH1 protein, a non-conserved, intrinsically disordered region connects two conserved, catalytically active structured domains of MLH1. This region has as yet been regarded as a flexible spacer, and missense alterations in this region have been considered non-pathogenic. However, we have identified and investigated a small motif (ConMot) in this linker which is conserved in eukaryotes. Deletion of the ConMot or scrambling of the motif abolished mismatch repair activity. A mutation from a cancer family within the motif (p.Arg385Pro) also inactivated MMR, suggesting that ConMot alterations can be causative for Lynch syndrome. Intriguingly, the mismatch repair defect of the ConMot variants could be restored by addition of a ConMot peptide containing the deleted sequence. This is the first instance of a DNA mismatch repair defect conferred by a mutation that can be overcome by addition of a small molecule. Based on the experimental data and AlphaFold2 predictions, we suggest that the ConMot may bind close to the C-terminal MLH1-PMS2 endonuclease and modulate its activation during the MMR process.
The retinoblastoma protein (Rb) and its homologs p107 and p130 are critical regulators of gene expression during the cell cycle and are commonly inactivated in cancer. Rb proteins use their "pocket domain"to bind an LxCxE sequence motif in other proteins, many of which function with Rb proteins to co-regulate transcrip-tion. Here, we present binding data and crystal structures of the p107 pocket domain in complex with LxCxE peptides from the transcriptional co-repressor proteins HDAC1, ARID4A, and EID1. Our results explain why Rb and p107 have weaker affinity for cellular LxCxE proteins compared with the E7 protein from human papil-lomavirus, which has been used as the primary model for understanding LxCxE motif interactions. Our struc-tural and mutagenesis data also identify and explain differences in Rb and p107 affinities for some LxCxE-containing sequences. Our study provides new insights into how Rb proteins bind their cell partners with varying affinity and specificity.
Retinitis pigmentosa (RP) is a genetically heterogeneous form of inherited retinal disease that leads to progressive visual impairment. One genetic subtype of RP, RP54, has been linked to mutations in PCARE (photoreceptor cilium actin regulator). We have recently shown that PCARE recruits WASF3 to the tip of a primary cilium, and thereby activates an Arp2/3 complex which results in the remodeling of actin filaments that drives the expansion of the ciliary tip membrane. On the basis of these findings, and the lack of proper photoreceptor development in mice lacking Pcare, we postulated that PCARE plays an important role in photoreceptor outer segment disk formation. In this study, we aimed to decipher the relationship between predicted structural and function amino acid motifs within PCARE and its function. Our results show that PCARE contains a predicted helical coiled coil domain together with evolutionary conserved binding sites for photoreceptor kinase MAK (type RP62), as well as EVH1 domain-binding linear motifs. Upon deletion of the helical domain, PCARE failed to localize to the cilia. Furthermore, upon deletion of the EVH1 domain-binding motifs separately or together, co-expression of mutant protein with WASF3 resulted in smaller ciliary tip membrane expansions. Finally, inactivation of the lipid modification on the cysteine residue at amino acid position 3 also caused a moderate decrease in the sizes of ciliary tip expansions. Taken together, our data illustrate the importance of amino acid motifs and domains within PCARE in fulfilling its physiological function.
Almost twenty years after its initial release, the Eukaryotic Linear Motif (ELM) resource remains an invaluable source of information for the study of motif-mediated protein-protein interactions. ELM provides a comprehensive, regularly updated and well-organised repository of manually curated, experimentally validated short linear motifs (SLiMs). An increasing number of SLiM-mediated interactions are discovered each year and keeping the resource up-to-date continues to be a great challenge. In the current update, 30 novel motif classes have been added and five existing classes have undergone major revisions. The update includes 411 new motif instances mostly focused on cell-cycle regulation, control of the actin cytoskeleton, membrane remodelling and vesicle trafficking pathways, liquid-liquid phase separation and integrin signalling. Many of the newly annotated motif-mediated interactions are targets of pathogenic motif mimicry by viral, bacterial or eukaryotic pathogens, providing invaluable insights into the molecular mechanisms underlying infectious diseases. The current ELM release includes 317 motif classes incorporating 3934 individual motif instances manually curated from 3867 scientific publications. ELM is available at: http://elm.eu.org.
The first reported receptor for SARS-CoV-2 on host cells was the angiotensin-converting enzyme 2 (ACE2). However, the viral spike protein also has an RGD motif, suggesting that cell surface integrins may be co-receptors. We examined the sequences of ACE2 and integrins with the Eukaryotic Linear Motif (ELM) resource and identified candidate short linear motifs (SLiMs) in their short, unstructured, cytosolic tails with potential roles in endocytosis, membrane dynamics, autophagy, cytoskeleton, and cell signaling. These SLiM candidates are highly conserved in vertebrates and may interact with the μ2 subunit of the endocytosis-associated AP2 adaptor complex, as well as with various protein domains (namely, I-BAR, LC3, PDZ, PTB, and SH2) found in human signaling and regulatory proteins. Several motifs overlap in the tail sequences, suggesting that they may act as molecular switches, such as in response to tyrosine phosphorylation status. Candidate LC3-interacting region (LIR) motifs are present in the tails of integrin β3 and ACE2, suggesting that these proteins could directly recruit autophagy components. Our findings identify several molecular links and testable hypotheses that could uncover mechanisms of SARS-CoV-2 attachment, entry, and replication against which it may be possible to develop host-directed therapies that dampen viral infection and disease progression. Several of these SLiMs have now been validated to mediate the predicted peptide interactions.
The Protein Ensemble Database (PED) (https://proteinensemble.org), which holds structural ensembles of intrinsically disordered proteins (IDPs), has been significantly updated and upgraded since its last release in 2016. The new version, PED 4.0, has been completely redesigned and reimplemented with cutting-edge technology and now holds about six times more data (162 versus 24 entries and 242 versus 60 structural ensembles) and a broader representation of state of the art ensemble generation methods than the previous version. The database has a completely renewed graphical interface with an interactive feature viewer for region-based annotations, and provides a series of descriptors of the qualitative and quantitative properties of the ensembles. High quality of the data is guaranteed by a new submission process, which combines both automatic and manual evaluation steps. A team of biocurators integrate structured metadata describing the ensemble generation methodology, experimental constraints and conditions. A new search engine allows the user to build advanced queries and search all entry fields including cross-references to IDP-related resources such as DisProt, MobiDB, BMRB and SASBDB. We expect that the renewed PED will be useful for researchers interested in the atomic-level understanding of IDP function, and promote the rational, structure-based design of IDP-targeting drugs.
Bacterial pathogens have developed complex strategies to successfully survive and proliferate within their hosts. Throughout the infection cycle, direct interaction with host cells occurs. Many bacteria have been found to secrete proteins, such as effectors and toxins, directly into the host cell with the potential to interfere with cell regulatory processes, either enzymatically or through protein-protein interactions (PPIs). Short linear motifs (SLiMs) are abundant peptide modules in cell signaling proteins. Here, we cover the reported examples of eukaryotic-like SLiM mimicry being used by pathogenic bacteria to hijack host cell machinery and discuss how drugs targeting SLiM-regulated cell signaling networks are being evaluated for interference with bacterial infections. This emerging anti-infective opportunity may become an essential contributor to antibiotic replacement strategies.
Modern biology produces data at a staggering rate. Yet, much of these biological data is still isolated in the text, figures, tables and supplementary materials of articles. As a result, biological information created at great expense is significantly underutilised. The protein motif biology field does not have sufficient resources to curate the corpus of motif-related literature and, to date, only a fraction of the available articles have been curated. In this study, we develop a set of tools and a web resource, 'articles.ELM', to rapidly identify the motif literature articles pertinent to a researcher's interest. At the core of the resource is a manually curated set of about 8000 motif-related articles. These articles are automatically annotated with a range of relevant biological data allowing in-depth search functionality. Machine-learning article classification is used to group articles based on their similarity to manually curated motif classes in the Eukaryotic Linear Motif resource. Articles can also be manually classified within the resource. The 'articles.ELM' resource permits the rapid and accurate discovery of relevant motif articles thereby improving the visibility of motif literature and simplifying the recovery of valuable biological insights sequestered within scientific articles. Consequently, this web resource removes a critical bottleneck in scientific productivity for the motif biology field. Database URL: http://slim.icr.ac.uk/articles/.
The eukaryotic linear motif (ELM) resource is a repository of manually curated experimentally validated short linear motifs (SLiMs). Since the initial release almost 20 years ago, ELM has become an indispensable resource for the molecular biology community for investigating functional regions in many proteins. In this update, we have added 21 novel motif classes, made major revisions to 12 motif classes and added >400 new instances mostly focused on DNA damage, the cytoskeleton, SH2-binding phosphotyrosine motifs and motif mimicry by pathogenic bacterial effector proteins. The current release of the ELM database contains 289 motif classes and 3523 individual protein motif instances manually curated from 3467 scientific publications. ELM is available at: http://elm.eu.org.
Over the past few years, it has become apparent that approximately 35% of the human proteome consists of intrinsically disordered regions. Many of these disordered regions are rich in short linear motifs (SLiMs) which mediate protein-protein interactions. Although these motifs are short and often partially conserved, they are involved in many important aspects of protein function, including cleavage, targeting, degradation, docking, phosphorylation, and other posttranslational modifications. The Eukaryotic Linear Motif resource (ELM) was established over 15 years ago as a repository to store and catalogue the scientific discoveries of motifs. Each motif in the database is annotated and curated manually, based on the experimental evidence gathered from publications. The entries themselves are submitted to ELM by filling in two annotation templates designed for motif class and motif instance annotation. In this protocol, we describe the steps involved in annotating new motifs and how to submit them to ELM.
The postsynaptic density extends across the postsynaptic dendritic spine with Discs large (DLG) as the most abundant scaffolding protein. DLG dynamically alters the structure of the postsynaptic density, thus controlling the function and distribution of specific receptors at the synapse. PDZ domains make up one of the most abundant protein interaction domain families in animals. One important interaction governing postsynaptic architecture is that between the PDZ3 domain from DLG and cysteine-rich interactor of PDZ3 (CRIPT). However, little is know regarding functional evolution of the PDZ3:CRIPT interaction. Here, we subjected PDZ3 and CRIPT to ancestral sequence reconstruction, resurrection and biophysical experiments. We show that the PDZ3:CRIPT interaction is an ancient interaction, which was present in the last common ancestor of Eukaryotes, and that high affinity is maintained in most extant animal phyla. However, affinity is low in nematodes and insects, raising questions about the physiological function of the interaction in species from these animal groups. Our findings demonstrate how an apparently established protein-protein interaction involved in cellular scaffolding in bilaterians can suddenly be subject to dynamic evolution including possible loss of function.