Establishing a coherent mapping of the relationships among all known proteins is crucial for elucidating processes of protein emergence and evolution. Yet the capacity to fully capture relationships of protein similarity is complicated by the nonstraightforward interplay between sequence and structure; indeed, proteins with unrelated sequences can adopt similar structures, and, conversely, proteins with similar or identical sequences can manifest radically different structures. Here, we introduce Contrastive Learning Sequence-Structure (CLSS), a contrastive protein language model (PLM) trained to coembed sequence and structure information in a self-supervised manner, facilitating a holistic representation of protein relatedness. CLSS represents the structures and sequences of full domains and domain subsequences as vectors in the same high-dimensional latent space. We show that this approach yields meaningful shared representations, which recapitulate the extensive structure- and sequence-based knowledge encoded in human-curated hierarchical protein classification systems (ECOD and CATH). Moreover, the representations generated by CLSS outperform those generated by alternative state-of-the-art PLMs in downstream classification tasks. Notably, we show that even the far larger space of domain subsequences is successfully coembedded, establishing a PLM tailored to these evolutionarily meaningful objects. CLSS embeddings produce informative representations of the protein universe without further downstream processing, as we demonstrate by analyzing preferential associations between protein architectures and ligand types across protein space.
We have seen more progress in computational biology for macromolecules in the last five years than we experienced in the five preceding decades. Thus, it is very challenging to forecast future progress. It is possible that we have reached a plateau, and we will be stuck with similar problems as we have today. Still, it is also possible that the field will continue its rapid progress and completely transform other fields, such as biochemistry, molecular and cell biology, and medicine. It is also possible that general AI will take over, and all scientific endeavours will be conducted without human input. To be honest, we do not know what will happen, but we will highlight a few of the challenges and the most critical research questions that we face today. Hopefully, these will be resolved within the following decades, or hopefully much earlier. Looking back over the last decade, we can see that machine learning and deep learning have become significantly more popular (T-test residual > 2) among the papers published within our section of PlosCB. We do believe that this trend will continue; therefore, we focus on the challenges that must be overcome for it to make significant and notable contributions. The future of computational biology for macromolecules in 20 years is likely to be characterised by transformative advances in accuracy, automation, integration, and explainability, with AI playing a role in one form or another.
A key prerequisite of transporter proteins' function is their trafficking to the target cellular membranes where they fulfill distinct physiological roles. Cornichon proteins (CNIH/Erv14) represent a highly conserved family of coat protein complex II (COPII)-coated vesicle cargo receptors that facilitate the exit of numerous transporters from the endoplasmic reticulum (ER) to proceed via the secretory pathway. Despite their biomedical significance, the cargo specificities of the four human cornichons (CNIH1-4) remain largely unexplored. Here, we conducted a bioinformatics analysis of the CNIH/Erv14 family, revealing evolutionary conservation profiles of the family based on an alignment of 1879 sequences. AlphaFold3 modeling predicts that residues identified as the most evolutionarily conserved in cornichon family interact with Sec24 proteins of COPII vesicles. We also demonstrate the suitability of the model yeast Saccharomyces cerevisiae for studying the properties and putative interactors of human cornichons. We engineered S. cerevisiae strains in which the endogenous cornichon gene (ERV14) was replaced with human CNIH1, CNIH2, or CNIH4 coding sequences or CNIH coding sequences were expressed from multi-copy plasmids. The studied human cornichons were functional in S. cerevisiae cells and, to varying extents, complemented the differing phenotypes related to yeast ScErv14 roles in monovalent-cation homeostasis. The presence of human CNIHs supported the functioning of the yeast plasma-membrane Na+, K+/H+ antiporter Nha1, a known cargo of ScErv14. Both yeast ScErv14 and human CNIH cornichons improved the plasma-membrane targeting and functioning of the human Na+/H+ antiporter NHA2 in yeast cells, identifying NHA2 as a novel cargo of cornichon COPII cargo receptors.
Metal ions are essential for a broad range of biochemical processes in living organisms, with zinc being the second most abundant transition metal ion. Zinc has catalytic, structural, and regulatory functions in proteins, impacting virtually all aspects of cell biology. Currently, there are notable challenges in performing a large-scale accurate systematic analysis of the as-yet unexplored occurrences of zinc ion in nature. To address this, we developed ZincSight for predicting zinc-binding sites. ZincSight performs on par with existing structure-based tools in terms of the precision-recall curve for zinc ion detection, and the accuracy of spatial positioning of the ion, yet it is significantly faster, and offers a straightforward reasoning for its predictions, which is missing even in the best alternatives. Tests using a panel of metals show that, while trained on zinc-binding sites, ZincSight in fact detects all transition metal binding sites alike - a reflection of the similarity in coordination among the transition metals. It also detects binding sites for calcium and other alkaline-earth metals with lower accuracy, but not alkali metal binding sites. Suitable for exploring the usage of zinc and other transition metals in large sets of protein structures, or models thereof, ZincSight is available as a free-to-download open-source software at: https://github.com/MECHTI1/ZincSight. A Google Colab notebook is available at: https://colab.research.google.com/github/MECHTI1/ZincSight/blob/master/ZincSight.ipynb.
Amino acid sequence dictates the three-dimensional structure and biological function of proteins. Yet, despite decades of research, our understanding of the interplay between sequence and structure is incomplete. To meet this challenge, we introduce Contrastive Learning Sequence-Structure (CLSS), an AI-based contrastive learning model trained to co-embed sequence and structure information in a self-supervised manner. We trained CLSS on large and diverse sets of protein building blocks called domains. CLSS represents both sequences and structures as vectors in the same high-dimensional space, where distance relates to sequence-structure similarity. Thus, CLSS provides a natural way to represent the protein universe, reflecting evolutionary relationships, as well as structural changes. We find that CLSS refines expert knowledge about the global organization of protein space, and highlights transitional forms that resist hierarchical classification. CLSS reveals linkage between domains of seemingly separate lineages, thereby significantly improving our understanding of evolutionary design. ### Competing Interest Statement The authors have declared no competing interest. Israel Science Foundation, https://ror.org/04sazxf24, 1764/21 HSFP, RGEC29/2025
Recent findings increasingly suggest the emergence of proteins by mix and match of short peptides, or 'building blocks'. What are these building blocks, and how did they evolve into contemporary proteins? We review two complementary approaches to tackling these questions. First, a bottom-up approach that involves identifying putative components of primordial peptides, and the synthetic routes through which these peptides may have emerged. Second, searches in protein space to reveal building blocks that make up the contemporary protein repertoire; proteins that are not closely related to one another may nevertheless have certain parts in common, suggesting common ancestry. Identifying such shared building blocks, and characterizing their functions, can shed light on the ancient molecules from which proteins emerged, and hint at the mechanisms that govern their evolution. A key challenge lies in merging these two approaches to create a cohesive narrative of how proteins emerged and continue to evolve.
How does enzymatic activity emerge? To shed light on this fundamental question, we study type B dihydrofolate reductases (DfrB), which were discovered for their role in antibiotic resistance. These rudimentary enzymes are evolutionarily distinct from the ubiquitous, monomeric FolA dihydrofolate reductases targeted by the antibiotic trimethoprim. DfrB is unique: it homotetramerizes to form a highly symmetrical central tunnel that accommodates its substrates in close proximity and the right orientation, thus promoting the metabolically essential production of tetrahydrofolate. It is the only known enzyme built from the ancient Src Homology 3 fold, typically a binding module. Strikingly, by studying the evolution of this enzyme family, we observe that no active-site residues are conserved across catalytically active homologs. Integrating experimental and computational analyses, we identify an intricate relationship between homotetramerization and catalytic activity, where formation of a tunnel featuring positive electrostatic potential proves to be a powerful predictor of activity. We demonstrate that the DfrB enzymes have not evolved in response to the synthetic antibiotic to which they confer strong resistance, and propose that DfrB domains evolved the capacity for rudimentary catalysis from a binding capacity. That (rudimentary) catalysis can emerge from the homotetramerization of a binding domain, and that it has been recently recruited by pathogenic bacteria, manifests the opportunistic nature of evolution.
Transition metals (e.g., Fe2/3+, Zn2+, Mn2+) are essential enzymatic cofactors in all organisms. Their environmental scarcity led to the evolution of high-affinity uptake systems. Our research focuses on two bacterial manganese ABC importers, Streptococcus pneumoniae PsaBC and Bacillus anthracis MntBC, both critical for virulence. Both importers share a similar homodimeric structure, where each protomer comprises a transmembrane domain (TMD) linked to a cytoplasmic nucleotide-binding domain (NBD). Due to their size and slow turnover rates, the utility of conventional molecular simulation approaches to reveal functional dynamics is limited. Thus, we employed a novel, computationally efficient method integrating Gaussian Network Models (GNM) with information theory Transfer Entropy (TE) calculations. Our calculations are in remarkable agreement with previous functional studies. Furthermore, based on the calculations, we generated 10 point-mutations and experimentally tested their effects, finding excellent concordance between computational predictions and experimental results. We identified "allosteric hotspots" in both transporters, in the transmembrane translocation pathway, at the coupling helices linking the TMDs and NBDs, and in the ATP binding sites. In both PsaBC and MntBC, we observed bi-directional information flow between the two TMDs, with minimal allosteric transmission to the NBDs. Conversely, the NBDs exhibited almost no NBD-NBD allosteric crosstalk but showed pronounced information flow from the NBD of one protomer towards the TMD of the other protomer. This unique allosteric "footprint" distinguishes ABC importers of transition metals from other members of the ABC transporter superfamily establishing them as a distinct functional class. This study offers the first comprehensive insight into the conformational dynamics of these vital virulence determinants, providing potential avenues for developing urgently needed novel antibacterial agents.
Clonal propagation of plants by induction of adventitious roots (ARs) from stem cuttings is a requisite step in breeding programs. A major barrier exists for propagating valuable plants that naturally have low capacity to form ARs. Due to the central role of auxin in organogenesis, indole-3-butyric acid is often used as part of commercial rooting mixtures, yet many recalcitrant plants do not form ARs in response to this treatment. Here we describe the synthesis and screening of a focused library of synthetic auxin conjugates in Eucalyptus grandis cuttings and identify 4-chlorophenoxyacetic acid–l-tryptophan-OMe as a competent enhancer of adventitious rooting in a number of recalcitrant woody plants, including apple and argan. Comprehensive metabolic and functional analyses reveal that this activity is engendered by prolonged auxin signaling due to initial fast uptake and slow release and clearance of the free auxin 4-chlorophenoxyacetic acid. This work highlights the utility of a slow-release strategy for bioactive compounds for more effective plant growth regulation. Adventitious roots are induced in various woody plants, enabling clonal propagation.
Protein space is characterized by extensive recurrence, or "reuse," of parts, suggesting that new proteins and domains can evolve by mixing-and-matching of existing segments. From an evolutionary perspective, for a given combination to persist, the protein segments should presumably not only match geometrically but also dynamically communicate with each other to allow concerted motions that are key to function. Evidence from protein space supports the premise that domains indeed combine in this manner; we explore whether a similar phenomenon can be observed at the sub-domain level. To this end, we use Gaussian Network Models (GNMs) to calculate the so-called soft modes, or low-frequency modes of motion for a dataset of 150 protein domains. Modes of motion can be used to decompose a domain into segments of consecutive amino acids that we call "dynamic elements", each of which belongs to one of two parts that move in opposite senses. We find that, in many cases, the dynamic elements, detected based on GNM analysis, correspond to established "themes": Sub-domain-level segments that have been shown to recur in protein space, and which were detected in previous research using sequence similarity alone (i.e. completely independently of the GNM analysis). This statistically significant correlation hints at the importance of dynamics in evolution. Overall, the results are consistent with an evolutionary scenario where proteins have emerged from themes that need to match each other both geometrically and dynamically, e.g. to facilitate allosteric regulation.
By far, the biggest recent breakthrough for studying protein evolution has been the advent of new computational methods for predicting protein structures and assemblies. This development is proving transformational for evolutionary studies because proteins' amino acid sequences evolve faster than their 3D structures. Predicting 3D structures with accuracy lets us see farther back into evolutionary time and find more with our ability to sequence DNA, thanks to tools like AlphaFold, RosettaFold, ESMfold, and others. Within just 3 years, we've moved from having detailed structures for hundreds of thousands of proteins to predicting hundreds of millions with reasonable accuracy. This leap has been enhanced by new, fast structure comparison algorithms, such as FoldSeek, that let us use these vast datasets effectively. With these tools in hand, we can explore deep time, uncovering entirely new protein families, 3D folds, least understood organisms on the planet: think of all of the strange fungi, uncultivatable and viruses, known only by DNA sequencing. We've essentially just been given entirely new computational lenses to peer into their proteomes and begin understanding some
Cation/proton antiporters (CPAs) regulate cells' salt concentration and pH. Their malfunction is associated with a range of human pathologies, yet only a handful of CPA-targeting therapeutics are presently in clinical development. Here, we discuss how recently published mammalian protein structures and emerging computational technologies may help to bridge this gap.
The ConSurf web-sever for the analysis of proteins, RNA, and DNA provides a quick and accurate estimate of the per-site evolutionary rate among homologues. The analysis reveals functionally important regions, such as catalytic and ligand-binding sites, which often evolve slowly. Since the last report in 2016, ConSurf has been improved in multiple ways. It now has a user-friendly interface that makes it easier to perform the analysis and to visualize the results. Evolutionary rates are calculated based on a set of homologous sequences, collected using hidden Markov model-based search tools, recently embedded in the pipeline. Using these, and following the removal of redundancy, ConSurf assembles a representative set of effective homologues for protein and nucleic acid queries to enable informative analysis of the evolutionary patterns. The analysis is particularly insightful when the evolutionary rates are mapped on the macromolecule structure. In this respect, the availability of AlphaFold model structures of essentially all UniProt proteins makes ConSurf particularly relevant to the research community. The UniProt ID of a query protein with an available AlphaFold model can now be used to start a calculation. Another important improvement is the Python re-implementation of the entire computational pipeline, making it easier to maintain. This Python pipeline is now available for download as a standalone version. We demonstrate some of ConSurf's key capabilities by the analysis of caveolin-1, the main protein of membrane invaginations called caveolae.
Malfunction of the CFTR protein results in cystic fibrosis, one of the most common hereditary diseases. CFTR functions as an anion channel, the gating of which is controlled by long-range allosteric communications. Allostery also has direct bearings on CF treatment: the most effective CFTR drugs modulate its activity allosterically. Herein, we integrated Gaussian network model, transfer entropy, and anisotropic normal mode-Langevin dynamics and investigated the allosteric communications network of CFTR. The results are in remarkable agreement with experimental observations and mutational analysis and provide extensive novel insight. We identified residues that serve as pivotal allosteric sources and transducers, many of which correspond to disease-causing mutations. We find that in the ATP-free form, dynamic fluctuations of the residues that comprise the ATP-binding sites facilitate the initial binding of the nucleotide. Subsequent binding of ATP then brings to the fore and focuses on dynamic fluctuations that were present in a latent and diffuse form in the absence of ATP. We demonstrate that drugs that potentiate CFTR's conductance do so not by directly acting on the gating residues, but rather by mimicking the allosteric signal sent by the ATP-binding sites. We have also uncovered a previously undiscovered allosteric 'hotspot' located proximal to the docking site of the phosphorylated regulatory (R) domain, thereby establishing a molecular foundation for its phosphorylation-dependent excitatory role. This study unveils the molecular underpinnings of allosteric connectivity within CFTR and highlights a novel allosteric 'hotspot' that could serve as a promising target for the development of novel therapeutic interventions.
Multiple sequence alignments (MSAs) are the workhorse of molecular evolution and structural biology research. From MSAs, the amino acids that are tolerated at each site during protein evolution can be inferred. However, little is known regarding the repertoire of tolerated amino acids in proteins when only a few or no sequence homologs are available, such as orphan and de novo designed proteins. Here we present EvoRator2, a deep-learning algorithm trained on over 15,000 protein structures that can predict which amino acids are tolerated at any given site, based exclusively on protein structural information mined from atomic coordinate files. We show that EvoRator2 obtained satisfying results for the prediction of position-weighted scoring matrices (PSSM). We further show that EvoRator2 obtained near state-of-the-art performance on proteins with high quality structures in predicting the effect of mutations in deep mutation scanning (DMS) experiments and that for certain DMS targets, EvoRator2 outperformed state-of-the-art methods. We also show that by combining EvoRator20s predictions with those obtained by a state-of-the-art deep-learning method that accounts for the information in the MSA, the prediction of the effect of mutation in DMS experiments was improved in terms of both accuracy and stability. EvoRator2 is designed to predict which amino-acid substitutions are tolerated in such proteins without many homologous sequences, including orphan or de novo designed proteins. We implemented our approach in the EvoRator web server (https://evorator.tau.ac.il).(c) 2023 Published by Elsevier Ltd.
The opportunistic fungus Aspergillus fumigatus is the primary invasive mold pathogen in humans, and is responsible for an estimated 200,000 yearly deaths worldwide. Most fatalities occur in immunocompromised patients who lack the cellular and humoral defenses necessary to halt the pathogen's advance, primarily in the lungs. One of the cellular responses used by macrophages to counteract fungal infection is the accumulation of high phagolysosomal Cu levels to destroy ingested pathogens. A. fumigatus responds by activating high expression levels of crpA, which encodes a Cu+ P-type ATPase that actively transports excess Cu from the cytoplasm to the extracellular environment. In this study, we used a bioinformatics approach to identify two fungal-unique regions in CrpA that we studied by deletion/replacement, subcellular localization, Cu sensitivity in vitro, killing by mouse alveolar macrophages, and virulence in a mouse model of invasive pulmonary aspergillosis. Deletion of CrpA fungal-unique amino acids 1-211 containing two N-terminal Cu-binding sites, moderately increased Cu-sensitivity but did not affect expression or localization to the endoplasmic reticulum (ER) and cell surface. Replacement of CrpA fungal-unique amino acids 542-556 consisting of an intracellular loop between the second and third transmembrane helices resulted in ER retention of the protein and strongly increased Cu-sensitivity. Deleting CrpA N-terminal amino acids 1-211 or replacing amino acids 542-556 also increased sensitivity to killing by mouse alveolar macrophages. Surprisingly, the two mutations did not affect virulence in a mouse model of infection, suggesting that even weak Cu-efflux activity by mutated CrpA preserves fungal virulence.
Amyloids, protein, and peptide assemblies in various organisms are crucial in physiological and pathological processes. Their intricate structures, however, present significant challenges, limiting our understanding of their functions, regulatory mechanisms, and potential applications in biomedicine and technology. This study evaluated the AlphaFold2 ColabFold method's structure predictions for antimicrobial amyloids, using eight antimicrobial peptides (AMPs), including those with experimentally determined structures and AMPs known for their distinct amyloidogenic morphological features. Additionally, two well-known human amyloids, amyloid-β and islet amyloid polypeptide, were included in the analysis due to their disease relevance, short sequences, and antimicrobial properties. Amyloids typically exhibit tightly mated β-strand sheets forming a cross-β configuration. However, certain amphipathic α-helical subunits can also form amyloid fibrils adopting a cross-α structure. Some AMPs in the study exhibited a combination of cross-α and cross-β amyloid fibrils, adding complexity to structure prediction. The results showed that the AlphaFold2 ColabFold models favored α-helical structures in the tested amyloids, successfully predicting the presence of α-helical mated sheets and a hydrophobic core resembling the cross-α configuration. This implies that the AI-based algorithms prefer assemblies of the monomeric state, which was frequently predicted as helical, or capture an α-helical membrane-active form of toxic peptides, which is triggered upon interaction with lipid membranes.