Protein embedding is a protein representation that carries along the information derived from filtering large volumes of sequences stored in large archives. Routinely, the protein is represented by a matrix in which each residue is a context-specific vector whose dimensions reflect the size of the large architectures of neural networks (transformers) trained with deep learning algorithms on large volumes of sequences. A recently introduced method (Embedding-Based Alignment, EBA) is particularly suited for pairwise embedding comparisons and, as we report here, allows for remote homolog detection under specific constraints, including protein sequence length similarity. Multifunctional proteins are present in different species. However, particularly in humans, the problem of their structural and functional annotation is urgent since, according to recent statistics, they comprise up to 50% of the human reference proteome. In this paper we show that when EBA is applied to a set of randomly selected multifunctional human proteins, it retrieves, after a clustering procedure and rigorous validation on the reference Swiss-Prot database, proteins that are remote homologs to each other and carry similar structural and functional features as the query protein.
The human reference proteome is routinely modeled with predictive tools such as AlphaFold2 and ESMFold. The two methods, based on different procedures, can behave differently depending on the experimental information available for a protein. We previously released a public database that stores pairs of predicted models, allowing us to obtain insights into the two methods and providing a resource where users can select the better model for downstream analysis. Here, we update the database after the latest release of UniProt (2025_04), we functionally characterize the models by mapping Pfam entries on the 3D structures, and we introduce external quality assessment metrics to evaluate and compare the models. We observe that, regardless of the quality and similarity of the predicted models, both AlphaFold2 and ESMFold converge with high pLDDT values in regions covered by Pfam entries. Alpha&ESMhFolds, including all its features, is freely available at https://alpha-esmhfolds.biocomp.unibo.it/.
Glutathione S-transferase (GST) is an enzyme superfamily of particular interest for human health with many functional roles, and it is involved in several cancer types. Under cell stress conditions, their concentrations can increase by up to 10% of cell protein content. Recently, a study describing the landscape of RNA-binding proteins in mammalian spermatogenesis reported evidence of canonical GST-RNA interactions in three mouse mu GSTs. Prompted by this, we searched for available databases and found that RBP2GO, which collects candidate RNA-binding proteins (RBPs) detected in recent human proteomic studies, also lists a few human GST-RNA interactions, without any molecular details. To highlight the molecular features of the GST-RNA interaction, we applied recently developed predictors of RNA-binding sites and validated the results with AutoDock Vina, a docking program that computes binding affinity. Overall, our findings support the notion that a GST-RNA interaction can exist and suggest a potential overlap between RNA binding sites and residues responsible for binding glutathione (GSH), which is the most common GST substrate. Our computational analysis supports the notion that human GSTs can bind RNA that shares the binding region with the glutathione-binding pocket.
Background:The Critical Assessment of Functional Annotation (CAFA) is a community effort held to understand the field of computational protein function prediction. Every three years, since 2010, the organizers initiate an experiment to collect function predictions on a large set of proteins and then evaluate the performance of predicting methods on a subset of proteins that have accumulated experimental annotations between the submission deadline and the evaluation time. CAFA provides an independent and rigorous assessment of the current state of the art, thus leveling the playing field, highlighting successes, revealing bottlenecks, and offering a forum for the exchange of ideas in protein science. Here, we report the results of the fourth CAFA experiment (CAFA4). Results:CAFA4 featured the participation of 148 methods from 70 research groups on a total of 46,205 unique proteins over a 5-year annotation accumulation phase, the longest in any CAFA. In a comparison across CAFA2-CAFA4 methods, the prediction of Gene Ontology (GO) terms has clearly improved across all three GO aspects and traditional evaluation settings. While not achieving the first rank, several CAFA2 and CAFA3 methods featured in the top ten methods in many evaluations, suggesting that earlier methods still hold relevance. The performance is weaker in the newly introduced "partial knowledge" evaluation category (proteins with experimental annotations before submission deadline that gained additional annotations in the same GO aspect during the annotation accumulation phase), highlighting the need for a new class of methods. The rankings of the methods were stable over the years in traditional evaluation settings, but less so in the new partial knowledge evaluation. Overall, the field continues to progress with some influx of new participants. Sustained efforts will be necessary to substantially advance it.
In the context of the Critical Assessment of the Genome Interpretation, 6th edition (CAGI6), the Genetics of Neurodevelopmental Disorders Lab in Padua proposed a new ID-challenge to give the opportunity of developing computational methods for predicting patient's phenotype and the causal variants. Eight research teams and 30 models had access to the phenotype details and real genetic data, based on the sequences of 74 genes (VCF format) in 415 pediatric patients affected by Neurodevelopmental Disorders (NDDs). NDDs are clinically and genetically heterogeneous conditions, with onset in infant age. In this study we evaluate the ability and accuracy of computational methods to predict comorbid phenotypes based on clinical features described in each patient and causal variants. Finally, we asked to develop a method to find new possible genetic causes for patients without a genetic diagnosis. As already done for the CAGI5, seven clinical features (ID, ASD, ataxia, epilepsy, microcephaly, macrocephaly, hypotonia), and variants (causative, putative pathogenic and contributing factors) were provided. Considering the overall clinical manifestation of our cohort, we give out the variant data and phenotypic traits of the 150 patients from CAGI5 ID-Challenge as training and validation for the prediction methods development.
The human reference proteome is routinely modelled with predictive tools such as AlphaFold2. We recently released a database in which, for each human protein, the AlphaFold2 model is paired with its ESMFold counterpart. The two predictive methods take advantage of different procedures and it is interesting to compare them in relation to their quality, particularly when an experimental protein structure is not available. Here, we select three state-of-the-art quality assessment methods and we adopt them to compare 42,942 pairs of models. This procedure helps to find the most reliable models for human proteins, particularly for the set of proteins for which structure prediction methods give dissimilar results. We obtain that when predicted structures are similar, AlphaFold2 models consistently receive higher scores than the ESMFold counterparts. When predicted structures differ, the ESMFold model is the best choice for 49 % of the proteins according to a consensus of the three QA tools.
New thermodynamic and functional studies have been recently conducted to evaluate the impact of amino acid substitutions on the Mitogen Activated Protein Kinases 1 and 3 (MAPK1/3). The Critical Assessment of Genome Interpretation (CAGI) data provider, at Sapienza University of Rome, measured the unfolding free energy and the enzymatic activity of a set of variants (MAPK challenge dataset). Thermodynamic measurements for the denaturant-induced equilibrium unfolding of the phosphorylated and unphosphorylated forms of the MAPKs were obtained by monitoring the far-UV circular dichroism and intrinsic fluorescence changes as a function of denaturant concentration. These values have been used to calculate the change in unfolding free energy between the variant and wild-type proteins at zero concentration of denaturant ( ΔΔ G^H_2O ). The enzymatic activity of the phosphorylated MAPKs variants was also measured using Chelation-Enhanced Fluorescence to monitor the phosphorylation of a peptide substrate. The MAPK challenge dataset, composed of a total of 23 single amino acid substitutions (11 and 12 for MAPK1 and MAPK3, respectively), was used to assess the effectiveness of the computational methods in predicting the ΔΔ G^H_2O values, associated with the variants, and categorize them as destabilizing and not destabilizing. The data on the enzymatic activity of the MAPKs mutants were used to assess the performance of the methods for predicting the functional impact of the variants. For the sixth edition of CAGI, thirteen independent research groups from four continents (Asia, Australia, Europe and North America) submitted > 80 sets of predictions, obtained from different approaches. In this manuscript, we summarized the results of our assessment to highlight the possible limitations of the available algorithms.
Metabolomics opens novel avenues to study the basic biological mechanisms underlying complex traits, starting from characterization of metabolites. Metabolites and their levels in a biofluid represent simple molecular phenotypes (metabotypes) that are direct products of enzyme activities and relate to all metabolic pathways, including catabolism and anabolism of nutrients. In this study, we demonstrated the utility of merging metabolomics and genomics in pigs to uncover a large list of genetic factors that influence mammalian metabolism. We obtained targeted characterization of the plasma metabolome of more than 1300 pigs from two populations of Large White and Duroc pig breeds. The metabolomic profiles of these pigs were used to identify genetically influenced metabolites by estimating the heritability of the level of 188 metabolites. Then, combining breed-specific genome-wide association studies of single metabolites and their ratios and across breed meta-analyses, we identified a total of 97 metabolite quantitative trait loci (mQTL), associated with 126 metabolites. Using these results, we constructed a human-pig comparative catalog of genetic factors influencing the metabolomic profile. Whole genome resequencing data identified several putative causative mutations for these mQTL. Additionally, based on a major mQTL for kynurenine level, we designed a nutrigenetic study feeding piglets that carried different genotypes at the candidate gene kynurenine 3-monooxygenase (KMO) varying levels of tryptophan and demonstrated the effect of this genetic factor on the kynurenine pathway. Furthermore, we used metabolomic profiles of Large White and Duroc pigs to reconstruct metabolic pathways using Gaussian Graphical Models, which included perturbation of the identified mQTL. This study has provided the first catalog of genetic factors affecting molecular phenotypes that describe the pig blood metabolome, with links to important metabolic pathways, opening novel avenues to merge genetics and nutrition in this livestock species. The obtained results are relevant for basic and applied biology and to evaluate the pig as a biomedical model. Genetically influenced metabolites can be further exploited in nutrigenetic approaches in pigs. The described molecular phenotypes can be useful to dissect complex traits and design novel feeding, breeding and selection programs in pigs.
Identifying protective antigens (PAs), i.e., targets for bacterial vaccines, is challenging as conducting in-vivo tests at the proteome scale is impractical. Reverse Vaccinology (RV) aids in narrowing down the pool of candidates through computational screening of proteomes. Within RV, one prominent approach is to train Machine Learning (ML) models to classify PAs. These models can be used to predict unseen protein sequences and assist researchers in selecting promising candidates. Traditionally, proteins are fed into these models as vectors of biological and physico-chemical descriptors derived from their residue sequences. However, this method relies on multiple third-party software packages, which may be unreliable, difficult to use, or no longer maintained. Furthermore, selecting descriptors is susceptible to biases. Hence, Protein Sequence Embeddings (PSEs)-high-dimensional vectorial representations of protein sequences obtained from pretrained deep neural networks-have emerged as an alternative to descriptors, offering data-driven feature extraction and a streamlined computational pipeline. We introduce PSEs as a descriptor-free representation of protein sequences for ML in RV. We conducted a thorough comparison of PSE-based and descriptor-based pipelines for PA classification across 10 bacterial species evaluated independently. Our results show that the PSE-based pipeline, which leverages the FAIR ESM-2 protein language model, outperformed the descriptor-based pipeline in 9 out of 10 species, with a mean Area Under the Receiver Operating Characteristics curve (AUROC) of 0.875 versus 0.855. Additionally, it achieved superior performance on the iBPA benchmark (0.86 AUROC vs. 0.82) compared to other methods in the literature. Lastly, we applied the pipeline to rank unseen proteomes based on protective potential to guide candidate selection for pre-clinical testing. Compared to the standard RV practice of ranking candidates according to their biological descriptors, our approach reduces the number of pre-clinical tests needed to identify PAs by up to 83% on average.
Regular, systematic, and independent assessments of computational tools that are used to predict the pathogenicity of missense variants are necessary to evaluate their clinical and research utility and guide future improvements. The Critical Assessment of Genome Interpretation (CAGI) conducts the ongoing Annotate-All-Missense (Missense Marathon) challenge, in which missense variant effect predictors (also called variant impact predictors) are evaluated on missense variants added to disease-relevant databases following the prediction submission deadline. Here we assess predictors submitted to the CAGI 6 Annotate-All-Missense challenge, predictors commonly used in clinical genetics, and recently developed deep learning methods. We examine performance across a range of settings relevant for clinical and research applications, focusing on different subsets of the evaluation data as well as high-specificity and high-sensitivity regimes. Our evaluations reveal notable advances in current methods relative to older, well-cited tools in the field. While meta-predictors tend to outperform their constituent individual predictors, several newer individual predictors perform comparably to commonly used meta-predictors. Predictor performance varies between high-specificity and high-sensitivity regimes, highlighting that different methods may be optimal for different use cases. We also characterize two potential sources of bias. Predictors that incorporate allele frequency as a predictive feature tend to have reduced performance when distinguishing pathogenic variants from very rare benign variants, and predictors trained on pathogenicity labels from curated variant databases often inherit gene-level label imbalances. Our findings help illuminate the clinical and research utility of modern missense variant effect predictors and identify potential areas for future development.
AlphaFold2 predicts protein structures from structural and functional knowledge. Alternatively, ESMFold does the same adopting protein language models. Here, we map available Pfam domains on pairs of models of the human reference proteome computed with both procedures and we compare the mapped regions relevant for functional annotation. We find that, rather irrespectively of the global superimposition of the pairwise models, Pfam-containing regions overlap with a TM-score above 0.8 and a predicted local distance difference test (pLDDT) which is higher than the rest of the modeled sequence. This indicates that both methods are similarly performing in modeled regions that overlap Pfam domains, carrying structural and functional information, with pLDDT values slightly higher for AlphaFold2. The mapping of 9,834 Pfam domains also allows the location of 2,578 active sites in 3,382 enzymes of the human proteome, including 807 proteins for which the active site is not reported in UniProt.
Critical evaluation of computational tools for predicting variant effects is important considering their increased use in disease diagnosis and driving molecular discoveries. In the sixth edition of the Critical Assessment of Genome Interpretation (CAGI) challenge, a dataset of 28 STK11 rare variants (27 missense, 1 single amino acid deletion), identified in primary non-small cell lung cancer biopsies, was experimentally assayed to characterize computational methods from four participating teams and five publicly available tools. Predictors demonstrated a high level of performance on key evaluation metrics, measuring correlation with the assay outputs and separating loss-of-function (LoF) variants from wildtype-like (WT-like) variants. The best participant model, 3Cnet, performed competitively with well-known tools. Unique to this challenge was that the functional data was generated with both biological and technical replicates, thus allowing the assessors to realistically establish maximum predictive performance based on experimental variability. Three out of the five publicly available tools and 3Cnet approached the performance of the assay replicates in separating LoF variants from WT-like variants. Surprisingly, REVEL, an often-used model, achieved a comparable correlation with the real-valued assay output as that seen for the experimental replicates. Performing variant interpretation by combining the new functional evidence with computational and population data evidence led to 16 new variants receiving a clinically actionable classification of likely pathogenic (LP) or likely benign (LB). Overall, the STK11 challenge highlights the utility of variant effect predictors in biomedical sciences and provides encouraging results for driving research in the field of computational genome interpretation.
The pathogenicity of human variants is an important annotation feature that may help in understanding, at a molecular level, the propensity for a human being to develop a certain disease or pathology. Recently, protein sequence embedding associated with machine and/or deep learning has been proven useful in improving results in this area. Different aspects of pathogenic variants can help in understanding the molecular mechanisms of the disease at a molecular level. These include solvent accessibility in the folded gene, the effect on the protein stability, and eventually the perturbation on interaction networks important for biological processes. Here, we describe how, once a variant is predicted "pathogenic", other important structural and functional properties can be derived computationally at the same website ( https://bioinformaticsweeties.biocomp.unibo.it/ ), including the protein structure, if not available. All the properties can help to understand variant effects within the complex context of the cell environment.
Continued advances in variant effect prediction are necessary to demonstrate the ability of machine learning methods to accurately determine the clinical impact of variants of unknown significance (VUS). Towards this goal, the ARSA Critical Assessment of Genome Interpretation (CAGI) challenge was designed to characterize progress by utilizing 219 experimentally assayed missense VUS in the Arylsulfatase A (ARSA) gene to assess the performance of community-submitted predictions of variant functional effects. The challenge involved 15 teams, and evaluated additional predictions from established and recently released models. Notably, a model developed by participants of a genetics and coding bootcamp, trained with standard machine-learning tools in Python, demonstrated superior performance among submissions. Furthermore, the study observed that state-of-the-art deep learning methods provided small but statistically significant improvement in predictive performance compared to less elaborate techniques. These findings underscore the utility of variant effect prediction, and the potential for models trained with modest resources to accurately classify VUS in genetic and clinical research.
Recent thermodynamic and functional studies have been conducted to evaluate the impact of amino acid substitutions on Calmodulin (CaM). The Critical Assessment of Genome Interpretation (CAGI) data provider at University of Verona (Italy) measured the melting temperature (Tm) and the percentage of unfolding (%unfold) of a set of CaM variants (CaM challenge dataset). Thermodynamic measurements for the equilibrium unfolding of CaM were obtained by monitoring far-UV Circular Dichroism as a function of temperature. These measurements were used to determine the Tm and the percentage of protein remaining unfolded at the highest temperature. The CaM challenge dataset, comprising a total of 15 single amino acid substitutions, was used to evaluate the effectiveness of computational methods in predicting the Tm and unfolding percentages associated with the variants, and categorizing them as destabilizing or not. For the sixth edition of CAGI, nine independent research groups from four continents (Asia, Australia, Europe, and North America) submitted over 52 sets of predictions, derived from various approaches. In this manuscript, we summarize the results of our assessment to highlight the potential limitations of current algorithms and provide insights into the future development of more accurate prediction tools. By evaluating the thermodynamic stability of CaM variants, this study aims to enhance our understanding of the relationship between amino acid substitutions and protein stability, ultimately contributing to more accurate predictions of the effects of genetic variants.
This paper reports the evaluation of predictions for the "CALM1" challenge in the fifth round of the Critical Assessment of Genome Interpretation held in 2018. In the challenge, the participants were asked to predict effects on yeast growth caused by missense variants of human calmodulin, a highly conserved protein in eukaryotic cells sensing calcium concentration. The performance of predictors implementing different algorithms and methods is similar. Most predictors are able to identify the deleterious or tolerated variants with modest accuracy, with a baseline predictor based purely on sequence conservation slightly outperforming the submitted predictions. Nevertheless, we think that the accuracy of predictions remains far from satisfactory, and the field awaits substantial improvements. The most poorly predicted variants in this round surround functional CALM1 sites that bind calcium or peptide, which suggests that better incorporation of structural analysis may help improve predictions.
Piero Fariselli合作论文数University of Bologna167
Emidio Capriotti合作论文数42