The SARS-CoV-2 Omicron variants show different behavior compared to the previous variants, especially with respect to the Delta variant, which promotes a lower morbidity despite being much more contagious. In this perspective, we performed molecular dynamics (MD) simulations of the different spike RBD/hACE2 complexes corresponding to the WT, Delta and four Omicron variants. Carrying out a comprehensive analysis of residue interactions within and between the two partners allowed us to draw the profile of each variant by using complementary methods (PairInt, hydrophobic potential, contact PCA). PairInt calculations highlighted the residues most involved in electrostatic interactions, which make a strong contribution to the binding with highly stable interactions between spike RBD and hACE2. Apolar contacts made a substantial and complementary contribution in Omicron with the detection of two hydrophobic patches. Contact networks and cross-correlation matrices were able to detect subtle changes at point mutations as the S375F mutation occurring in all Omicron variants, which is likely to confer an advantage in binding stability. This study brings new highlights on the dynamic binding of spike RBD to hACE2, which may explain the final persistence of Omicron over Delta.
Motivation:The human leukocyte antigen (HLA) system is the main cause of organ transplant loss through the recognition of HLAs present on the graft by donor-specific antibodies raised by the recipient. It is therefore of key importance to identify all potentially immunogenic B-cell epitopes on HLAs in order to refine organ allocation. Such HLAs epitopes are currently characterized by the presence of polymorphic residues called "eplets". However, many polymorphic positions in HLAs sequences are not yet experimentally confirmed as eplets associated with a HLA epitope. Moreover, structural studies of these epitopes only consider 3D static structures. Results:We present here a machine-learning approach for predicting HLA epitopes, based on 3D-surface patches and molecular dynamics simulations. A collection of 3D-surface patches labeled as Epitope (2117) or Nonepitope (4769) according to Human Leukocyte Antigen Eplet Registry information was derived from 207 HLAs (61 solved and 146 predicted structures). Descriptors derived from static and dynamic patch properties were computed and three tree-based models were trained on a reduced non-redundant dataset. HLA-Epicheck is the prediction system formed by the three models. It leverages dynamic descriptors of 3D-surface patches for more than half of its prediction performance. Epitope predictions on unconfirmed eplets (absent from the initial dataset) are compared with experimental results and notable consistency is found. Availability and implementation:Structural data and MD trajectories are deposited as open data under doi: 10.57745/GXZHH8. In-house scripts and machine-learning models for HLA-EpiCheck are available from https://gitlab.inria.fr/capsid.public_codes/hla-epicheck.
Abstract Motivation Protein domains can be viewed as building blocks, essential for understanding structure–function relationships in proteins. However, each domain database classifies protein domains using its own methodology. Thus, in many cases, domain models and boundaries differ from one domain database to the other, raising the question of domain definition and enumeration of true domain instances. Results We propose an automated iterative workflow to assess protein domain classification by cross-mapping domain structural instances between domain databases and by evaluating structural alignments. CroMaSt (for Cross-Mapper of domain Structural instances) will classify all experimental structural instances of a given domain type into four different categories (‘Core’, ‘True’, ‘Domain-like’ and ‘Failed’). CroMast is developed in Common Workflow Language and takes advantage of two well-known domain databases with wide coverage: Pfam and CATH. It uses the Kpax structural alignment tool with expert-adjusted parameters. CroMaSt was tested with the RNA Recognition Motif domain type and identifies 962 ‘True’ and 541 ‘Domain-like’ structural instances for this domain type. This method solves a crucial issue in domain-centric research and can generate essential information that could be used for synthetic biology and machine-learning approaches of protein domain engineering. Availability and implementation The workflow and the Results archive for the CroMaSt runs presented in this article are available from WorkflowHub (doi: 10.48546/workflowhub.workflow.390.2). Supplementary information Supplementary data are available at Bioinformatics Advances online.
The search for an effective drug is still urgent for COVID-19 as no drug with proven clinical efficacy is available. Finding the new purpose of an approved or investigational drug, known as drug repurposing, has become increasingly popular in recent years. We propose here a new drug repurposing approach for COVID-19, based on knowledge graph (KG) embeddings. Our approach learns “ensemble embeddings” of entities and relations in a COVID-19 centric KG, in order to get a better latent representation of the graph elements. Ensemble KG-embeddings are subsequently used in a deep neural network trained for discovering potential drugs for COVID-19. Compared to related works, we retrieve more in-trial drugs among our top-ranked predictions, thus giving greater confidence in our prediction for out-of-trial drugs. For the first time to our knowledge, molecular docking is then used to evaluate the predictions obtained from drug repurposing using KG embedding. We show that Fosinopril is a potential ligand for the SARS-CoV-2 nsp13 target. We also provide explanations of our predictions thanks to rules extracted from the KG and instanciated by KG-derived explanatory paths. Molecular evaluation and explanatory paths bring reliability to our results and constitute new complementary and reusable methods for assessing KG-based drug repurposing.
1 Abstract The Human Leukocyte Antigen (HLA) system is the main cause of organ transplant loss through the recognition of HLA proteins by Donor-Specific Antibodies (DSA). Therefore, the identification of potentially immunogenic epitopes is a key task to refine organ allocation and then to improve the survival of transplanted organs. Here, we present HLA-EpiCheck, a machine learning predictor for B-cell epitopes on HLA proteins that leverages an unprecedented dataset of high-quality molecular dynamics simulations of 207 HLA proteins. Candidate epitopes are represented as surface patches centered on solvent-accessible residues and described by a set of 18 descriptors. The descriptors include both static and dynamic properties, such as hydrophobicity, electrostatic charges, relative solvent-accessible surface area and side-chain flexibility. The HLA-EpiCheck was trained using an Extra Trees ensemble learning method and was compared to DiscoTope-3.0, a state-of-the-art B-cell epitope predictor. HLA-EpiCheck largely outperformed DiscoTope-3.0 in the task of predicting HLA epitopes. HLA-EpiCheck was also used to assess the epitope status of a subset of non-confirmed eplets. The predictions were compared to experimental data and a notable consistency was found. These results suggest that HLA-EpiCheck could be used to better define HLA matching between donor and recipient to reduce de novo DSA formation and graft rejection.
The current rise of Open Science and Reproducibility in the Life Sciences requires the creation of rich, machine-actionable metadata in order to better share and reuse biological digital resources such as datasets, bioinformatics tools, training materials, etc. For this purpose, FAIR principles have been defined for both data and metadata and adopted by large communities, leading to the definition of specific metrics. However, automatic FAIRness assessment is still difficult because computational evaluations frequently require technical expertise and can be time-consuming. As a first step to address these issues, we propose FAIR-Checker, a web-based tool to assess the FAIRness of metadata presented by digital resources. FAIR-Checker offers two main facets: a "Check" module providing a thorough metadata evaluation and recommendations, and an "Inspect" module which assists users in improving metadata quality and therefore the FAIRness of their resource. FAIR-Checker leverages Semantic Web standards and technologies such as SPARQL queries and SHACL constraints to automatically assess FAIR metrics. Users are notified of missing, necessary, or recommended metadata for various resource categories. We evaluate FAIR-Checker in the context of improving the FAIRification of individual resources, through better metadata, as well as analyzing the FAIRness of more than 25 thousand bioinformatics software descriptions.
Background: Data management is fast becoming an essential part of scientific practice, driven by open science and FAIR (findable, accessible, interoperable, and reusable) data sharing requirements. Whilst data management plans (DMPs) are clear to data management experts and data stewards, understandings of their purpose and creation are often obscure to the producers of the data, which in academic environments are often PhD students. Methods: Within the RNAct EU Horizon 2020 ITN project, we engaged the 10 RNAct early-stage researchers (ESRs) in a training project aimed at formulating a DMP. To do so, we used the Data Stewardship Wizard (DSW) framework and modified the existing Life Sciences Knowledge Model into a simplified version aimed at training young scientists, with computational or experimental backgrounds, in core data management principles. We collected feedback from the ESRs during this exercise. Results: Here, we introduce our new life-sciences training DMP template for young scientists. We report and discuss our experiences as principal investigators (PIs) and ESRs during this project and address the typical difficulties that are encountered in developing and understanding a DMP. Conclusions: We found that the DS-wizard can also be an appropriate tool for DMP training, to get terminology and concepts across to researchers. A full training in addition requires an upstream step to present basic DMP concepts and a downstream step to publish a dataset in a (public) repository. Overall, the DS-Wizard tool was essential for our DMP training and we hope our efforts can be used in other projects.
Machine learning is now an essential part of any biomedical study but its integration into real effective Learning Health Systems, including the whole process of Knowledge Discovery from Data (KDD), is not yet realised. We propose an original extension of the KDD process model that involves an inductive database. We designed for the first time a generic model of Inductive Clinical DataBase (ICDB) aimed at hosting both patient data and learned models. We report experiments conducted on patient data in the frame of a project dedicated to fight heart failure. The results show how the ICDB approach allows to identify biomarker combinations, specific and predictive of heart fibrosis phenotype, that put forward hypotheses relative to underlying mechanisms. Two main scenarios were considered, a local-to-global KDD scenario and a trans-cohort alignment scenario. This promising proof of concept enables us to draw the contours of a next-generation Knowledge Discovery Environment (KDE).
Aims:End-stage renal disease (ESRD) treated by chronic hemodialysis (HD) is associated with poor cardiovascular (CV) outcomes, with no available evidence-based therapeutics. A multiplexed proteomic approach may identify new pathophysiological pathways associated with CV outcomes, potentially actionable for precision medicine.Methods and results:The AURORA trial was an international, multicentre, randomized, double-blind trial involving 2776 patients undergoing maintenance HD. Rosuvastatin vs. placebo had no significant effect on the composite primary endpoint of death from CV causes, nonfatal myocardial infarction or nonfatal stroke. We first compared CV risk-matched cases and controls (n = 410) to identify novel biomarkers using a multiplex proximity extension immunoassay (276 proteomic biomarkers assessed with OlinkTM). We replicated our findings in 200 unmatched cases and 200 controls. External validation was conducted from a multicentre real-life Danish cohort [Aarhus-Aalborg (AA), n = 331 patients] in which 92 OlinkTM biomarkers were assessed. In AURORA, only N-terminal pro-brain natriuretic peptide (NT-proBNP, positive association) and stem cell factor (SCF) (negative association) were found consistently associated with the trial's primary outcome across exploration and replication phases, independently from the baseline characteristics. Stem cell factor displayed a lower added predictive ability compared with NT-ProBNP. In the AA cohort, in multivariable analyses, BNP was found significantly associated with major CV events, while higher SCF was associated with less frequent CV deaths.Conclusions:Our findings suggest that NT-proBNP and SCF may help identify ESRD patients with respectively high and low CV risk, beyond classical clinical predictors and also point at novel pathways for prevention and treatment.
Automatic protein function annotation is a challenging task that is fundamental in many medical applications. Indeed, the capability to predict whether a protein has a given function is a key step for disease understanding and drug design. For such reasons, many authors have proposed computational methods for protein function prediction. One key element that is present in many proposals is similarity functions. Such functions are often used to compute the pairwise similarity between two proteins. It is commonly accepted that proteins with similar structures share the same function. Nevertheless, no previous works have focused on proposing a similarity function that is specifically designed for protein function annotation. In this work, we analyze the best similarity functions for the protein function annotation task and propose a new one. We performed experiments in a simple pairwise similarity scenario and also using our proposal as part of a more complex protein function annotation method. Based on the results, we can state that our proposal is a valid alternative as a building block of many protein function annotation methods.
Emerging SARS-CoV-2 variants raise concerns about our ability to withstand the Covid-19 pandemic, and therefore, understanding mechanistic differences of those variants is crucial. In this study, we investigate disparities between the SARS-CoV-2 wild type and five variants that emerged in late 2020, focusing on the structure and dynamics of the spike protein interface with the human angiotensin-converting enzyme 2 (ACE2) receptor, by using crystallographic structures and extended analysis of microsecond molecular dynamics simulations. Dihedral angle principal component analysis (PCA) showed the strong similarities in the spike receptor binding domain (RBD) dynamics of the Alpha, Beta, Gamma, and Delta variants, in contrast with those of WT and Epsilon. Dynamical perturbation networks and contact PCA identified the peculiar interface dynamics of the Delta variant, which cannot be directly imputable to its specific L452R and T478K mutations since those residues are not in direct contact with the human ACE2 receptor. Our outcome shows that in the Delta variant the L452R and T478K mutations act synergistically on neighboring residues to provoke drastic changes in the spike/ACE2 interface; thus a singular mechanism of action eventually explains why it dominated over preceding variants.
The protein function annotation based on functional properties like the Enzyme Commission (EC) numbers is a very challenging task that aims to understand life at the molecular level. Especially, the size of features for each protein is very huge and the number of labeled samples is limited, which can significantly affect the annotation accuracy. To address these issues, we propose a novel semi-supervised graph deep learning model that aims to learn better latent representations for each protein/node by taking into account the neighborhood information in order to improve the annotation. Firstly, we extract a set of features from raw protein data. Each protein is associated with a 1-D feature vector that represents its InterPro domain composition. As D, the number of possible interPro domains, is very high (>11,000), we design a deep autoencoder model (DAE) that seeks to find an efficient representation of the domain composition of proteins in a lower dimensional latent space. Then, we construct a protein graph where each node is a protein associated with its latent representation vector and each edge is weighted by the Euclidean distance between the two nodes it connects. Finally, we train a semi-supervised graph neural network (SGNN) for the automatic protein function annotation using the constructed protein graph. Experiments are conducted on four reference proteomes in UniProtKB/SwissProt, including Human, Arabidopsis Thaliana, Mouse, and Rat. Experimental results show that the proposed model is competitive for protein function annotation compared to existing methods.
BackgroundAutomatic functional annotation of proteins is an open research problem in bioinformatics. The growing number of protein entries in public databases, for example in UniProtKB, poses challenges in manual functional annotation. Manual annotation requires expert human curators to search and read related research articles, interpret the results, and assign the annotations to the proteins. Thus, it is a time-consuming and expensive process. Therefore, designing computational tools to perform automatic annotation leveraging the high quality manual annotations that already exist in UniProtKB/SwissProt is an important research problemResultsIn this paper, we extend and adapt the GrAPFI (graph-based automatic protein function inference) (Sarker et al. in BMC Bioinform 21, 2020; Sarker et al., in: Proceedings of 7th international conference on complex networks and their applications, Cambridge, 2018) method for automatic annotation of proteins with gene ontology (GO) terms renaming it as GrAPFI-GO. The original GrAPFI method uses label propagation in a similarity graph where proteins are linked through the domains, families, and superfamilies that they share. Here, we also explore various types of similarity measures based on common neighbors in the graph. Moreover, GO terms are arranged in a hierarchical manner according to semantic parent-child relations. Therefore, we propose an efficient pruning and post-processing technique that integrates both semantic similarity and hierarchical relations between the GO terms. We produce experimental results comparing the GrAPFI-GO method with and without considering common neighbors similarity. We also test the performance of GrAPFI-GO and other annotation tools for GO annotation on a benchmark of proteins with and without the proposed pruning and post-processing procedure.ConclusionOur results show that the proposed semantic hierarchical post-processing potentially improves the performance of GrAPFI-GO and of other annotation tools as well. Thus, GrAPFI-GO exposes an original efficient and reusable procedure, to exploit the semantic relations among the GO terms in order to improve the automatic annotation of protein functions
Efficiently discovering causal relations from data and representing them in a way that facilitates their use is an important problem in science that has received much attention. In this paper, we propose an adaptation of the Formal Concept Analysis formalism to the problem of discovering and representing causal relations. We show that Formal Concept Analysis structures and algorithms are well-suited to this problem.
Many biological processes are mediated by protein-protein interactions (PPIs). Because protein domains are the building blocks of proteins, PPIs likely rely on domain-domain interactions (DDIs). Several attempts exist to infer DDIs from PPI networks but the produced datasets are heterogeneous and sometimes not accessible, while the PPI interactome data keeps growing. We describe a new computational approach called “PPIDM” (Protein-Protein Interactions Domain Miner) for inferring DDIs using multiple sources of PPIs. The approach is an extension of our previously described “CODAC” (Computational Discovery of Direct Associations using Common neighbors) method for inferring new edges in a tripartite graph. The PPIDM method has been applied to seven widely used PPI resources, using as “Gold-Standard” a set of DDIs extracted from 3D structural databases. Overall, PPIDM has produced a dataset of 84,552 non-redundant DDIs. Statistical significance (p-value) is calculated for each source of PPI and used to classify the PPIDM DDIs in Gold (9,175 DDIs), Silver (24,934 DDIs) and Bronze (50,443 DDIs) categories. Dataset comparison reveals that PPIDM has inferred from the 2017 releases of PPI sources about 46% of the DDIs present in the 2020 release of the 3did database, not counting the DDIs present in the Gold-Standard. The PPIDM dataset contains 10,229 DDIs that are consistent with more than 13,300 PPIs extracted from the IMEx database, and nearly 23,300 DDIs (27.5%) that are consistent with more than 214,000 human PPIs extracted from the STRING database. Examples of newly inferred DDIs covering more than 10 PPIs in the IMEx database are provided. Further exploitation of the PPIDM DDI reservoir includes the inventory of possible partners of a protein of interest and characterization of protein interactions at the domain level in combination with other methods. The result is publicly available at http://ppidm.loria.fr/ .
In this paper, we are interested in explaining datasets used in supervised classification problems by identifying the features that are most predictive and discriminant of the class. We first propose a definition of predictive and discriminant features based on the impact of the features on the performance of classification models. We then propose an approach to identifying the most predictive and discriminant features using multicriteria decision making. Finally, we present and discuss an experiment on a public dataset illustrating the potential of the approach.
The choice of the most appropriate unsupervised machine-learning method for “heterogeneous” or “mixed” data, i.e. with both continuous and categorical variables, can be challenging. Our aim was to examine the performance of various clustering strategies for mixed data using both simulated and real-life data. We conducted a benchmark analysis of “ready-to-use” tools in R comparing 4 model-based (Kamila algorithm, Latent Class Analysis, Latent Class Model [LCM] and Clustering by Mixture Modeling) and 5 distance/dissimilarity-based (Gower distance or Unsupervised Extra Trees dissimilarity followed by hierarchical clustering or Partitioning Around Medoids, K-prototypes) clustering methods. Clustering performances were assessed by Adjusted Rand Index (ARI) on 1000 generated virtual populations consisting of mixed variables using 7 scenarios with varying population sizes, number of clusters, number of continuous and categorical variables, proportions of relevant (non-noisy) variables and degree of variable relevance (low, mild, high). Clustering methods were then applied on the EPHESUS randomized clinical trial data (a heart failure trial evaluating the effect of eplerenone) allowing to illustrate the differences between different clustering techniques. The simulations revealed the dominance of K-prototypes, Kamila and LCM models over all other methods. Overall, methods using dissimilarity matrices in classical algorithms such as Partitioning Around Medoids and Hierarchical Clustering had a lower ARI compared to model-based methods in all scenarios. When applying clustering methods to a real-life clinical dataset, LCM showed promising results with regard to differences in (1) clinical profiles across clusters, (2) prognostic performance (highest C-index) and (3) identification of patient subgroups with substantial treatment benefit. The present findings suggest key differences in clustering performance between the tested algorithms (limited to tools readily available in R). In most of the tested scenarios, model-based methods (in particular the Kamila and LCM packages) and K-prototypes typically performed best in the setting of heterogeneous data.
Insulin-like Growth Factor Binding Protein 2 (IGFBP2) is a member of the IGFBP family which is present in the heart; cardiac IGFBP2 mRNA levels have furthermore been shown to rise in animal models of ischemia-induced heart failure (HF) [ [1] Berry M. Galinier M. Delmas C. et al. Proteomics analysis reveals IGFBP2 as a candidate diagnostic biomarker for heart failure. IJC Metabolic & Endocrine. 2015; 6: 5-12 Crossref Scopus (13) Google Scholar ]. The group of Dr. P. Rouet previously showed that IGFBP2 is a powerful diagnostic biomarker for HF whose accuracy for identifying acute HF is a valuable adjunct to brain natriuretic peptide (BNP) [ [1] Berry M. Galinier M. Delmas C. et al. Proteomics analysis reveals IGFBP2 as a candidate diagnostic biomarker for heart failure. IJC Metabolic & Endocrine. 2015; 6: 5-12 Crossref Scopus (13) Google Scholar ]: The area under the curve of the receiver operating characteristic (ROC) curve for HF was 0.908 (0.830–0.958) for IGFBP2 versus 0.857 (0.770–0.921) for BNP. In addition, IGFBP2 improved the model's accuracy in predicting HF as well as the net reclassification index (NRI) (0.574 (0.340–0.750) p < .001) on top of BNP. Insulin-like Growth Factor Binding Protein 2 predicts mortality risk in heart failureInternational Journal of CardiologyVol. 300PreviewInsulin-like Growth Factor Binding Protein 2 (IGFBP2) showed greater heart failure (HF) diagnostic accuracy than the "grey zone" B-type natriuretic peptides, and may have prognostic utility as well. Full-Text PDF
Background Many patients with heart failure with preserved ejection fraction (HFpEF) are women. Exploring mechanisms underlying the sex differences may improve our understanding of the pathophysiology of HFpEF. Studies focusing on sex differences in circulating proteins in HFpEF patients are scarce. Methods A total of 415 proteins were analyzed in 392 HFpEF patients included in The Metabolic Road to Diastolic Heart Failure: Diastolic Heart Failure study (MEDIA-DHF). Sex differences in these proteins were assessed using adjusted logistic regression analyses. The associations between candidate proteins and cardiovascular (CV) death or CV hospitalization (with sex interaction) were assessed using Cox regression models. Results We found 9 proteins to be differentially expressed between female and male patients. Women expressed more LPL and PLIN1, which are markers of lipid metabolism; more LHB, IGFBP3, and IL1RL2 as markers of transcriptional regulation; and more Ep-CAM as marker of hemostasis. Women expressed less MMP-3, which is a marker associated with extracellular matrix organization; less NRP1, which is associated with developmental processes; and less ACE2, which is related to metabolism. Sex was not associated with the study outcomes (adj. HR 1.48, 95% CI 0.83–2.63), p = 0.18. Conclusion In chronic HFpEF, assessing sex differences in a wide range of circulating proteins led to the identification of 9 proteins that were differentially expressed between female and male patients. These findings may help further investigations into potential pathophysiological processes contributing to HFpEF.
Malika Smaïl-Tabbone合作论文数Batiment B - Equipe Orapilleur;Campus Scientifique91
Yannick Toussaint合作论文数INRIA Nancy Grand-Est and LORIA, Orpailleur Team6