Mycobacterium tuberculosis (Mtb) is the causative agent of tuberculosis (TB), an infectious disease that is a major killer worldwide. Due to selection pressure caused by the use of antibacterial drugs, Mtb is characterised by mutational events that have given rise to multi drug resistant (MDR) and extensively drug resistant (XDR) phenotypes. The rate at which mutations occur is an important factor in the study of molecular evolution, and it helps understand gene evolution. Within the same species, different protein-coding genes evolve at different rates. To estimate the rates of molecular evolution of protein-coding genes, a commonly used parameter is the ratio d N/ d S, where d N is the rate of non-synonymous substitutions and d S is the rate of synonymous substitutions. Here, we determined the estimated rates of molecular evolution of select biological processes and molecular functions across 264 strains of Mtb. We also investigated the molecular evolutionary rates of core genes of Mtb by computing the d N/ d S values, and estimated the pan genome of the 264 strains of Mtb. Our results show that the cellular amino acid metabolic process and the kinase activity function evolve at a significantly higher rate, while the carbohydrate metabolic process evolves at a significantly lower rate for M. tuberculosi s. These high rates of evolution correlate well with Mtb physiology and pathogenicity. We further propose that the core genome of M. tuberculosis likely experiences varying rates of molecular evolution which may drive an interplay between core genome and accessory genome during M. tuberculosis evolution.
Mycobacterium tuberculosis (Mtb) is the causative agent of tuberculosis (TB), an infectious disease that is a major killer worldwide. Due to selection pressure caused by the use of antibacterial drugs, Mtb is characterised by mutational events that have given rise to multi drug resistant (MDR) and extensively drug resistant (XDR) phenotypes. The rate at which mutations occur is an important factor in the study of molecular evolution, and it helps understand gene evolution. Within the same species, different protein-coding genes evolve at different rates. To estimate the rates of molecular evolution of protein-coding genes, a commonly used parameter is the ratio d N/ d S, where d N is the rate of non-synonymous substitutions and d S is the rate of synonymous substitutions. Here, we determined the estimated rates of molecular evolution of select biological processes and molecular functions across 264 strains of Mtb. We also investigated the molecular evolutionary rates of core genes of Mtb by computing the d N/ d S values, and estimated the pan genome of the 264 strains of Mtb. Our results show that the cellular amino acid metabolic process and the kinase activity function evolve at a significantly higher rate, while the carbohydrate metabolic process evolves at a significantly lower rate for M. tuberculosi s. These high rates of evolution correlate well with Mtb physiology and pathogenicity. We further propose that the core genome of M. tuberculosis likely experiences varying rates of molecular evolution which may drive an interplay between core genome and accessory genome during M. tuberculosis evolution. Keywords Comparative genomics , N/ , S , molecular evolution , ,
In the past decade, the decreased cost of advanced high-throughput technologies has revolutionized biomedical sciences in terms of data volume and diversity. To handle the sheer volumes of sequencing data, quantitative techniques such as machine learning have been employed to handle and find meaning in these data. The need for the integration of complex and multidimensional datasets poses one of the grand challenges of modern bioinformatics. Integrating data from various sources to create larger datasets can allow for greater knowledge transfer and reuse following publication, whether data are submitted to a public repository or shared directly. Standardized procedures, data formats, and comprehensive quality management considerations are the cornerstones of data integration. Combining data from multiple sources can expand the knowledge of a subject. This chapter discusses the importance of incorporating data standardization and good data governance practices in the biomedical sciences. The chapter also describes existing standardization resources and efforts, as well as the challenges related to these practices, emphasizing the critical role of standardization in the omics era. The discussion has been supplemented with practical examples from different “omics” fields.
Recent technological advances have allowed the unprecedented generation of large data sets in the biological sciences. Gaining the most value from this generation requires the data to be distributed and shared more widely so that multiple groups may make use of it. This brings about a number of technical and social challenges, and different approaches have been developed to resolve them. In this chapter, we introduce the concept and principles of data sharing, we discuss two data sharing methods, sharing through an archive and sharing through a data commons, we then provide a case example from the Human Heredity and Health in Africa (H3Africa) consortium, sharing data in resource-limited regions. We also discuss the overall challenges associated with data sharing, as well as Beacons and the associated security and privacy concerns.
Genomics data are currently being produced at unprecedented rates, resulting in increased knowledge discovery and submission to public data repositories. Despite these advances, genomic information on African-ancestry populations remains significantly low compared with European- and Asian-ancestry populations. This information is typically segmented across several different biomedical data repositories, which often lack sufficient fine-grained structure and annotation to account for the diversity of African populations, leading to many challenges related to the retrieval, representation and findability of such information. To overcome these challenges, we developed the African Genomic Medicine Portal (AGMP), a database that contains metadata on genomic medicine studies conducted on African-ancestry populations. The metadata is curated from two public databases related to genomic medicine, PharmGKB and DisGeNET. The metadata retrieved from these source databases were limited to genomic variants that were associated with disease aetiology or treatment in the context of African-ancestry populations. Over 2000 variants relevant to populations of African ancestry were retrieved. Subsequently, domain experts curated and annotated additional information associated with the studies that reported the variants, including geographical origin, ethnolinguistic group, level of association significance and other relevant study information, such as study design and sample size, where available. The AGMP functions as a dedicated resource through which to access African-specific information on genomics as applied to health research, through querying variants, genes, diseases and drugs. The portal and its corresponding technical documentation, implementation code and content are publicly available.
Pathogenic Vibrio spp. are largely responsible for human diseases caused through consumption of contaminated seafood. The aim of this study was to determine the prevalence, population densities, species diversity and molecular characteristics of pathogenic Vibrio in various seafood commodities and its associated health risks. Samples of finfish and shellfish (oysters and sea urchins) were collected from different regions and analyzed for Vibrio using the Most Probable Number (MPN) technique. Genomic DNA of putative Vibrio isolates was analyzed by whole genome sequencing (WGS) for taxonomic identification and identification of genes responsible for virulence and antimicrobial resistance. The risk of vibrio-related illnesses due to the consumption of contaminated seafood was assessed using Risk Ranger. Population densities of presumptive Vibrio fell in the range of 2.6 - 4.4 Log MPN/g and correlated with seasonality, with the summer season favoring significantly (p < 0.05) higher Vibrio counts. A total of 15 Vibrio isolates were identified as V. alginolyticus (5), V . parahaemolyticus (6), V. harveyi (2) or V. diabolicus (2). Two of the six V. parahaemolyticus isolates (ST 2504 and ST 2505) originating from oysters were found to be either tdh + or trh + and thus considered a human pathogen due to elaboration of Thermostable Direct Hemolysin (TDH) or TDH-related hemolysin (TRH). In addition to virulence genes, the shellfish isolates also harbored genes encoding resistance to multiple antibiotics including tetracycline, penicillin, quinolone and beta-lactam antibiotics, thus arousing concern. The risk assessment exercise pointed to an estimated 21 annual cases of V. parahaemolyticus -associated gastroenteritis in the general population attributed to consumption of contaminated oysters. This study highlights not only the wide prevalence and diversity of Vibrio in seafood, but also the potential of certain strains to threaten public health.
Background The wealth of biological information available nowadays in public databases has triggered an unprecedented rise in multi-database search and data retrieval for obtaining detailed information about key functional and structural entities. This concerns investigations ranging from gene or genome analysis to protein structural analysis. However, the retrieval of interconnected data from a number of different databases is very often done repeatedly in an unsystematic way. Results Here, we present TAxonomy, Gene, Ontology, Protein, Structure INtegrated (TAGOPSIN), a command line program written in Java for rapid and systematic retrieval of select data from seven of the most popular public biological databases relevant to comparative genomics and protein structure studies. The program allows a user to retrieve organism-centred data and assemble them in a single data warehouse which constitutes a useful resource for several biological applications. TAGOPSIN was tested with a number of organisms encompassing eukaryotes, prokaryotes and viruses. For example, it successfully integrated data for about 17,000 UniProt entries of Homo sapiens and 21 UniProt entries of human coronavirus. Conclusion TAGOPSIN demonstrates efficient data integration whereby manipulation of interconnected data is more convenient than doing multi-database queries. The program facilitates for instance interspecific comparative analyses of protein-coding genes in a molecular evolutionary study, or identification of taxa-specific protein domains and three-dimensional structures. TAGOPSIN is available as a JAR file at https://github.com/ebundhoo/TAGOPSIN and is released under the GNU General Public License.
A Correction to this paper has been published: https://doi.org/10.1038/s41586-021-03286-9.
As the genomic profile across cancers varies from person to person, patient prognosis and treatment may differ based on the mutational signature of each tumour. Thus, it is critical to understand genomic drivers of cancer and identify potential mutational commonalities across tumors originating at diverse anatomical sites. Large-scale cancer genomics initiatives, such as TCGA, ICGC and GENIE have enabled the analysis of thousands of tumour genomes. Our goal was to identify new cancer-causing mutations that may be common across tumour sites using mutational and gene expression profiles. Genomic and transcriptomic data from breast, ovarian, and prostate cancers were aggregated and analysed using differential gene expression methods to identify the effect of specific mutations on the expression of multiple genes. Mutated genes associated with the most differentially expressed genes were considered to be novel candidates for driver mutations, and were validated through literature mining, pathway analysis and clinical data investigation. Our driver selection method successfully identified 116 probable novel cancer-causing genes, with 4 discovered in patients having no alterations in any known driver genes: MXRA5, OBSCN, RYR1, and TG. The candidate genes previously not officially classified as cancer-causing showed enrichment in cancer pathways and in cancer diseases. They also matched expectations pertaining to properties of cancer genes, for instance, showing larger gene and protein lengths, and having mutation patterns suggesting oncogenic or tumor suppressor properties. Our approach allows for the identification of novel putative driver genes that are common across cancer sites using an unbiased approach without any a priori knowledge on pathways or gene interactions and is therefore an agnostic approach to the identification of putative common driver genes acting at multiple cancer sites.
Atherosclerosis and rheumatoid arthritis are chronic inflammatory diseases of high incidence worldwide. Both diseases are driven by a complex pathogenesis involving molecular interactions between the myriad of genes or proteins which can be modeled as protein-protein interaction networks (PPIN). Here, we identify the common genes and their functional implications in pathways that underlie the common pathophysiology of RA and atherosclerosis using network analysis of PPIN. By integrating protein-protein interaction (PPI) data and gene-disease associations, we constructed a common PPIN to identify the common genes and their topological significance. We then performed network analysis to identify hub genes based on degree and betweenness centrality. Functional modules were identified and mapped to Gene Ontology (GO) annotations. The resulting PPIN consists of 1379 nodes and 1739 interactions with 26 hub genes representing the union of atherosclerosis and RA related genes. Of these 26 genes, PTGS2, VCAM-1, ICAM-1, TNF-α, ALOX5, TGFβ1 and FOS are identified as both hub and common inflammation genes. Five most significant functional modules are detected and GO analysis reveals that these genes are involved in inflammatory response, regulation of T cell proliferation and response to lipopolysaccharide as biological processes. Three significant pathways validated by KEGG are TNF, NFkB and Interleukin-17 signaling pathways. Dysregulation in both lipid metabolism and T cell differentiation and proliferation were found to be the main triggers causing an inflammatory response which was found to be the common shared pathology that links atherosclerosis and RA. These genes could represent potential biomarkers and drug targets for future research on RA and atherosclerosis and their associated comorbidities.
Extensive biological data are currently readily available in public databases, making it possible for a specific research problem to be addressed by multi-database search and data retrieval. However, getting the correct interconnected data from a number of different databases can be cumbersome. Here, we present TAGOPSIN (TAxonomy, Gene, Ontology, Protein, Structure INtegrated), a command line program written in Java whose purpose is to retrieve data from seven public biological repositories and assemble them in a single data warehouse managed by PostgreSQL. Using the object-oriented paradigm, organisms, genomes, genes, proteins, biological functions, protein domain families and protein 3D structures are modelled as real-world interrelated entities. Accordingly, TAGOPSIN retrieves selected data from their respective FTP or HTTP servers and gathers them in a unified local repository. The program was tested with several model prokaryotic organisms. For example, about 1.1 million coding nucleotide sequences and 1,706 PDB entries were retrieved for 264 strains of Mycobacterium tuberculosis . TAGOPSIN constitutes a valuable tool in molecular evolutionary and other comparative genomics studies as well as structure-based studies. TAGOPSIN is released under the GNU General Public License and is available as a JAR file at https://github.com/ebundhoo/TAGOPSIN.
Drug Repositioning is the use of existing drugs to treat new diseases. Drug molecules exert their actions by binding to specific 3D biological molecules. Working and reasoning with 3D structures is complex, thus researchers prefer working with 1D or text data. Furthermore, drug-repositioning studies often use data sets from various independent sources, which make data processing and analysis time consuming due to different file formats, missing data, and complex cross-referencing. Here, we integrate 12 publicly available data sets on various biological/chemical entities like disease, gene, protein, pathway, drug, and side effect and 5 ontologies to provide an abstraction paradigm. The resulting integrated repository, which we called Sirius (for shedding light on drug-disease relationships) contains 7,321 disease related phenotypes, 47,063 protein functions, 2,226 drugs functions and 72,787 drug side effects having 12, 11, 9 and 4 abstraction levels, respectively. We illustrate the usefulness of our repository by studying the relationships between drugs and diseases, using side effect and pathway data. Our study predicted 117 associations, of which 93 are confirmed by the CTD database. The database is available on request from the authors as an SQL dump file.
Comparing and classifying protein domain interactions according to their three-dimensional (3D) structures can help to understand protein structure-function and evolutionary relationships. Additionally, structural knowledge of existing domain-domain interactions can provide a useful way to find structural templates with which to model the 3D structures of unsolved protein complexes. Here we present a straightforward guide to using the "Kbdock" protein domain structure database and its associated web site for exploring and comparing protein domain-domain interactions (DDIs) and domain-peptide interactions (DPIs) at the Pfam domain family level. We also briefly explain how the Kbdock web site works, and we provide some notes and suggestions which should help to avoid some common pitfalls when working with 3D protein domain structures.
ABSTRACT Many of the modeling targets in the blind CASP‐11/CAPRI‐30 experiment were protein homo‐dimers and homo‐tetramers. Here, we perform a retrospective docking‐based analysis of the perfectly symmetrical CAPRI Round 30 targets whose crystal structures have been published. Starting from the CASP “stage‐2” fold prediction models, we show that using our recently developed “SAM” polar Fourier symmetry docking algorithm combined with NAMD energy minimization often gives acceptable or better 3D models of the target complexes. We also use SAM to analyze the overall quality of all CASP structural models for the selected targets from a docking‐based perspective. We demonstrate that docking only CASP “center” structures for the selected targets provides a fruitful and economical docking strategy. Furthermore, our results show that many of the CASP models are dockable in the sense that they can lead to acceptable or better models of symmetrical complexes. Even though SAM is very fast, using docking and NAMD energy minimization to pull out acceptable docking models from a large ensemble of docked CASP models is computationally expensive. Nonetheless, thanks to our SAM docking algorithm, we expect that applying our docking protocol on a modern computer cluster will give us the ability to routinely model 3D structures of symmetrical protein complexes from CASP‐quality models. Proteins 2017; 85:463–469. © 2016 Wiley Periodicals, Inc.
ABSTRACTWe present the results for CAPRI Round 30, the first joint CASP‐CAPRI experiment, which brought together experts from the protein structure prediction and protein–protein docking communities. The Round comprised 25 targets from amongst those submitted for the CASP11 prediction experiment of 2014. The targets included mostly homodimers, a few homotetramers, and two heterodimers, and comprised protein chains that could readily be modeled using templates from the Protein Data Bank. On average 24 CAPRI groups and 7 CASP groups submitted docking predictions for each target, and 12 CAPRI groups per target participated in the CAPRI scoring experiment. In total more than 9500 models were assessed against the 3D structures of the corresponding target complexes. Results show that the prediction of homodimer assemblies by homology modeling techniques and docking calculations is quite successful for targets featuring large enough subunit interfaces to represent stable associations. Targets with ambiguous or inaccurate oligomeric state assignments, often featuring crystal contact‐sized interfaces, represented a confounding factor. For those, a much poorer prediction performance was achieved, while nonetheless often providing helpful clues on the correct oligomeric state of the protein. The prediction performance was very poor for genuine tetrameric targets, where the inaccuracy of the homology‐built subunit models and the smaller pair‐wise interfaces severely limited the ability to derive the correct assembly mode. Our analysis also shows that docking procedures tend to perform better than standard homology modeling techniques and that highly accurate models of the protein components are not always required to identify their association modes with acceptable accuracy. Proteins 2016; 84(Suppl 1):323–348. © 2016 The Authors Proteins: Structure, Function, and Bioinformatics Published by Wiley Periodicals, Inc.
While the number of solved 3D protein structures continues to grow rapidly, the structural rules that distinguish protein-protein interactions between different structural families are still not clear. Here, we classify and analyse the secondary structural features and promiscuity of a comprehensive non-redundant set of domain family binding sites (DFBSs) and hetero domain-domain interactions (DDIs) extracted from our updated KBDOCK resource. We have partitioned 4001 DFBSs into five classes using their propensities for three types of secondary structural elements (“α” for helices, “β” for strands, and “γ” for irregular structure) and we have analysed how frequently these classes occur in DDIs. Our results show that β elements are not highly represented in DFBSs compared to α and γ elements. At the DDI level, all classes of binding sites tend to preferentially bind to the same class of binding sites and α/β contacts are significantly disfavored. Very few DFBSs are promiscuous: 80% of them interact with just one Pfam domain. About 50% of our Pfam domains bear only one single-partner DFBS and are therefore monogamous in their interactions with other domains. Conversely, promiscuous Pfam domains bear several DFBSs among which one or two are promiscuous, thereby multiplying the promiscuity of the concerned protein.
Comparing, classifying and modelling protein structural interactions can enrich our understanding of many biomolecular processes. This contribution describes Kbdock (http://kbdock.loria.fr/), a database system that combines the Pfam domain classification with coordinate data from the PDB to analyse and model 3D domain-domain interactions (DDIs). Kbdock can be queried using Pfam domain identifiers, protein sequences or 3D protein structures. For a given query domain or pair of domains, Kbdock retrieves and displays a non-redundant list of homologous DDIs or domain-peptide interactions in a common coordinate frame. Kbdock may also be used to search for and visualize interactions involving different, but structurally similar, Pfam families. Thus, structural DDI templates may be proposed even when there is little or no sequence similarity to the query domains.
Protein docking algorithms aim to calculate the three-dimensional (3D) structure of a protein complex starting from its unbound components. Although ab initio docking algorithms are improving, there is a growing need to use homology modeling techniques to exploit the rapidly increasing volumes of structural information that now exist. However, most current homology modeling approaches involve finding a pair of complete single-chain structures in a homologous protein complex to use as a 3D template, despite the fact that protein complexes are often formed from one or more domain-domain interactions (DDIs). To model 3D protein complexes by domain-domain homology, we have developed a case-based reasoning approach called KBDOCK which systematically identifies and reuses domain family binding sites from our database of nonredundant DDIs. When tested on 54 protein complexes from the Protein Docking Benchmark, our approach provides a near-perfect way to model single-domain protein complexes when full-homology templates are available, and it extends our ability to model more difficult cases when only partial or incomplete templates exist. These promising early results highlight the need for a new and diverse docking benchmark set, specifically designed to assess homology docking approaches.
Malika Smaïl-Tabbone合作论文数Batiment B - Equipe Orapilleur;Campus Scientifique8