Genome-wide phenotypic screens in the budding yeast Saccharomyces cerevisiae , enabled by its knockout collection, have produced the largest, richest, and most systematic phenotypic description of any organism. However, integrative analyses of this rich data source have been virtually impossible because of the lack of a central data repository and consistent metadata annotations. Here, we describe the aggregation, harmonization, and analysis of ~14,500 yeast knockout screens, which we call Yeast Phenome. Using this unique dataset, we characterized two unknown genes ( YHR045W and YGL117W ) and showed that tryptophan starvation is a by-product of many chemical treatments. Furthermore, we uncovered an exponential relationship between phenotypic similarity and intergenic distance, which suggests that gene positions in both yeast and human genomes are optimized for function.
The BioGRID (Biological General Repository for Interaction Datasets, thebiogrid.org) is an open-access database resource that houses manually curated protein and genetic interactions from multiple species including yeast, worm, fly, mouse, and human. The ~1.93 million curated interactions in BioGRID can be used to build complex networks to facilitate biomedical discoveries, particularly as related to human health and disease. All BioGRID content is curated from primary experimental evidence in the biomedical literature, and includes both focused low-throughput studies and large high-throughput datasets. BioGRID also captures protein post-translational modifications and protein or gene interactions with bioactive small molecules including many known drugs. A built-in network visualization tool combines all annotations and allows users to generate network graphs of protein, genetic and chemical interactions. In addition to general curation across species, BioGRID undertakes themed curation projects in specific aspects of cellular regulation, for example the ubiquitin-proteasome system, as well as specific disease areas, such as for the SARS-CoV-2 virus that causes COVID-19 severe acute respiratory syndrome. A recent extension of BioGRID, named the Open Repository of CRISPR Screens (ORCS, orcs.thebiogrid.org), captures single mutant phenotypes and genetic interactions from published high throughput genome-wide CRISPR/Cas9-based genetic screens. BioGRID-ORCS contains datasets for over 1,042 CRISPR screens carried out to date in human, mouse and fly cell lines. The biomedical research community can freely access all BioGRID data through the web interface, standardized file downloads, or via model organism databases and partner meta-databases.
A major obstacle to treating Alzheimer’s disease (AD) is our lack of understanding of the molecular mechanisms underlying selective neuronal vulnerability, which is a key characteristic of the disease. Here we present a framework to integrate high-quality neuron-type specific molecular profiles across the lifetime of the healthy mouse, which we generated using bacTRAP, with postmortem human functional genomics and quantitative genetics data. We demonstrate human-mouse conservation of cellular taxonomy at the molecular level for AD vulnerable and resistant neurons, identify specific genes and pathways associated with AD pathology, and pinpoint a specific functional gene module underlying selective vulnerability, enriched in processes associated with axonal remodeling, and affected by both amyloid accumulation and aging. Overall, our study provides a molecular framework for understanding the complex interplay between Aβ, aging, and neurodegeneration within the most vulnerable neurons in AD.
A key challenge for the diagnosis and treatment of complex human diseases is identifying their molecular basis. Here, we developed a unified computational framework, URSAHD (Unveiling RNA Sample Annotation for Human Diseases), that leverages machine learning and the hierarchy of anatomical relationships present among diseases to integrate thousands of clinical gene expression profiles and identify molecular characteristics specific to each of the hundreds of complex diseases. URSAHD can distinguish between closely related diseases more accurately than literature-validated genes or traditional differential-expression-based computational approaches and is applicable to any disease, including rare and understudied ones. We demonstrate the utility of URSAHD in classifying related nervous system cancers and experimentally verifying novel neuroblastoma-associated genes identified by URSAHD. We highlight the applications for potential targeted drug-repurposing and for quantitatively assessing the molecular response to clinical therapies. URSAHD is freely available for public use, including the use of underlying models, at ursahd.princeton.edu.
A great deal of information on the molecular genetics and biochemistry of model organisms has been reported in the scientific literature. However, this data is typically described in free text form and is not readily amenable to computational analyses. To this end, the BioGRID database systematically curates the biomedical literature for genetic and protein interaction data. This data is provided in a standardized computationally tractable format and includes structured annotation of experimental evidence. BioGRID curation necessarily involves substantial human effort by expert curators who must read each publication to extract the relevant information. Computational text-mining methods offer the potential to augment and accelerate manual curation. To facilitate the development of practical text-mining strategies, a new challenge was organized in BioCreative V for the BioC task, the collaborative Biocurator Assistant Task. This was a non-competitive, cooperative task in which the participants worked together to build BioC-compatible modules into an integrated pipeline to assist BioGRID curators. As an integral part of this task, a test collection of full text articles was developed that contained both biological entity annotations (gene/protein and organism/species) and molecular interaction annotations (protein–protein and genetic interactions (PPIs and GIs)). This collection, which we call the BioC-BioGRID corpus, was annotated by four BioGRID curators over three rounds of annotation and contains 120 full text articles curated in a dataset representing two major model organisms, namely budding yeast and human. The BioC-BioGRID corpus contains annotations for 6409 mentions of genes and their Entrez Gene IDs, 186 mentions of organism names and their NCBI Taxonomy IDs, 1867 mentions of PPIs and 701 annotations of PPI experimental evidence statements, 856 mentions of GIs and 399 annotations of GI evidence statements. The purpose, characteristics and possible future uses of the BioC-BioGRID corpus are detailed in this report. Database URL:http://bioc.sourceforge.net/BioC-BioGRID.html
The Precision Medicine Track in BioCreative VI aims to bring together the biomedical text mining community for a novel challenge: mining the biomedical literature in search of information of value to precision medicine initiatives such as mutations disrupting/affecting protein-protein interactions (PPI). The Precision Medicine track is organized into two tasks: 1) the triage task – focusing on selection of relevant PubMed articles describing PPI affected by mutations, and 2) the relation extraction task – focusing on extracting the interacting gene pairs for the interactions that are affected by the presence of a mutation. To support this track with an effective training dataset and limited curator time, the track organizers used a two-staged approach. First, for the creation of the training dataset, the organizers and curators worked on leveraging the information from expertly curated and publicly available PPI databases, augmenting it with a set of articles selected via publicly available state-of-the-art text mining tools. 4,082 PubMed articles were thus carefully reviewed, annotated and released for system development. They contained 1,729 articles labelled positive for curation, out of which, 597 contained 752 curated relations. The second stage pertained to the creation of the testing dataset, which consisted of 1,464 PubMed articles, previously not curated in any of the known PPI databases. These articles were highly likely to describe PPI and sequence variants according to several text mining tests. Each article in the testing dataset was annotated by at least two curators, for relevance relation extraction. Five BioGRID annotators participated and reviewed more than 600 articles each. The testing set contained 730 articles labelled positive for curation, out of which, 688 articles contained 930 curated relations. We detail here the data collection, manual review and annotation process. We give a report on the precision medicine track corpus characteristics. This analysis will provide useful information to developers and researchers for comparing and developing innovative text mining approaches for the https://thebiogrid.org/ BioCreative VI challenge and other Precision Medicine related applications. Keywords—corpus creation, manual annotation, protein-protein interaction, mutation, relation extraction, information extraction.
BioC is a simple XML format for text, annotations and relations, and was developed to achieve interoperability for biomedical text processing. Following the success of BioC in BioCreative IV, the BioCreative V BioC track addressed a collaborative task to build an assistant system for BioGRID curation. In this paper, we describe the framework of the collaborative BioC task and discuss our findings based on the user survey. This track consisted of eight subtasks including gene/protein/organism named entity recognition, protein-protein/genetic interaction passage identification and annotation visualization. Using BioC as their data-sharing and communication medium, nine teams, world-wide, participated and contributed either new methods or improvements of existing tools to address different subtasks of the BioC track. Results from different teams were shared in BioC and made available to other teams as they addressed different subtasks of the track. In the end, all submitted runs were merged using a machine learning classifier to produce an optimized output. The biocurator assistant system was evaluated by four BioGRID curators in terms of practical usability. The curators' feedback was overall positive and highlighted the user-friendly design and the convenient gene/protein curation tool based on text mining.Database URL: http://www.biocreative.org/tasks/biocreative-v/track-1-bioc/.
BioGRID has recently integrated chemical-protein data, thereby allowing the association of small molecules with key genetic and protein networks. Visualization of the drug-target associations has also been made possible by BioGRID’s newly improved interactive Network Viewer, which overlays the small molecule data on the gene/protein interactions. Combining curated genetic, protein and chemical interaction data in one resource should facilitate network-based approaches to drug discovery and drug repurposing.
The BioGRID database is an extensive repository of curated genetic and protein interactions for the budding yeast Saccharomyces cerevisiae, the fission yeast Schizosaccharomyces pombe, and the yeast Candida albicans SC5314, as well as for several other model organisms and humans. This protocol describes how to use the BioGRID website to query genetic or protein interactions for any gene of interest, how to visualize the associated interactions using an embedded interactive network viewer, and how to download data files for either selected interactions or the entire BioGRID interaction data set.
The Biological General Repository for Interaction Datasets (BioGRID) is a freely available public database that provides the biological and biomedical research communities with curated protein and genetic interaction data. Structured experimental evidence codes, an intuitive search interface, and visualization tools enable the discovery of individual gene, protein, or biological network function. BioGRID houses interaction data for the major model organism species—including yeast, nematode, fly, zebrafish, mouse, and human—with particular emphasis on the budding yeast Saccharomyces cerevisiae and the fission yeast Schizosaccharomyces pombe as pioneer eukaryotic models for network biology. BioGRID has achieved comprehensive curation coverage of the entire literature for these two major yeast models, which is actively maintained through monthly curation updates. As of September 2015, BioGRID houses approximately 335,400 biological interactions for budding yeast and approximately 67,800 interactions for fission yeast. BioGRID also supports an integrated posttranslational modification (PTM) viewer that incorporates more than 20,100 yeast phosphorylation sites curated through its sister database, the PhosphoGRID.
Abstract LAMHDI.org, a free web-based resource, helps researchers identify model systems to investigate disease mechanisms and therapies by bringing together information about diseases and model organisms. The LAMHDI portal allows search of disease models across species (non-human primates, zebrafish, mice, rats, flies, and yeast, with others in the pipeline) using gene orthology and pathway membership as key linkages. New work includes matching specific phenotypes and common pathways, a zebrafish atlas, and a graphical search based on spatial models of the brain. LAMHDI’s collaboration with BioGRID permits the exploration of networks linked to immune response, such as the NFkappaB and Interferon gamma signaling pathways. Interactions involving homologous members of these pathways provide a valuable resource for exploration across animal models. The goal is to facilitate the identification of models for disease research, make better use of existing model organisms and data about them, and provide the ability to discover new relationships between disease, phenotypes and genes that will further our understanding of disease. Companion efforts will speed the discovery and validation of novel drug candidates in areas as diverse as infectious disease, neuroscience, cardiovascular and metabolic disorders, autoimmunity, and cancer.
The Biological General Repository for Interaction Datasets (BioGRID: http://thebiogrid.org) is an open access database that houses genetic and protein interactions curated from the primary biomedical literature for all major model organism species and humans. As of September 2014, the BioGRID contains 749,912 interactions as drawn from 43,149 publications that represent 30 model organisms. This interaction count represents a 50% increase compared to our previous 2013 BioGRID update. BioGRID data are freely distributed through partner model organism databases and meta-databases and are directly downloadable in a variety of formats. In addition to general curation of the published literature for the major model species, BioGRID undertakes themed curation projects in areas of particular relevance for biomedical sciences, such as the ubiquitin-proteasome system and various human disease-associated interaction networks. BioGRID curation is coordinated through an Interaction Management System (IMS) that facilitates the compilation interaction records through structured evidence codes, phenotype ontologies, and gene annotation. The BioGRID architecture has been improved in order to support a broader range of interaction and post-translational modification types, to allow the representation of more complex multi-gene/protein interactions, to account for cellular phenotypes through structured ontologies, to expedite curation through semi-automated text-mining approaches, and to enhance curation quality control.
During development, Sonic hedgehog (Shh) regulates the proliferation of cerebellar granule neuron precursors (GNPs) in part via expression of Nmyc. We present evidence supporting a novel role for the Mad family member Mad3 in the Shh pathway to regulate Nmyc expression and GNP proliferation. Mad3 mRNA is transiently expressed in GNPs during proliferation. Cultured GNPs express Mad3 in response to Shh stimulation in a cyclopamine-dependent manner. Mad3 is necessary for Shh-dependent GNP proliferation as measured by bromodeoxyuridine incorporation and Nmyc expression. Furthermore, Mad3 overexpression, but not that of other Mad proteins, is sufficient to induce GNP proliferation in the absence of Shh. Structure-function analysis revealed that Max dimerization and recruitment of the mSin3 corepressor are required for Mad3-mediated GNP proliferation. Surprisingly, basic-domain-dependent DNA binding of Mad3 is not required, suggesting that Mad3 interacts with other DNA binding proteins to repress transcription. Interestingly, cerebellar tumors and pretumor cells derived from patched heterozygous mice express high levels of Mad3 compared with adjacent normal cerebellar tissue. Our studies support a novel role for Mad3 in cerebellar GNP proliferation and possibly tumorigenesis, and they challenge the current paradigm that Mad3 should antagonize Nmyc by competition for direct DNA binding via Max dimerization.