The development process for the beta cell genomics application ontology (BCGO) is described. This process should be generally applicable and consists of integration of a subset of reference ontologies. A key element is use of the Ontology for Biomedical Investigation (OBI) as an ontology framework. Another element is enriching ontologies using existing patterns when needed. The ontology is validated in three aspects based on our needs including data annotation, queries and automated classification. The BCGO is available on: http://purl.obolibrary.org/obo/bcgo.owl.
INTRODUCTION: The National Institute on Aging Genetics of Alzheimer's Disease Data Storage Site Alzheimer's Genomics Database (GenomicsDB) is a public knowledge base of Alzheimer's disease (AD) genetic datasets and genomic annotations. METHODS: GenomicsDB uses a custom systems architecture to adopt and enforce rigorous standards that facilitate harmonization of AD-relevant genome-wide association study summary statistics datasets with functional annotations, including over 230 million annotated variants from the AD Sequencing Project. RESULTS: GenomicsDB generates interactive reports compiled from the harmonized datasets and annotations. These reports contextualize AD-risk associations in a broader functional genomic setting and summarize them in the context of functionally annotated genes and variants. DISCUSSION: Created to make AD-genetics knowledge more accessible to AD researchers, the GenomicsDB is designed to guide users unfamiliar with genetic data in not only exploring but also interpreting this ever-growing volume of data. Scalable and interoperable with other genomics resources using data technology standards, the GenomicsDB can serve as a central hub for research and data analysis on AD and related dementias.
The Alzheimer’s Genomics Database (GenomicsDB) is an interactive knowledgebase for Alzheimer's disease (AD) genetics that offers a platform for data sharing, discovery, and analysis to accelerate research on AD and AD related dementias (ADRD). The resource provides unrestricted access to GWAS summary statistics datasets, variant annotations, and meta-analysis results deposited at the NIA Genetics of Alzheimer’s Disease Data Storage Site (NIAGADS). Here we introduce an updated version of the GenomicsDB, now mapped against the latest human genome assembly, GRCh38. The GenomicsDB allows users to interactively mine AD/ADRD summary statistics datasets or visually inspect them and compare with annotated variant tracks and GRCh38 aligned functional genomics tracks from the FILER functional genomics repository (Kuksa et al. 2022). Available tracks include key AD datasets from large-scale sequencing projects such as IGAP (Lambert et al. 2013; Kunkle et al. 2019) and the ADGC (e.g., Naj et al. 2011) that have been lifted over from earlier genome builds. Available annotated variants include those from the AD Sequencing Project’s (ADSP) Release3 WGS, called from nearly 17k samples. All variants reported in the GenomicsDB are annotated using an assembly specific version of ADSP annotation pipeline (Butkiewicz et al. 2018). The GenomicsDB provides a REST API for programmatic access. Supported queries return complete annotations for single or bulk lookups of GenomicsDB entries (genes, variants, genomic loci) and ranged queries of datasets as JSON documents. The GenomicsDB currently hosts >70 datasets and reports >150 million annotated variants from the ADSP and 23 AD/ADRD GWAS studies, mapped to GRCh38. Keeping up with community standards allows this resource to provide ongoing support for AD/ADRD research by making it possible to integrate AD/ADRD relevant `omics datasets generated by the ADSP and others who have adopted GRCh38 as their reference genome. The NIAGADS Alzheimer’s Genomics Database v. GRCh38 (www.niagads.org/genomics) provides an up-to-date and accessible platform that facilitates unrestricted sharing of genetic knowledge underpinning AD/ADRD. To best serve our user community, NIAGADS will continue to provide the GRCh37.p19 version of the Alzheimer’s Genomics Database (including programmatic access) in archival format.
NIAGADS is a national genomics data repository that facilitates access of genotypic and sequencing data to qualified investigators for the study of the genetics of Alzheimer’s disease (AD) and related neurological diseases. Collaborations with large consortia and centers such as the Alzheimer’s Disease Genetics Consortium (ADGC), Cohorts for Heart and Aging Research in Genomic Epidemiology (CHARGE) Consortium, the Alzheimer’s Disease Sequencing Project (ADSP), and the Genome Center for Alzheimer’s Disease (GCAD) allow NIAGADS to lead the effort in managing large AD datasets that can be easily accessed and fully utilized by the research community. NIAGADS is supported by National Institute on Aging (NIA) under a cooperative agreement. All data derived from NIA funded AD genetics studies are expected to be deposited in NIAGADS or another NIA approved site. NIAGADS manages a Data Sharing Service (DSS) that facilitates the deposition and sharing of genomic data and association results with approved users in the neurodegenerative research community. In addition, researchers are able to freely use the NIAGADS Alzheimer’s Genomics Database ( www.niagads.org/genomics/ ) to search annotation resources that link published AD studies to AD-relevant sequence features and genome-wide annotations. As of January 2022, NIAGADS houses 82 datasets comprised of >93,000 samples including GWAS, sequencing, gene expression, annotations, deep phenotypes, and summary statistics. Qualified investigators can retrieve ADSP sequencing data with ease and flexibility through the NIAGADS DSS. As of January 2022, the ADSP and other contributing studies have completed whole exome sequencing (WES) of 20,503 samples and whole-genome sequencing (WGS) of 16,905 samples. Raw WES and WGS files, quality controlled VCF files, and phenotype data files are available via qualified access. The next round of sequencing currently underway will generate around 18,000 additional genomes to be released at the middle of 2022. NIAGADS is a rich resource for AD researchers, with the goal of facilitating advances in Alzheimer’s genetics research. By housing datasets from many projects and institutions, NIAGADS enables AD researchers to meet their research goals more efficiently. Datasets, guidelines, and new features are available on our website at https://www.niagads.org .
ABSTRACT To gain a deeper understanding of pancreatic β-cell development, we used iterative weighted gene correlation network analysis to calculate a gene co-expression network (GCN) from 11 temporally and genetically defined murine cell populations. The GCN, which contained 91 distinct modules, was then used to gain three new biological insights. First, we found that the clustered protocadherin genes are differentially expressed during pancreas development. Pcdhγ genes are preferentially expressed in pancreatic endoderm, Pcdhβ genes in nascent islets, and Pcdhα genes in mature β-cells. Second, after extracting sub-networks of transcriptional regulators for each developmental stage, we identified 81 zinc finger protein (ZFP) genes that are preferentially expressed during endocrine specification and β-cell maturation. Third, we used the GCN to select three ZFPs for further analysis by CRISPR mutagenesis of mice. Zfp800 null mice exhibited early postnatal lethality, and at E18.5 their pancreata exhibited a reduced number of pancreatic endocrine cells, alterations in exocrine cell morphology, and marked changes in expression of genes involved in protein translation, hormone secretion and developmental pathways in the pancreas. Together, our results suggest that developmentally oriented GCNs have utility for gaining new insights into gene regulation during organogenesis.
Newly differentiated pancreatic β cells lack proper insulin secretion profiles of mature functional β cells. The global gene expression differences between paired immature and mature β cells have been studied, but the dynamics of transcriptional events, correlating with temporal development of glucose-stimulated insulin secretion (GSIS), remain to be fully defined. This aspect is important to identify which genes and pathways are necessary for β-cell development or for maturation, as defective insulin secretion is linked with diseases such as diabetes. In this study, we assayed through RNA sequencing the global gene expression across six β-cell developmental stages in mice, spanning from β-cell progenitor to mature β cells. A computational pipeline then selected genes differentially expressed with respect to progenitors and clustered them into groups with distinct temporal patterns associated with biological functions and pathways. These patterns were finally correlated with experimental GSIS, calcium influx, and insulin granule formation data. Gene expression temporal profiling revealed the timing of important biological processes across β-cell maturation, such as the deregulation of β-cell developmental pathways and the activation of molecular machineries for vesicle biosynthesis and transport, signal transduction of transmembrane receptors, and glucose-induced Ca2+ influx, which were established over a week before β-cell maturation completes. In particular, β cells developed robust insulin secretion at high glucose several days after birth, coincident with the establishment of glucose-induced calcium influx. Yet the neonatal β cells displayed high basal insulin secretion, which decreased to the low levels found in mature β cells only a week later. Different genes associated with calcium-mediated processes, whose alterations are linked with insulin resistance and deregulation of glucose homeostasis, showed increased expression across β-cell stages, in accordance with the temporal acquisition of proper GSIS. Our temporal gene expression pattern analysis provided a comprehensive database of the underlying molecular components and biological mechanisms driving β-cell maturation at different temporal stages, which are fundamental for better control of the in vitro production of functional β cells from human embryonic stem/induced pluripotent cell for transplantation-based type 1 diabetes therapy.
Although NLP has been used to support cancer research more broadly, the development of NLP algorithms to extract evidence of progression from clinical notes to support lung cancer research is still in its infancy. In this study, we trained supervised machine learning classifiers using rich semantic features to detect and classify statements of progression status from radiology exams. Our progression status classifier achieves high F1-scores for detecting and discerning progression (0.80), stable (0.82), and not relevant (0.92) sentences, demonstrating promising performance. We are actively integrating these extractions with structured electronic health record data using ontologies to instantiate a longitudinal model of progression among non-small cell lung cancer patients.
NIAGADS is a national genomics data repository that facilitates access of genotypic and sequencing data to qualified investigators for the study of the genetics of Alzheimer’s disease (AD) and related neurological diseases. Collaborations with large consortia and centers such as the Alzheimer’s Disease Genetics Consortium (ADGC), Cohorts for Heart and Aging Research in Genomic Epidemiology (CHARGE) Consortium, the Alzheimer’s Disease Sequencing Project (ADSP), and the Genome Center for Alzheimer’s Disease (GCAD) allow NIAGADS to lead the effort in managing large AD datasets that can be easily accessed and fully utilized by the research community. NIAGADS is supported by National Institute on Aging (NIA) under a cooperative agreement. All data derived from NIA funded AD genetics studies are expected to be deposited in NIAGADS or another NIA approved site. NIAGADS manages a Data Sharing Service (DSS) that facilitates the deposition and sharing of genomic data and association results with approved users in the neurodegenerative research community. In addition, researchers are able to freely use the NIAGADS Alzheimer’s Genomics Database (www.niagads.org/genomics/) to search annotation resources that link published AD studies to AD-relevant sequence features and genome-wide annotations. As of January 2021, NIAGADS houses 74 datasets comprised of >90,000 samples including GWAS, sequencing, gene expression, annotations, deep phenotypes, and summary statistics. Qualified investigators can retrieve ADSP sequencing data with ease and flexibility through the NIAGADS DSS. As of February 2021, the ADSP and other contributing studies have completed whole exome sequencing (WES) of 20,504 samples and whole-genome sequencing (WGS) of 16,908 samples. Raw WES and WGS files, quality controlled VCF files, and phenotype data files are available via qualified access. The next round of sequencing currently underway will generate around 18,000 additional genomes to be released at the end of 2021. NIAGADS is a rich resource for AD researchers, with the goal of facilitating advances in Alzheimer’s genetics research. By housing datasets from many projects and institutions, NIAGADS enables AD researchers to meet their research goals more efficiently. Datasets, guidelines, and new features are available on our website at https://www.niagads.org.
The Ontology for Biomedical Investigations (OBI) underwent a focused review of assay term annotations, logic and hierarchy with a goal to improve and standardize these terms. As a result, inconsistencies in W3C Web Ontology Language (OWL) expressions were identified and corrected, and additionally, standardized design patterns and a formalized template to maintain them were developed. We describe here this informative and productive process to describe the specific benefits and obstacles for OBI and the universal lessons for similar projects.
ABSTRACTOver the last two decades, molecular biology has been changed by the introduction of high-throughput technologies. Data sharing requirements have prompted the establishment of persistent data archives. A standardized approach for recording and managing these data was first proposed in the Minimal Information About a Microarray Experiment (MIAME) guidelines. The Minimal Information about a high throughput nucleotide Sequencing Experiment (MINSEQE) proposal was introduced in 2008 as a logical extension of the guidelines to next-generation sequencing (NGS) technologies used for transcriptome analysis.We present a historical snapshot of the data-sharing situation focusing on transcriptomics data from both microarray and RNA-sequencing experiments published between 2009 and 2013, a period during which RNA-seq studies became increasingly popular for transcriptome analysis. We assess how much data from RNA-seq based experiments is actually available in persistent data archives, compared to data derived from microarray based experiments, and evaluate how these types of data differ. Based on this analysis, we provide recommendations to improve RNA-seq data availability, reusability, and reproducibility.
Biological ontologies are used to organize, curate, and interpret the vast quantities of data arising from biological experiments. While this works well when using a single ontology, integrating multiple ontologies can be problematic, as they are developed independently, which can lead to incompatibilities. The Open Biological and Biomedical Ontologies (OBO) Foundry was created to address this by facilitating the development, harmonization, application, and sharing of ontologies, guided by a set of overarching principles. One challenge in reaching these goals was that the OBO principles were not originally encoded in a precise fashion, and interpretation was subjective. Here we show how we have addressed this by formally encoding the OBO principles as operational rules and implementing a suite of automated validation checks and a dashboard for objectively evaluating each ontology’s compliance with each principle. This entailed a substantial effort to curate metadata across all ontologies and to coordinate with individual stakeholders. We have applied these checks across the full OBO suite of ontologies, revealing areas where individual ontologies require changes to conform to our principles. Our work demonstrates how a sizable federated community can be organized and evaluated on objective criteria that help improve overall quality and interoperability, which is vital for the sustenance of the OBO project and towards the overall goals of making data FAIR.
The Alzheimer's Genomics Database (NIAGADS GenomicsDB) provides public access to GWAS summary statistics datasets deposited at the NIA Genetics of Alzheimer's Disease Data Storage Site (NIAGADS). The NIAGADS GenomicsDB makes available 69 summary statistics datasets and reports >150 million annotated variants from 23 GWAS studies of Alzheimer's disease and related dementias (AD/ADRD). Programmatic access to the database is essential for many bioinformatics applications dealing with AD/ADRD genetics.The NIAGADS GenomicsDB is powered by a big data optimized relational database system. It uses ontologies to consistently annotate datasets, facilitating data harmonization and efficient real-time data mining. AD/ADRD-risk associated variants from GWAS datasets are annotated using the Alzheimer's Disease Sequencing Project's (ADSP) annotation pipeline (Butkiewicz et al. 2018). Variants are also mapped to proximal and co-located genes and regulatory regions using the NIAGADS FILER functional genomics data repository (https://lisanwanglab.org/FILER). We have created a RESTful Application Programming Interface (API) that facilitates the integration of NIAGADS GenomicsDB datasets and annotations into software applications. The web service supports queries that return annotated GenomicsDB entries (genes, variants, genomic loci) and GWAS summary statistics results in the form of JSON documents. Results may contain functional annotations, genetic evidence for AD/ADRD risk, and/or sequence information associated with an entry. Calls to the API are based on simple URLs that specify the entry type and data required. The API also provides endpoints for querying variant linkage disequilibrium blocks, variant tracks (annotations and summary statistics) compatible with the IGV.js genome browser (https://github.com/igvteam/igv.js/), and LocusZoom.js (https://github.com/statgen/locuszoom/) renderings of the GWAS summary statistics datasets.The NIAGADS GenomicsDB API establishes a simple framework that enables sharing and integration of annotated NIAGADS GenomicsDB entries and datasets. Use of the API requires only a HTTPS library and JSON parser, permitting large-scale programmatic access and analysis of AD/ADRD-linked variants independent of any specific programming language.The NIAGADS GenomicsDB API enables the easy retrieval of a wide range of genetic evidence for AD/ADRD, promoting data sharing and reuse among the AD/ADRD research community. Information about accessing the REST API as it becomes available can be found on the NIAGADS GenomicsDB website (https://www.niagads.org/genomics).
Erythroid cell formation critically depends upon signals transduced via EPO/EPOR/JAK2 complexes. This includes not only core response modules (e.g., JAK2/STAT5, RAS/MEK/ERK), but also specialized effectors (e.g., Erythroferrone, ASCT2 glutamine transport, Spi2A). By employing phospho-proteomics and a human erythroblastic cell model, we presently identify 121 new EPO target proteins, together with their EPO- modulated domains and phosphosites. Gene Ontology enrichment for ‘Molecular Function’ identified adaptor proteins as one top EPO target category. This includes a novel EPOR/JAK2-coupled network of actin assemblage modifiers, with adaptors DLG-1, DLG-3, WAS, WASL and CD2AP as prime components. ‘Cellular Component’ GO analysis further identified 19 new EPO- modulated cytoskeletal targets including the erythroid cytoskeletal targets SPECTRIN-A, SPECTRIN–B, ADDUCIN-2 and GLYCOPHORIN-C. In each, EPO-induced phosphorylation occurred at p-Y sites and subdomains that suggest coordinated regulation by EPO of the erythroid cytoskeleton. GO analysis of ‘Biological AUTHOR CONTRIBUTIONS For phospho-PTM based LC-MS/MS analyses of EPO target proteins and associated data mining, DMW, MAH, and MPS were prime contributors. In bioinformatic studies, EA and CS assembled phospho-PTM data into an upgraded Erythron Database platform, performed gene ontology enrichment of EPO target sets, and contributed to Cytoscape-based graphics. Tabulated data (Supplemental Tables 1–7) were assembled by EA and MAH. Studies of EPO effects on TXNIP phosphorylation were performed by MAH and AW (including the preparation of antibodies to p-TXNIP T349 by AW and Cell Signaling Technology). TXNIP knockdown studies were via DMW and MAH, with technical contributions by Ruth Asch. In cell phenotyping analyses of primary human erythroid progenitors, DMW together with EJ established cell culture systems and performed flow cytometry phenotyping of dexamethasone cultures. For primary human erythroid cell culture phase I-III systems, cell phenotype analyses were performed by MAH. For manuscript construction, all authors contributed to data analysis, interpretations and writing. Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain. CONFLICTS OF INTERESTS: The authors declare no competing interests. Processes’ further revealed metabolic regulators as a likewise unexpected EPO target set. Targets included ALDOLASE-A, PYRUVATE DEHYDROGENASE-A1 and THIOREDOXIN-INTERACTING PROTEIN (TXNIP), with EPO-modulated pY-sites in each occurring within functional subdomains. In TXNIP, EPO- induced phosphorylation occurred at novel p-T349 and p-S358 sites, and was paralleled by rapid increases in TXNIP levels. In UT7epo-E and primary human HSC-derived erythroid progenitor cells, lentivirus-mediated shRNA knockdown studies revealed novel pro-erythropoietic roles for TXNIP. Specifically, TXNIP’s knockdown sharply inhibited c-KIT expression; compromised EPO dose-dependent erythroblast proliferation and survival; and delayed late-stage erythroblast formation. Overall, new insight is provided into EPO’s diverse action mechanisms, and TXNIP’s contributions to EPO-dependent human erythropoiesis. Graphical phosphorylation of p-Y440 within a CALM interacting domain. In GYPC, EPO induced the phosphorylation of p-Y126 (C-terminal cytoplasmic domain). D,E: Within EZRIN (EZR), EPO regulated the phosphorylation of p-Y499 phosphorylation within an actin-interacting ERMAD subdomain. In CALM-1, EPO regulated the phosphorylation of p-Y100 within calcium binding domain III. F: Findings define a network of interacting cytoskeletal factors that are rapidly and coordinately regulated by EPO at unique novel p-Y sites within functionally important subdomains.
Although stroke is an established risk factor for Alzheimer’s disease (AD), the role vascular factors play in driving AD remains uncertain. Here we leverage GWAS summary statistics to assess the genetic correlation (pleiotropy) between stroke and AD and introduce a novel fine‐mapping approach to pinpoint shared causal variants in the context of higher‐order genome organization.
Standardizing clinical information in a common data model is important for promoting interoperability and facilitating high quality research. Semantic Web technologies such as Resource Description Framework can be utilized to their full potential when a clinical data model accurately reflects the reality of the clinical situation it describes. To this end, the Open Biomedical Ontologies Foundry provides a set of ontologies that conform to the principles of realism and can be used to create a realism-based clinical data model. However, the challenge of programmatically defining such a model and loading data from disparate sources into the model has not been addressed by pre-existing software solutions. The PennTURBO Semantic Engine is a tool developed at the University of Pennsylvania that works in conjunction with data aggregation software to transform source-specific RDF data into a source-independent, realism-based data model. This system sources classes from an application ontology and specifically defines how instances of those classes may relate to each other. Additionally, the system defines and executes RDF data transformations by launching dynamically generated SPARQL update statements. The Semantic Engine was designed as a generalizable RDF data standardization tool, and is able to work with various data models and incoming data sources. Its human-readable configuration files can easily be shared between institutions, providing the basis for collaboration on a standard realism-based clinical data model.
The TURBO Medication Mapper (TMM) identifies terms from RxNorm that best represent a list of medication strings, like one would find in a Clinical Data Warehouse (CDW). TMM has several differentiating characteristics, compared to other tools: the machine learning component does not require a human-curated gold standard for training; normalizations are applied to source-specific language in the strings (instead of just excluding them); the confidence of each mapping is represented as a relationship to the absolute truth, along with a 0.0-1.0 score; the results, along with supporting knowledge, are saved into an RDF graph and a Solr document database is generated. Queries for drug classes like “statins” are based on OBO foundry ontologies like the Drug Ontology (DrOn) and ChEBI, and they return more results than multiple SQL search strategies over the CDW, with few false positives. TMM is available for download from GitHub.
The concept of open data has been gaining traction as a mechanism to increase data use, ensure that data are preserved over time, and accelerate discovery. While epidemiology data sets are increasingly deposited in databases and repositories, barriers to access still remain. ClinEpiDB was constructed as an open-access online resource for clinical and epidemiologic studies by leveraging the extensive web toolkit and infrastructure of the Eukaryotic Pathogen Database Resources (EuPathDB; a collection of databases covering 170+ eukaryotic pathogens, relevant related species, and select hosts) combined with a unified semantic web framework. Here we present an intuitive point-and-click website that allows users to visualize and subset data directly in the ClinEpiDB browser and immediately explore potential associations. Supporting study documentation aids contextualization, and data can be downloaded for advanced analyses. By facilitating access and interrogation of high-quality, large-scale data sets, ClinEpiDB aims to spur collaboration and discovery that improves global health.
Jonathan Schug合作论文数the University of Pennsylvania16
Eileen Kraemer合作论文数Computer Science Department;University of Georgia9