The Human Phenotype Ontology (HPO) is a widely used resource that comprehensively organizes and defines the phenotypic features of human disease, enabling computational inference and supporting genomic and phenotypic analyses through semantic similarity and machine learning algorithms. The HPO has widespread applications in clinical diagnostics and translational research, including genomic diagnostics, gene-disease discovery, and cohort analytics. In recent years, groups around the world have developed translations of the HPO from English to other languages, and the HPO browser has been internationalized, allowing users to view HPO term labels and in many cases synonyms and definitions in ten languages in addition to English. Since our last report, a total of 2239 new HPO terms and 49235 new HPO annotations were developed, many in collaboration with external groups in the fields of psychiatry, arthrogryposis, immunology and cardiology. The Medical Action Ontology (MAxO) is a new effort to model treatments and other measures taken for clinical management. Finally, the HPO consortium is contributing to efforts to integrate the HPO and the GA4GH Phenopacket Schema into electronic health records (EHRs) with the goal of more standardized and computable integration of rare disease data in EHRs.
Bridging the gap between genetic variations, environmental determinants, and phenotypic outcomes is critical for supporting clinical diagnosis and understanding mechanisms of diseases. It requires integrating open data at a global scale. The Monarch Initiative advances these goals by developing open ontologies, semantic data models, and knowledge graphs for translational research. The Monarch App is an integrated platform combining data about genes, phenotypes, and diseases across species. Monarch's APIs enable access to carefully curated datasets and advanced analysis tools that support the understanding and diagnosis of disease for diverse applications such as variant prioritization, deep phenotyping, and patient profile-matching. We have migrated our system into a scalable, cloud-based infrastructure; simplified Monarch's data ingestion and knowledge graph integration systems; enhanced data mapping and integration standards; and developed a new user interface with novel search and graph navigation features. Furthermore, we advanced Monarch's analytic tools by developing a customized plugin for OpenAI’s ChatGPT to increase the reliability of its responses about phenotypic data, allowing us to interrogate the knowledge in the Monarch graph using state-of-the-art Large Language Models. The resources of the Monarch Initiative can be found at monarchinitiative.org and its corresponding code repository at github.com/monarch-initiative/monarch-app.
Knowledge graphs (KGs) are a powerful approach for integrating heterogeneous data and making inferences in biology and many other domains, but a coherent solution for constructing, exchanging, and facilitating the downstream use of knowledge graphs is lacking. Here we present KG-Hub, a platform that enables standardized construction, exchange, and reuse of knowledge graphs. Features include a simple, modular extract-transform-load (ETL) pattern for producing graphs compliant with Biolink Model (a high-level data model for standardizing biological data), easy integration of any OBO (Open Biological and Biomedical Ontologies) ontology, cached downloads of upstream data sources, versioned and automatically updated builds with stable URLs, web-browsable storage of KG artifacts on cloud infrastructure, and easy reuse of transformed subgraphs across projects. Current KG-Hub projects span use cases including COVID-19 research, drug repurposing, microbial-environmental interactions, and rare disease research. KG-Hub is equipped with tooling to easily analyze and manipulate knowledge graphs. KG-Hub is also tightly integrated with graph machine learning (ML) tools which allow automated graph machine learning, including node embeddings and training of models for link prediction and node classification.
Existing phenotype ontologies were originally developed to represent phenotypes that manifest as a character state in relation to a wild-type or other reference. However, these do not include the phenotypic trait or attribute categories required for the annotation of genome-wide association studies (GWAS), Quantitative Trait Loci (QTL) mappings or any population-focussed measurable trait data. The integration of trait and biological attribute information with an ever increasing body of chemical, environmental and biological data greatly facilitates computational analyses and it is also highly relevant to biomedical and clinical applications. The Ontology of Biological Attributes (OBA) is a formalised, species-independent collection of interoperable phenotypic trait categories that is intended to fulfil a data integration role. OBA is a standardised representational framework for observable attributes that are characteristics of biological entities, organisms, or parts of organisms. OBA has a modular design which provides several benefits for users and data integrators, including an automated and meaningful classification of trait terms computed on the basis of logical inferences drawn from domain-specific ontologies for cells, anatomical and other relevant entities. The logical axioms in OBA also provide a previously missing bridge that can computationally link Mendelian phenotypes with GWAS and quantitative traits. The term components in OBA provide semantic links and enable knowledge and data integration across specialised research community boundaries, thereby breaking silos.
Despite progress in the development of standards for describing and exchanging scientific information, the lack of easy-to-use standards for mapping between different representations of the same or similar objects in different databases poses a major impediment to data integration and interoperability. Mappings often lack the metadata needed to be correctly interpreted and applied. For example, are two terms equivalent or merely related? Are they narrow or broad matches? Are they associated in some other way? Such relationships between the mapped terms are often not documented, leading to incorrect assumptions and making them hard to use in scenarios that require a high degree of precision (such as diagnostics or risk prediction). Also, the lack of descriptions of how mappings were done makes it hard to combine and reconcile mappings, particularly curated and automated ones. The Simple Standard for Sharing Ontological Mappings (SSSOM) addresses these problems by: 1. Introducing a machine-readable and extensible vocabulary to describe metadata that makes imprecision, inaccuracy and incompleteness in mappings explicit. 2. Defining an easy to use table-based format that can be integrated into existing data science pipelines without the need to parse or query ontologies, and that integrates seamlessly with Linked Data standards. 3. Implementing open and community-driven collaborative workflows designed to evolve the standard continuously to address changing requirements and mapping practices. 4. Providing reference tools and software libraries for working with the standard. In this paper, we present the SSSOM standard, describe several use cases, and survey some existing work on standardizing the exchange of mappings, with the goal of making mappings Findable, Accessible, Interoperable, and Reusable (FAIR). The SSSOM specification is at http://w3id.org/sssom/spec.
Within clinical, biomedical, and translational science, an increasing number of projects are adopting graphs for knowledge representation. Graph-based data models elucidate the interconnectedness among core biomedical concepts, enable data structures to be easily updated, and support intuitive queries, visualizations, and inference algorithms. However, knowledge discovery across these "knowledge graphs" (KGs) has remained difficult. Data set heterogeneity and complexity; the proliferation of ad hoc data formats; poor compliance with guidelines on findability, accessibility, interoperability, and reusability; and, in particular, the lack of a universally accepted, open-access model for standardization across biomedical KGs has left the task of reconciling data sources to downstream consumers. Biolink Model is an open-source data model that can be used to formalize the relationships between data structures in translational science. It incorporates object-oriented classification and graph-oriented features. The core of the model is a set of hierarchical, interconnected classes (or categories) and relationships between them (or predicates) representing biomedical entities such as gene, disease, chemical, anatomic structure, and phenotype. The model provides class and edge attributes and associations that guide how entities should relate to one another. Here, we highlight the need for a standardized data model for KGs, describe Biolink Model, and compare it with other models. We demonstrate the utility of Biolink Model in various initiatives, including the Biomedical Data Translator Consortium and the Monarch Initiative, and show how it has supported easier integration and interoperability of biomedical KGs, bringing together knowledge from multiple sources and helping to realize the goals of translational science.
This article explores the planning and implementation process for a project with promise, a community partnership school for a historically low-performing elementary school using an asset-based community development approach. We offer insights into the community needs assessment process that enabled four key community partners to identify needs and projects for the school and surrounding community. The community partnership school draws its strength from four local organizations assimilating their expertise and resources on focal areas for community engagement. Beyond organizational resources, the partners also developed local networks and resources that could be useful for the community. Building on the asset-based community development model, insights and challenges are presented for others seeking to employ a similar approach to mobilize assets for student success and community engagement.
Wikidata is a community-maintained knowledge base that has been assembled from repositories in the fields of genomics, proteomics, genetic variants, pathways, chemical compounds, and diseases, and that adheres to the FAIR principles of findability, accessibility, interoperability and reusability. Here we describe the breadth and depth of the biomedical knowledge contained within Wikidata, and discuss the open-source tools we have built to add information to Wikidata and to synchronize it with source databases. We also demonstrate several use cases for Wikidata, including the crowdsourced curation of biomedical ontologies, phenotype-based diagnosis of disease, and drug repurposing.
In biology and biomedicine, relating phenotypic outcomes with genetic variation and environmental factors remains a challenge: patient phenotypes may not match known diseases, candidate variants may be in genes that haven’t been characterized, research organisms may not recapitulate human or veterinary diseases, environmental factors affecting disease outcomes are unknown or undocumented, and many resources must be queried to find potentially significant phenotypic associations. The Monarch Initiative (https://monarchinitiative.org) integrates information on genes, variants, genotypes, phenotypes and diseases in a variety of species, and allows powerful ontology-based search. We develop many widely adopted ontologies that together enable sophisticated computational analysis, mechanistic discovery and diagnostics of Mendelian diseases. Our algorithms and tools are widely used to identify animal models of human disease through phenotypic similarity, for differential diagnostics and to facilitate translational research. Launched in 2015, Monarch has grown with regards to data (new organisms, more sources, better modeling); new API and standards; ontologies (new Mondo unified disease ontology, improvements to ontologies such as HPO and uPheno); user interface (a redesigned website); and community development. Monarch data, algorithms and tools are being used and extended by resources such as GA4GH and NCATS Translator, among others, to aid mechanistic discovery and diagnostics.
The accelerating growth of genomic and proteomic information for Chlamydia species, coupled with unique biological aspects of these pathogens, necessitates bioinformatic tools and features that are not provided by major public databases. To meet these growing needs, we developed ChlamBase, a model organism database for Chlamydia that is built upon the WikiGenomes application framework, and Wikidata, a community-curated database. ChlamBase was designed to serve as a central access point for genomic and proteomic information for the Chlamydia research community. ChlamBase integrates information from numerous external databases, as well as important data extracted from the literature that are otherwise not available in structured formats that are easy to use. In addition, a key feature of ChlamBase is that it empowers users in the field to contribute new annotations and data as the field advances with continued discoveries. ChlamBase is freely and publicly available at chlambase.org.
Wikidata, a project of the Wikimedia Foundation, is an openly editable, semantic web-compatible framework for knowledge management. Wikidata has a large and active community that contributes to, maintains, and improves the quality of the data in Wikidata as well as deciding how the data itself should be represented. Our team has been populating Wikidata with a foundational semantic network linking genes, proteins, drugs, and diseases. Upon this foundation, we hope to stimulate the growth of this knowledge graph that can be used to build new knowledge-based applications that drive new discoveries. A cornerstone of the Wikidata knowledge graph is its built-in model for tracking the evidence underlying claims. For any claim in the graph (e.g. gene A regulates gene B), it is possible to provide evidence supporting or refuting that claim. The manner in which these evidence statements can be constructed is left open for the community to decide. Therefore, the patterns for representing the semantics of the associated evidence and the provenance trails linking back to the original sources of information must be defined and consistently used, such that this information is easily accessible by the end user or by software that exposes this information to end users.
With the advancement of genome sequencing technologies, new genomes are being sequenced daily. While these sequences are deposited in publicly available data warehouses, their functional and genomic annotations (beyond genes which are predicted automatically) mostly reside in the text of primary publications. Professional curators are hard at work extracting those annotations from the literature for the most studied organisms and depositing them in structured databases. However, the resources don’t exist to fund the comprehensive curation of the thousands of newly sequenced organisms in this manner. Here, we describe WikiGenomes ( wikigenomes.org ), a web application that facilitates the consumption and curation of genomic data by the entire scientific community. WikiGenomes is based on Wikidata, an openly editable knowledge graph with the goal of aggregating published knowledge into a free and open database. WikiGenomes empowers the individual genomic researcher to contribute their expertise to the curation effort and integrates the knowledge into Wikidata, enabling it to be accessed by anyone without restriction.
Background The biology of recurrent or long-term infections of humans by Chlamydia trachomatis is poorly understood. Because repeated or persistent infections are correlated with serious complications in humans, understanding these processes may improve clinical management and public health disease control. Methods We conducted whole-genome sequence analysis on C. trachomatis isolates collected from a previously described patient set in which individuals were shown to be infected with a single serovar over a lengthy period. Results Data from 5 of 7 patients showed compelling evidence for the ability of these patients to harbor the same strain for 3-5 years. Mutations in these strains were cumulative, very uncommon, and not linked to any single protein or pathway. Serovar J strains isolated from 1 patient 3 years apart did not accumulate a single base change across the genome. In contrast, the sequence results of 2 patients, each infected only with serovar Ia strains, revealed multiple same-serovar infections over 1-5 years. Conclusions These data demonstrate examples of long-term persistence in patients in the face of repeated antibiotic therapy and show that pathogen mutational strategies are not important in persistence of this pathogen in patients.
ABSTRACT Intracellular bacterial pathogens in the family Chlamydiaceae are causes of human blindness, sexually transmitted disease, and pneumonia. Genetic dissection of the mechanisms of chlamydial pathogenicity has been hindered by multiple limitations, including the inability to inactivate genes that would prevent the production of elementary bodies. Many genes are also Chlamydia-specific genes, and chlamydial genomes have undergone extensive reductive evolution, so functions often cannot be inferred from homologs in other organisms. Conditional mutants have been used to study essential genes of many microorganisms, so we screened a library of 4,184 ethyl methanesulfonate-mutagenized Chlamydia trachomatis isolates for temperature-sensitive (TS) mutants that developed normally at physiological temperature (37°C) but not at nonphysiological temperatures. Heat-sensitive TS mutants were identified at a high frequency, while cold-sensitive mutants were less common. Twelve TS mutants were mapped using a novel markerless recombination approach, PCR, and genome sequencing. TS alleles of genes that play essential roles in other bacteria and chlamydia-specific open reading frames (ORFs) of unknown function were identified. Temperature-shift assays determined that phenotypes of the mutants manifested at distinct points in the developmental cycle. Genome sequencing of a larger population of TS mutants also revealed that the screen had not reached saturation. In summary, we describe the first approach for studying essential chlamydial genes and broadly applicable strategies for genetic mapping in Chlamydia spp. and mutants that both define checkpoints and provide insights into the biology of the chlamydial developmental cycle. IMPORTANCE Study of the pathogenesis of Chlamydia spp. has historically been hampered by a lack of genetic tools. Although there has been recent progress in chlamydial genetics, the existing approaches have limitations for the study of the genes that mediate growth of these organisms in cell culture. We used a genetic screen to identify conditional Chlamydia mutants and then mapped these alleles using a broadly applicable recombination strategy. Phenotypes of the mutants provide fundamental insights into unexplored areas of chlamydial pathogenesis and intracellular biology. Finally, the reagents and approaches we describe are powerful resources for the investigation of these organisms.
—Wikidata is a world readable and writable knowledge base maintained by the Wikimedia Foundation. It offers the opportunity to collaboratively construct a fully open access knowledge graph spanning biology, medicine, and all other domains of knowledge. To meet this potential, social and technical challenges must be overcome most of which are familiar to the biocuration community. These include community ontology building, high precision information extraction, provenance, and license management. By working together with Wikidata now, we can help shape it into a trustworthy, unencumbered central node in the Semantic Web of biomedical data.
The last 20 years of advancement in sequencing technologies have led to sequencing thousands of microbial genomes, creating mountains of genetic data. While efficiency in generating the data improves almost daily, applying meaningful relationships between taxonomic and genetic entities on this scale requires a structured and integrative approach. Currently, knowledge is distributed across a fragmented landscape of resources from government-funded institutions such as National Center for Biotechnology Information (NCBI) and UniProt to topic-focused databases like the ODB3 database of prokaryotic operons, to the supplemental table of a primary publication. A major drawback to large scale, expert-curated databases is the expense of maintaining and extending them over time. No entity apart from a major institution with stable long-term funding can consider this, and their scope is limited considering the magnitude of microbial data being generated daily. Wikidata is an openly editable, semantic web compatible framework for knowledge representation. It is a project of the Wikimedia Foundation and offers knowledge integration capabilities ideally suited to the challenge of representing the exploding body of information about microbial genomics. We are developing a microbial specific data model, based on Wikidata's semantic web compatibility, which represents bacterial species, strains and the gene and gene products that define them. Currently, we have loaded 43,694 gene and 37,966 protein items for 21 species of bacteria, including the human pathogenic bacteriaChlamydia trachomatis.Using this pathogen as an example, we explore complex interactions between the pathogen, its host, associated genes, other microbes, disease and drugs using the Wikidata SPARQL endpoint. In our next phase of development, we will add another 99 bacterial genomes and their gene and gene products, totaling ∼900,000 additional entities. This aggregation of knowledge will be a platform for community-driven collaboration, allowing the networking of microbial genetic data through the sharing of knowledge by both the data and domain expert.
ABSTRACT Chlamydia trachomatis can enter a viable but nonculturable state in vitro termed persistence. A common feature of C. trachomatis persistence models is that reticulate bodies fail to divide and make few infectious progeny until the persistence-inducing stressor is removed. One model of persistence that has relevance to human disease involves tryptophan limitation mediated by the host enzyme indoleamine 2,3-dioxygenase, which converts l -tryptophan to N -formylkynurenine. Genital C. trachomatis strains can counter tryptophan limitation because they encode a tryptophan-synthesizing enzyme. Tryptophan synthase is the only enzyme that has been confirmed to play a role in interferon gamma (IFN-γ)-induced persistence, although profound changes in chlamydial physiology and gene expression occur in the presence of persistence-inducing stressors. Thus, we screened a population of mutagenized C. trachomatis strains for mutants that failed to reactivate from IFN-γ-induced persistence. Six mutants were identified, and the mutations linked to the persistence phenotype in three of these were successfully mapped. One mutant had a missense mutation in tryptophan synthase; however, this mutant behaved differently from previously described synthase null mutants. Two hypothetical genes of unknown function, ctl0225 and ctl0694 , were also identified and may be involved in amino acid transport and DNA damage repair, respectively. Our results indicate that C. trachomatis utilizes functionally diverse genes to mediate survival during and reactivation from persistence in HeLa cells.
Includes Figure S1: Schematic design of MyVariant.info; Figure S2: Histograms of the request time for MyGene.info; Figure S3: JSON annotation object examples; Table S3: Examples of HGVS nomenclature; Supplementary Note 1: IPython notebook for Miller syndrome study. (PDF 946 kb)