Clinical genetic laboratories must have access to clinically validated biomedical data for precision medicine. A lack of accessibility, normalized structure, and consistency in evaluation complicates interpretation of disease causality, resulting in confusion in assessing the clinical validity of genes and genetic variants for diagnosis. A key goal of the Clinical Genome Resource (ClinGen) is to fill the knowledge gap concerning the strength of evidence supporting the role of a gene in a monogenic disease, which is achieved through a process known as Gene-Disease Validity curation. Here we review the work of ClinGen in developing a curation infrastructure that supports the standardization, harmonization, and dissemination of Gene-Disease Validity data through the creation of frameworks and the utilization of common data standards. This infrastructure is based on several applications, including the ClinGen GeneTracker, Gene Curation Interface, Data Exchange, GeneGraph, and website.
The Clinical Genome Resource (ClinGen) serves as an authoritative resource on the clinical relevance of genes and variants. In order to support our curation activities and to disseminate our findings to the community, we have developed a Data Platform of informatics resources backed by standardized data models. In this workshop we demonstrate our publicly available resources including curation interfaces, (Variant Curation Interface, CIViC), supporting infrastructure (Allele Registry, Genegraph), and data models (SEPIO, GA4GH VRS, VA).
Methods: In total, 76 longitudinal cfDNA samples from 21 patients with advanced NSCLC receiving first-line treatment were sequenced at 10x coverage.In parallel, we performed deep-targeted sequencing (approx.1250x coverage) for 47 matched time points.By applying genome-wide fragmentation-based statistics, a trained model was used to determine tumor fractions at baseline and their dynamic changes throughout treatment.Results: Histological subtypes were: adenocarcinoma (n = 11), squamous cell carcinoma (n = 5), and others (n = 5).15/21 (71%) patients were treated with immune checkpoint inhibitors alone or in combination with chemotherapy.Genome-wide fragmentation profiles in nonoverlapping regions showed marked heterogeneity at baseline when compared to time points associated with stable disease or partial response.We observed a strong correlation between tumor fractions determined by DELFI-TF and MAF (Pearson, r = 0.92, p < 0.001).Longitudinal DELFI-TF dynamics overlapped MAF changes and were associated with treatment response as assessed by conventional CT scans.Conclusion: DELFI-TF accurately tracked tumor fractions in patients with advanced NSCLC receiving immune checkpoint blockade.Our results show the potential of using an inexpensive and reliable method for monitoring treatment responses independent of a mutation-based approach.
As the diversity of genomic variation data increases with our growing understanding of the role of variation in health and disease, it is critical to develop standards for precise inter-system exchange of these data for research and clinical applications. The Global Alliance for Genomics and Health (GA4GH) Variation Representation Specification (VRS) meets this need through a technical terminology and information model for disambiguating and concisely representing variation concepts. Here we discuss the recent Genotype model in VRS, which may be used to represent the allelic composition of a genetic locus. We demonstrate the use of the Genotype model and the constituent Haplotype model for the precise and interoperable representation of pharmacogenomic diplotypes, HGVS variants, and VCF records using VRS and discuss how this can be leveraged to enable interoperable exchange and search operations between assayed variation and genomic knowledgebases.
There are thousands of distinct disease entities and concepts, each of which are known by different and sometimes contradictory names. The Monarch Initiative aims to integrate genotype, phenotype, and disease knowledge from a large variety of sources in support of improved diagnostics and mechanism discovery through various algorithms and tools. However, the lack of a unified system for managing disease entities poses a major challenge for both machines and humans to predict causes and treatments for disease. The multitude of disease resources have not been well coordinated nor computationally integrated. Furthermore, the classification of phenotypes and their association with diseases is another source of disagreement across sources. The Human Phenotype Ontology has helped to standardize phenotypic features across knowledge sources, but there was no equivalent computationally-harmonized disease ontology. To address these problems, a community of disease resources worked together to create the Mondo Disease Ontology as an open, community-driven ontology that integrates key medical and biomedical terminologies and is iteratively and regularly updated via manual curation and through synchronization with external sources using a Bayesian algorithm. Mondo supports disease data integration to improve diagnosis, treatment, and translational research. It records the sources of all data and is continually updated, making it suitable for research and clinical applications that require up-to-date disease knowledge. Evidence before this study Many disease terminologies currently exist, but there is not a definitive standard for encoding diseases while addressing requirements for information exchange. Existing sources of disease definitions include the National Cancer Institute Thesaurus (NCIt), the Online Mendelian Inheritance in Man (OMIM), Orphanet, SNOMED CT, Disease Ontology (DO), ICD-10, MedGen, and numerous others. Each of these is designed for a particular purpose, and as such has different strengths. However, these standards only partially overlap and often conflict in the classification or mapping approach, making it difficult to align them with each other and/or with other knowledge sources. This need to integrate information has resulted in a proliferation of mappings between disease entries in different resources; these mappings lack completeness, accuracy, and precision, and are often inconsistent between resources. Added value of this study In order to computationally leverage the available knowledge sources for diagnostics and to reveal underlying mechanisms of diseases, we need to understand which terms are meaningfully equivalent across different resources. This will allow integration of associated information, such as treatments, genetics, phenotypes, etc. We therefore created the Mondo Disease Ontology to provide a logic-based structure for unifying multiple disease resources. Implications of all the available evidence Mondo can be leveraged by researchers and clinicians for disease annotations and data integration to aid in clinical diagnosis, treatment and advancement of human health care. Mondo is a freely available, open terminology that contains over 20,000 disease classes. Mondo is iteratively developed with contributions from the intended community and is under continuous revision, with future plans to further revise the top-level classes. Recently, efforts to classify rare diseases have centered on retrieving terms from various sources to provide a unified resource. Mondo can be explored using any of a variety of ontology browsers such as the Ontology Lookup Service (OLS) (ebi.ac.uk/ols/ontologies/mondo), and the ontology files and current releases are available on GitHub ([github.com/monarch-initiative/mondo][1]). ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement Mondo is generously supported by the NIH National Human Genome Research Institute Phenomics First Resource, NIH-NHGRI # 1 RM1 HG010860-01, a Center of Excellence in Genomic Science; and an NIH Office of the Director Grant #5R24OD011883 for the Monarch Initiative. Additional support for this research/work was supported in part by the National Center for Biotechnology Information of the National Library of Medicine (NLM), National Institutes of Health. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable. Yes All data produced are available online at <https://github.com/monarch-initiative/mondo>. <https://github.com/monarch-initiative/mondo> [1]: http://github.com/monarch-initiative/mondo
Precision oncology is the practice of interpreting the clinical significance of observed molecular changes in patient neoplasms, potentially impacting medical decision making and care. This process is labor-intensive and (among other challenges) involves accurately translating between variation representation conventions from one resource to the next. For example, differences in representations of Copy Number Variation (CNV) from genomic regions, cytogenomic bands, or gene features create challenges in knowledge matching due to lack of standards covering all of these modalities of observed variation.The Global Alliance for Genomics and Health (GA4GH; ga4gh.org) is an international collaborative of genomic data sharing initiatives (Driver Projects) developing genomic data sharing standards within a human rights framework. GA4GH recently published the Variation Representation Specification (VRS; pronounced “verse”), a standard for the computational representation of biomolecular variation. VRS is a terminology, schema, and associated conventions for creating uniquely identifiable and federatable representations of molecular variation. VRS has formal data classes well-suited to differentiating between variation on a single molecule (e.g. tandem duplications) from variation measured at a systemic level (e.g. genome-wide copy number variation). In addition to molecular sequence variation, VRS also supports variation on cytogenetic coordinate systems and genes, making it well-suited to representing variation associated with cancer biomarkers.We demonstrate the use of VRS to model reported gene-associated CNVs from the AACR Project GENIE cohort, to aid in the computational discovery of evidence from clinico-genomic knowledgebases with genomic or cytogenomic CNV representations. We highlight the use case of knowledge matching to the Atlas of Genetics and Cytogenetics in Oncology and Haematology (“the Atlas”; atlasgeneticsoncology.org), a cytogenetics resource historically driven by user website navigation. Using VRS search tools we developed for the Variant Interpretation for Cancer Consortium (VICC; cancervariants.org) GA4GH Driver Project, we found that 64% of GENIE samples with reported CNVs matched clinically relevant knowledge in the Atlas. This work was enabled by programmatic search tools leveraging standard VRS object structures, demonstrating how VRS enables collection of real-world evidence across more resources without manual interpretation or custom normalization methods. We conclude with a survey of open-source tools supporting this analysis as well as search of other clinico-genomic knowledgebases with VRS, including CIViC (civicdb.org), BRCA Exchange (brcaexchange.org), and the Molecular Oncology Almanac (moalmanac.org). Citation Format: Matthew Cannon, Kori Kuzma, James Stevenson, Jiachen Liu, Colin O'Sullivan, Bimal P. Chaudhari, Matthew Brush, Robert R. Freimuth, Tristan Nelson, Michael Baudis, Obi L. Griffith, Malachi Griffith, Lawrence Babb, Melissa S. Cline, Xuelu Liu, Brian Walsh, Alex H. Wagner. Introduction of the GA4GH Variation Representation Specification (VRS) and supporting tools for discovery and exchange of clinical genomic and cytogenomic knowledge in cancers [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2022; 2022 Apr 8-13. Philadelphia (PA): AACR; Cancer Res 2022;82(12_Suppl):Abstract nr 1177.
Maximizing the personal, public, research, and clinical value of genomic information will require the reliable exchange of genetic variation data. We report here the Variation Representation Specification (VRS, pronounced "verse"), an extensible framework for the computable representation of variation that complements contemporary human-readable and flat file standards for genomic variation representation. VRS provides semantically precise representations of variation and leverages this design to enable federated identification of biomolecular variation with globally consistent and unique computed identifiers. The VRS framework includes a terminology and information model, machine-readable schema, data sharing conventions, and a reference implementation, each of which is intended to be broadly useful and freely available for community use. VRS was developed by a partnership among national information resource providers, public initiatives, and diagnostic testing laboratories under the auspices of the Global Alliance for Genomics and Health (GA4GH).
The Global Alliance for Genomics and Health (GA4GH) aims to accelerate biomedical advances by enabling the responsible sharing of clinical and genomic data through both harmonized data aggregation and federated approaches. The decreasing cost of genomic sequencing (along with other genome-wide molecular assays) and increasing evidence of its clinical utility will soon drive the generation of sequence data from tens of millions of humans, with increasing levels of diversity. In this perspective, we present the GA4GH strategies for addressing the major challenges of this data revolution. We describe the GA4GH organization, which is fueled by the development efforts of eight Work Streams and informed by the needs of 24 Driver Projects and other key stakeholders. We present the GA4GH suite of secure, interoperable technical standards and policy frameworks and review the current status of standards, their relevance to key domains of research and clinical care, and future plans of GA4GH. Broad international participation in building, adopting, and deploying GA4GH standards and frameworks will catalyze an unprecedented effort in data sharing that will be critical to advancing genomic medicine and ensuring that all populations can access its benefits.
Conflict resolution in genomic variant interpretation is a critical step toward improving patient care. Evaluating interpretation discrepancies in copy number variants (CNVs) typically involves assessing overlapping genomic content with focus on genes/regions that may be subject to dosage sensitivity (haploinsufficiency (HI) and/or triplosensitivity (TS)). CNVs containing dosage sensitive genes/regions are generally interpreted as "likely pathogenic" (LP) or "pathogenic" (P), and CNVs involving the same known dosage sensitive gene(s) should receive the same clinical interpretation. We compared the Clinical Genome Resource (ClinGen) Dosage Map, a publicly available resource documenting known HI and TS genes/regions, against germline, clinical CNV interpretations within the ClinVar database. We identified 251 CNVs overlapping known dosage sensitive genes/regions but not classified as LP or P; these were sent back to their original submitting laboratories for re-evaluation. Of 246 CNVs re-evaluated, an updated clinical classification was warranted in 157 cases (63.8%); no change was made to the current classification in 79 cases (32.1%); and 10 cases (4.1%) resulted in other types of updates to ClinVar records. This effort will add curated interpretation data into the public domain and allow laboratories to focus attention on more complex discrepancies.
Effective exchange of information about genetic variants is currently hampered by the lack of readily available globally unique variant identifiers that would enable aggregation of information from different sources. The ClinGen Allele Registry addresses this problem by providing (1) globally unique “canonical” variant identifiers (CAids) on demand, either individually or in large batches; (2) access to variant-identifying information in a searchable Registry; (3) links to allele-related records in many commonly used databases; and (4) services for adding links to information about registered variants in external sources. A core element of the Registry is a canonicalization service, implemented using in-memory sequence alignment-based index, which groups variant identifiers denoting the same nucleotide variant and assigns unique and dereferenceable CAids. More than 650 million distinct variants are currently registered, including those from gnomAD, ExAC, dbSNP, and ClinVar, including a small number of variants registered by Registry users. The Registry is accessible both via a web interface and programmatically via well-documented Hypertext Transfer Protocol (HTTP) Representational State Transfer Application Programming Interface (REST-APIs). For programmatic interoperability, the Registry content is accessible in the JavaScript Object Notation for Linked Data (JSON-LD) format. We present several use cases and demonstrate how the linked information may provide raw material for reasoning about variant's pathogenicity.
BACKGROUND:The success of the clinical use of sequencing based tests (from single gene to genomes) depends on the accuracy and consistency of variant interpretation. Aiming to improve the interpretation process through practice guidelines, the American College of Medical Genetics and Genomics (ACMG) and the Association for Molecular Pathology (AMP) have published standards and guidelines for the interpretation of sequence variants. However, manual application of the guidelines is tedious and prone to human error. Web-based tools and software systems may not only address this problem but also document reasoning and supporting evidence, thus enabling transparency of evidence-based reasoning and resolution of discordant interpretations.RESULTS:In this report, we describe the design, implementation, and initial testing of the Clinical Genome Resource (ClinGen) Pathogenicity Calculator, a configurable system and web service for the assessment of pathogenicity of Mendelian germline sequence variants. The system allows users to enter the applicable ACMG/AMP-style evidence tags for a specific allele with links to supporting data for each tag and generate guideline-based pathogenicity assessment for the allele. Through automation and comprehensive documentation of evidence codes, the system facilitates more accurate application of the ACMG/AMP guidelines, improves standardization in variant classification, and facilitates collaborative resolution of discordances. The rules of reasoning are configurable with gene-specific or disease-specific guideline variations (e.g. cardiomyopathy-specific frequency thresholds and functional assays). The software is modular, equipped with robust application program interfaces (APIs), and available under a free open source license and as a cloud-hosted web service, thus facilitating both stand-alone use and integration with existing variant curation and interpretation systems. The Pathogenicity Calculator is accessible at http://calculator.clinicalgenome.org .CONCLUSIONS:By enabling evidence-based reasoning about the pathogenicity of genetic variants and by documenting supporting evidence, the Calculator contributes toward the creation of a knowledge commons and more accurate interpretation of sequence variants in research and clinical care.
Background: The Clinical Genome Resource (ClinGen) Electronic Health Record (EHR) Workgroup aims to integrate ClinGen resources with EHRs. A promising option to enable this integration is through the Health Level Seven (HL7) Infobutton Standard. EHR systems that are certified according to the US Meaningful Use program provide HL7-compliant infobutton capabilities, which can be leveraged to support clinical decision-making in genomics.Objectives: To integrate genomic knowledge resources using the HL7 infobutton standard. Two tactics to achieve this objective were: (1) creating an HL7-compliant search interface for ClinGen, and (2) proposing guidance for genomic resources on achieving HL7 Infobutton standard accessibility and compliance.Methods: We built a search interface utilizing OpenInfobutton, an open source reference implementation of the HL7 Infobutton standard. ClinGen resources were assessed for readiness towards HL7 compliance. Finally, based upon our experiences we provide recommendations for publishers seeking to achieve HL7 compliance.Results: Eight genomic resources and two sub-resources were integrated with the ClinGen search engine via OpenInfobutton and the HL7 infobutton standard. Resources we assessed have varying levels of readiness towards HL7-compliance. Furthermore, we found that adoption of standard terminologies used by EHR systems is the main gap to achieve compliance.Conclusion: Genomic resources can be integrated with EHR systems via the HL7 Infobutton standard using OpenInfobutton. Full compliance of genomic resources with the Infobutton standard would further enhance interoperability with EHR systems.