Rare diseases collectively affect millions of Americans, but less than 5% have approved treatments, and new drug development remains limited. For such diseases, drug repurposing may be an effective strategy to find new treatment options. In the rare genetic disorder community, drugs are frequently prescribed off-label. This information is rarely available for research, but if captured, could be leveraged to accelerate the identification of candidate drugs to be evaluated for safety and efficacy of the treatment of rare diseases. CURE ID is a publicly available treatment registry that collects real-world treatment data directly from healthcare providers, patients, and care partners in a consistent format. By aggregating this information, CURE ID can generate hypotheses for follow-up targeted research of repurposed drugs, potentially leading to the approval of these drugs for new indications. The success of the platform is predicated on its adoption in the rare disease community and routinely reporting treatment experiences to CURE ID.
Identifying rare disease (RD) patients in electronic health records (EHRs) is difficult, as most of the over 10,000 RDs are not adequately captured by standard coding systems. To address this, we developed a semi-automated workflow to map RDs to SNOMED-CT and ICD-10 codes, enabling improved RD identification across EHR systems. The optimized workflow yielded 88.4% true RD codes in a subset of 1,715 manually curated diseases. Using this workflow and starting with 12,003 GARD IDs mapped to ORPHANET, we obtained 12,081 SNOMED-CT and 357 ICD-10 codes representing 6,342 RDs, organized into 30 ORPHANET linearization classes. We applied these codes to the National COVID Cohort Collaborative (N3C) dataset of over 21 million patients. Among these patients, 8.46 million were identified as COVID-19 positive, of which 4.8 million were used in analyses. Among these, 316,836 (6.55%) had a preexisting RD. Logistic regression, adjusted for age and BMI, revealed that most RD classes were significantly associated with increased odds of severe COVID-19 outcomes. Notably high odds of mortality were observed for rare cardiac (OR = 4.07) and otorhinolaryngologic diseases (OR = 4.00). Hospitalization risk was also elevated across all RD classes, with the highest odds seen in otorhinolaryngologic (OR = 4.31) and endocrine diseases (OR = 3.38). This approach enables scalable RD patient identification in EHRs and highlights the need for tailored healthcare strategies to improve outcomes in RD populations.
Precision medicine aims to enhance diagnosis, treatment, and prognosis by integrating multimodal data at the point of care. However, challenges arise due to the vast number of diseases, differing methods of classification, and conflicting terminological coding systems and practices used to represent molecular definitions of disease. This lack of interoperability artificially constrains the potential for diagnosis, clinical decision support, care outcome analysis, as well as data linkage across research domains to support the development or repurposing of therapeutics. There is a clear and pressing need for a unified system for managing disease entities-including identifiers, synonyms, and definitions. To address these issues, we created the Mondo disease ontology-a community-driven, open-source, unified disease classification system that harmonizes diverse terminologies into a consistent, computable framework. Mondo integrates key medical and biomedical terminologies, including Online Mendelian Inheritance in Man (OMIM), Orphanet, Medical Subject Headings (MeSH), National Cancer Institute Thesaurus (NCIt), and more, to provide a comprehensive and accurate representation of disease concepts with fully provenanced and attributed links back to the sources. Mondo can be used as the handle for curation of gene-disease associations utilized in diagnostic applications, research applications such as computational phenotyping, and in clinical coding systems in clinical decision support by pointing the clinician to the numerous knowledge resources linked to the Mondo identifier. Mondo's community-centric approach, stewarded by the Monarch Initiative's expertise in ontologies, ensures that the ontology remains adaptable to the evolving needs of biomedical research and clinical communities, as well as the knowledge providers.
Background:Over 10,000 rare diseases (RDs) affect more than 300 million people globally, yet their influence on COVID-19 severity, reinfection risk, and long COVID remains poorly understood. This study evaluates the impact of RDs on these outcomes and examines the effectiveness of vaccination and antiviral treatments among individuals with and without RDs. Methods:We conducted a retrospective cohort study using harmonized electronic health records (EHRs) from the National COVID Cohort Collaborative (N3C), encompassing 21,704,702 individuals, including 4,825,605 with confirmed SARS-CoV-2 infection between Jan 1, 2020, and Jan 4, 2024. RDs were defined using 12,003 conditions curated from GARD and Orphanet, mapped to OMOP concepts, and classified into 18 RD classes based on medical specialty involvement. Primary outcomes included: (1) COVID-19 severity (hospitalization and life-threatening disease), (2) long COVID, and (3) SARS-CoV-2 reinfection. We applied multivariable logistic regression with inverse probability of treatment weighting and reported adjusted odds ratios with 95% confidence intervals and associated p-values. Models were controlled for demographics, comorbidities, and exposure to vaccination and antiviral treatments. Findings:Of 21,704,702 individuals, we identify 4,825,605 COVID-19 positive individuals, 6.36% had RDs, with markedly higher rates of rare disease (RD) patients that have life-threatening illness (16% vs. 6.1% without life-threatening illness) and that are hospitalized (13% vs. 6.0% without hospitalization). Otorhinolaryngologic diseases showed the highest risk of life-threatening outcomes (OR 4.51; 95% CI 3.81-5.33), followed by developmental defect during embryogenesis (OR 1.84; 95% CI 1.72-1.98) and cardiac conditions (OR 1.79; 95% CI 1.51-2.11). Hospitalization risk was highest for otorhinolaryngologic (OR 2.90; 95% CI 2.61-3.23), developmental defect during embryogenesis (OR 2.06; 95% CI 1.97-2.16), and hematologic and endocrine diseases (OR 1.81; 95% CI 1.75-1.87 and OR 1.81; 95% CI 1.64-1.99, respectively).In patients with RDs, vaccination alone or antiviral treatment alone was associated with reduced odds of life-threatening COVID-19 disease compared to non-vaccinated individuals (OR 0.71; 95% CI 0.66-0.77 and OR 0.33; 95% CI 0.26-0.42, respectively). The combination of both vaccination and antiviral treatment showed the greatest reduction in odds ratio (OR 0.24; 95% CI 0.20-0.27). Similar results were observed in patients without RDs. In contrast, vaccination or antiviral therapy alone, compared to no intervention, did not significantly reduce long COVID risk in RD patients, although these interventions alone did result in a lower odds ratio in patients without RD. However, their combination was protective in both groups. Vaccination alone, compared to no vaccination, also reduced the risk of reinfection across RD and non-RD populations. Interpretation:RD patients face elevated risks of severe COVID-19 outcomes. While vaccination and antivirals significantly reduce the acute severity of illness, their impact on long COVID appears limited in this population. Notably, vaccination was protective against COVID-19 reinfection in both RD and non-RD populations. These findings highlight the need for targeted strategies to protect RD patients beyond current interventions, particularly in preventing long-term complications.
The mission of the NCATS Division of Rare Diseases Research Innovation (DRDRI), formerly known as the Office of Rare Diseases Research, is to advance rare diseases research to benefit patients. DRDRI is part of the National Center for Advancing Translational Sciences, one of the 27 components of the US National Institutes of Health. DRDRI facilitates and coordinates NIH-wide activities involving rare diseases research, as well as directly supporting rare diseases research activities. These activities include the development and maintenance of a centralized database on rare diseases; collaboration and coordination with organizations focused on orphan products development and rare diseases research across the globe, advising the Office of the NIH Director on matters related to NIH-sponsored research involving rare diseases; and responding to information and policy requests about rare diseases within the NIH. DRDRI also supports various rare diseases research activities, including the Rare Diseases Clinical Research Network, rare disease-related conference grants, and assessment of the costs of untreated rare diseases. In addition, several of the projects DRDRI is supporting are “many diseases at a time” translational approaches for rare diseases, which emphasize leveraging commonalities across multiple rare diseases. These include the support of “basket trials” based on shared molecular etiologies across multiple rare diseases, as well as therapeutic platforms for the treatment of monogenic diseases, such as gene therapy and genome editing. This Perspective will provide an overview and summary of these various activities, noting where relevant our collaborative partnerships within the U.S. and internationally.
Drug development in rare diseases is challenging due to the limited availability of subjects with the diseases and recruiting from a small patient population. The high cost and low success rate of clinical trials motivate deliberate analysis of existing clinical trials to understand status of clinical development of orphan drugs and discover new insight for new trial. In this project, we aim to develop a user centered Rare disease based Clinical Trial Knowledge Graph (RCTKG) to integrate publicly available clinical trial data with rare diseases from the Genetic and Rare Disease (GARD) program in a semantic and standardized form for public use. To better serve and represent the interests of rare disease users, user stories were defined for three types of users, patients, healthcare providers and informaticians, to guide the RCTKG design in supporting the GARD program at NCATS/NIH and the broad clinical/research community in rare diseases.
BACKGROUND:Ontologies are fundamental components of informatics infrastructure in domains such as biomedical, environmental, and food sciences, representing consensus knowledge in an accurate and computable form. However, their construction and maintenance demand substantial resources and necessitate substantial collaboration between domain experts, curators, and ontology experts. We present Dynamic Retrieval Augmented Generation of Ontologies using AI (DRAGON-AI), an ontology generation method employing Large Language Models (LLMs) and Retrieval Augmented Generation (RAG). DRAGON-AI can generate textual and logical ontology components, drawing from existing knowledge in multiple ontologies and unstructured text sources. RESULTS:We assessed performance of DRAGON-AI on de novo term construction across ten diverse ontologies, making use of extensive manual evaluation of results. Our method has high precision for relationship generation, but has slightly lower precision than from logic-based reasoning. Our method is also able to generate definitions deemed acceptable by expert evaluators, but these scored worse than human-authored definitions. Notably, evaluators with the highest level of confidence in a domain were better able to discern flaws in AI-generated definitions. We also demonstrated the ability of DRAGON-AI to incorporate natural language instructions in the form of GitHub issues. CONCLUSIONS:These findings suggest DRAGON-AI's potential to substantially aid the manual ontology construction process. However, our results also underscore the importance of having expert curators and ontology editors drive the ontology generation process.
BACKGROUND:The United Nations recently made a call to address the challenges of an estimated 300 million persons worldwide living with a rare disease through the collection, analysis, and dissemination of disaggregated data. Epidemiologic Information (EI) regarding prevalence and incidence data of rare diseases is sparse and current paradigms of identifying, extracting, and curating EI rely upon time-intensive, error-prone manual processes. With these limitations, a clear understanding of the variation in epidemiology and outcomes for rare disease patients is hampered. This challenges the public health of rare diseases patients through a lack of information necessary to prioritize research, policy decisions, therapeutic development, and health system allocations.METHODS:In this study, we developed a newly curated epidemiology corpus for Named Entity Recognition (NER), a deep learning framework, and a novel rare disease epidemiologic information pipeline named EpiPipeline4RD consisting of a web interface and Restful API. For the corpus creation, we programmatically gathered a representative sample of rare disease epidemiologic abstracts, utilized weakly-supervised machine learning techniques to label the dataset, and manually validated the labeled dataset. For the deep learning framework development, we fine-tuned our dataset and adapted the BioBERT model for NER. We measured the performance of our BioBERT model for epidemiology entity recognition quantitatively with precision, recall, and F1 and qualitatively through a comparison with Orphanet. We demonstrated the ability for our pipeline to gather, identify, and extract epidemiology information from rare disease abstracts through three case studies.RESULTS:We developed a deep learning model to extract EI with overall F1 scores of 0.817 and 0.878, evaluated at the entity-level and token-level respectively, and which achieved comparable qualitative results to Orphanet's collection paradigm. Additionally, case studies of the rare diseases Classic homocystinuria, GRACILE syndrome, Phenylketonuria demonstrated the adequate recall of abstracts with epidemiology information, high precision of epidemiology information extraction through our deep learning model, and the increased efficiency of EpiPipeline4RD compared to a manual curation paradigm.CONCLUSIONS:EpiPipeline4RD demonstrated high performance of EI extraction from rare disease literature to augment manual curation processes. This automated information curation paradigm will not only effectively empower development of the NIH Genetic and Rare Diseases Information Center (GARD), but also support the public health of the rare disease community.
Rare diseases (RDs) are naturally associated with a low prevalence rate, which raises a big challenge due to there being less data available for supporting preclinical and clinical studies. There has been a vast improvement in our understanding of RD, largely owing to advanced big data analytic approaches in genetics/genomics. Consequently, a large volume of RD-related publications has been accumulated in recent years, which offers opportunities to utilize these publications for accessing the full spectrum of the scientific research and supporting further investigation in RD. In this study, we systematically analyzed, semantically annotated, and scientifically categorized RD-related PubMed articles, and integrated those semantic annotations in a knowledge graph (KG), which is hosted in Neo4j based on a predefined data model. With the successful demonstration of scientific contribution in RD via the case studies performed by exploring this KG, we propose to extend the current effort by expanding more RD-related publications and more other types of resources as a next step.
Rare diseases are naturally associated with low prevalence rate, which raises a big challenge due to less data available for supporting preclinical and clinical studies. Therefore, it is critical to fully utilize the accumulated scientific publications in rare diseases over years, in order to access full spectrum of scientific research and enable relevant scientific evidence extraction and generation. In this study, we obtained rare disease related PubMed articles, extracted multiple types of biomedical information, and semantically presented the data in a knowledge graph, which is hosted in Neo4j based on a predefined data model to support further rare disease research.
The National Center for Advancing Translational Science (NCATS) seeks to improve upon the translational process to advance research and treatment across all diseases and conditions and bring these interventions to all who need them. Addressing the racial/ethnic health disparities and health inequities that persist in screening, diagnosis, treatment, and health outcomes (e.g., morbidity, mortality) is central to NCATS' mission to deliver more interventions to all people more quickly. Working toward this goal will require enhancing diversity, equity, inclusion, and accessibility (DEIA) in the translational workforce and in research conducted across the translational continuum, to support health equity. This paper discusses how aspects of DEIA are integral to the mission of translational science (TS). It describes recent NIH and NCATS efforts to advance DEIA in the TS workforce and in the research we support. Additionally, NCATS is developing approaches to apply a lens of DEIA in its activities and research - with relevance to the activities of the TS community - and will elucidate these approaches through related examples of NCATS-led, partnered, and supported activities, working toward the Center's goal of bringing more treatments to all people more quickly.
Background: Limited knowledge and unclear underlying biology of many rare diseases pose significant challenges to patients, clinicians, and scientists. To address these challenges, there is an urgent need to inspire and encourage scientists to propose and pursue innovative research studies that aim to uncover the genetic and molecular causes of more rare diseases and ultimately to identify effective therapeutic solutions. A clear understanding of current research efforts, knowledge/research gaps, and funding patterns as scientific evidence is crucial to systematically accelerate the pace of research discovery in rare diseases, which is an overarching goal of this study. Methods: To semantically represent NIH funding data for rare diseases and advance its use of effectively promoting rare disease research, we identified NIH funded projects for rare diseases by mapping GARD diseases to the project based on project titles; subsequently we presented and managed those identified projects in a knowledge graph using Neo4j software, hosted at NCATS, based on a pre-defined data model that captures semantics among the data. With this developed knowledge graph, we were able to perform several case studies to demonstrate scientific evidence generation for supporting rare disease research discovery. Results: Of 5001 rare diseases belonging to 32 distinct disease categories, we identified 1294 diseases that are mapped to 45,647 distinct, NIH-funded projects obtained from the NIH ExPORTER by implementing semantic annotation of project titles. To capture semantic relationships presenting amongst mapped research funding data, we defined a data model comprised of seven primary classes and corresponding object and data properties. A Neo4j knowledge graph based on this predefined data model has been developed, and we performed multiple case studies over this knowledge graph to demonstrate its use in directing and promoting rare disease research. Conclusion: We developed an integrative knowledge graph with rare disease funding data and demonstrated its use as a source from where we can effectively identify and generate scientific evidence to support rare disease research. With the success of this preliminary study, we plan to implement advanced computational approaches for analyzing more funding related data, e.g., project abstracts and PubMed article abstracts, and linking to other types of biomedical data to perform more sophisticated research gap analysis and identify opportunities for future research in rare diseases.
Rare diseases affect between 25 and 30 million people in the United States, and understanding their epidemiology is critical to focusing research efforts. However, little is known about the prevalence of many rare diseases. Given a lack of automated tools, current methods to identify and collect epidemiological data are managed through manual curation. To accelerate this process systematically, we developed a novel predictive model to programmatically identify epidemiologic studies on rare diseases from PubMed. A long short-term memory recurrent neural network was developed to predict whether a PubMed abstract represents an epidemiologic study. Our model performed well on our validation set (precision = 0.846, recall = 0.937, AUC = 0.967), and obtained satisfying results on the test set. This model thus shows promise to accelerate the pace of epidemiologic data curation in rare diseases and could be extended for use in other types of studies and in other disease domains.
The increasing availability of data in translational science affords an unprecedented opportunity to improve health information access for rare diseases. To that end, we leverage our recent effort in building a comprehensive knowledge graph to provide a holistic view of rare diseases so as to empower patients and their families. We illustrate this holistic view through a preview of our upcoming updates to our rare disease information portal GARD.
BACKGROUND Although many efforts have been made to develop comprehensive disease resources that capture rare disease information for the purpose of clinical decision making and education, there is no standardized protocol for defining and harmonizing rare diseases across multiple resources. This introduces data redundancy and inconsistency that may ultimately increase confusion and difficulty for the wide use of these resources. To overcome such encumbrances, we report our preliminary study to identify phenotypical similarity among genetic and rare diseases (GARD) that are presenting similar clinical manifestations, and support further data harmonization. OBJECTIVE To support rare disease data harmonization, we aim to systematically identify phenotypically similar GARD diseases from a disease-oriented integrative knowledge graph and determine their similarity types. METHODS We identified phenotypically similar GARD diseases programmatically with 2 methods: (1) We measured disease similarity by comparing disease mappings between GARD and other rare disease resources, incorporating manual assessment; 2) we derived clinical manifestations presenting among sibling diseases from disease classifications and prioritized the identified similar diseases based on their phenotypes and genotypes. RESULTS For disease similarity comparison, approximately 87% (341/392) identified, phenotypically similar disease pairs were validated; 80% (271/392) of these disease pairs were accurately identified as phenotypically similar based on similarity score. The evaluation result shows a high precision (94%) and a satisfactory quality (86% F measure). By deriving phenotypical similarity from Monarch Disease Ontology (MONDO) and Orphanet disease classification trees, we identified a total of 360 disease pairs with at least 1 shared clinical phenotype and gene, which were applied for prioritizing clinical relevance. A total of 662 phenotypically similar disease pairs were identified and will be applied for GARD data harmonization. CONCLUSIONS We successfully identified phenotypically similar rare diseases among the GARD diseases via 2 approaches, disease mapping comparison and phenotypical similarity derivation from disease classification systems. The results will not only direct GARD data harmonization in expanding translational science research but will also accelerate data transparency and consistency across different disease resources and terminologies, helping to build a robust and up-to-date knowledge resource on rare diseases.
Objective In this study, we aimed to evaluate the capability of the Unified Medical Language System (UMLS) as one data standard to support data normalization and harmonization of datasets that have been developed for rare diseases. Through analysis of data mappings between multiple rare disease resources and the UMLS, we propose suggested extensions of the UMLS that will enable its adoption as a global standard in rare disease. Methods We analyzed data mappings between the UMLS and existing datasets on over 7,000 rare diseases that were retrieved from four publicly accessible resources: Genetic And Rare Diseases Information Center (GARD), Orphanet, Online Mendelian Inheritance in Men (OMIM), and the Monarch Disease Ontology (MONDO). Two types of disease mappings were assessed, (1) curated mappings extracted from those four resources; and (2) established mappings generated by querying the rare disease-based integrative knowledge graph developed in the previous study. Results We found that 100% of OMIM concepts, and over 50% of concepts from GARD, MONDO, and Orphanet were normalized by the UMLS and accurately categorized into the appropriate UMLS semantic groups. We analyzed 58,636 UMLS mappings, which resulted in 3,876 UMLS concepts across these resources. Manual evaluation of a random set of 500 UMLS mappings demonstrated a high level of accuracy (99%) of developing those mappings, which consisted of 414 mappings of synonyms (82.8%), 76 are subtypes (15.2%), and five are siblings (1%). Conclusion The mapping results illustrated in this study that the UMLS was able to accurately represent rare disease concepts, and their associated information, such as genes and phenotypes, and can effectively be used to support data harmonization across existing resources developed on collecting rare disease data. We recommend the adoption of the UMLS as a data standard for rare disease to enable the existing rare disease datasets to support future applications in a clinical and community settings.
BACKGROUND:The Genetic and Rare Diseases (GARD) Information Center was established by the National Institutes of Health (NIH) to provide freely accessible consumer health information on over 6500 genetic and rare diseases. As the cumulative scientific understanding and underlying evidence for these diseases have expanded over time, existing practices to generate knowledge from these publications and resources have not been able to keep pace. Through determining the applicability of computational approaches to enhance or replace manual curation tasks, we aim to both improve the sustainability and relevance of consumer health information, but also to develop a foundational database, from which translational science researchers may start to unravel disease characteristics that are vital to the research process.RESULTS:We developed a meta-ontology based integrative knowledge graph for rare diseases in Neo4j. This integrative knowledge graph includes a total of 3,819,623 nodes and 84,223,681 relations from 34 different biomedical data resources, including curated drug and rare disease associations. Semi-automatic mappings were generated for 2154 unique FDA orphan designations to 776 unique GARD diseases, and 3322 unique FDA designated drugs to UNII, as well as 180,363 associations between drug and indication from Inxight Drugs, which were integrated into the knowledge graph. We conducted four case studies to demonstrate the capabilities of this integrative knowledge graph in accelerating the curation of scientific understanding on rare diseases through the generation of disease mappings/profiles and pathogenesis associations.CONCLUSIONS:By integrating well-established database resources, we developed an integrative knowledge graph containing a large volume of biomedical and research data. Demonstration of several immediate use cases and limitations of this process reveal both the potential feasibility and barriers of utilizing graph-based resources and approaches to support their use by providers of consumer health information, such as GARD, that may struggle with the needs of maintaining knowledge reliant on an evolving and growing evidence-base. Finally, the successful integration of these datasets into a freely accessible knowledge graph highlights an opportunity to take a translational science view on the field of rare diseases by enabling researchers to identify disease characteristics, which may play a role in the translation of discover across different research domains.
This special communication describes activities, products, and lessons learned from a recent hackathon that was funded by the National Center for Advancing Translational Sciences via the Biomedical Data Translator program ('Translator'). Specifically, Translator team members self-organized and worked together to conceptualize and execute, over a five-day period, a multi-institutional clinical research study that aimed to examine, using open clinical data sources, relationships between sex, obesity, diabetes, and exposure to airborne fine particulate matter among patients with severe asthma. The goal was to develop a proof of concept that this new model of collaboration and data sharing could effectively produce meaningful scientific results and generate new scientific hypotheses. Three Translator Clinical Knowledge Sources, each of which provides open access (via Application Programming Interfaces) to data derived from the electronic health record systems of major academic institutions, served as the source of study data. Jupyter Python notebooks, shared in GitHub repositories, were used to call the knowledge sources and analyze and integrate the results. The results replicated established or suspected relationships between sex, obesity, diabetes, exposure to airborne fine particulate matter, and severe asthma. In addition, the results demonstrated specific differences across the three Translator Clinical Knowledge Sources, suggesting cohort- and/or environment-specific factors related to the services themselves or the catchment area from which each service derives patient data. Collectively, this special communication demonstrates the power and utility of intense, team-oriented hackathons and offers general technical, organizational, and scientific lessons learned.
Display Omitted • The Biomedical Data Translator Program was launched in October 2016. • The Biomedical Data Translator Consortium comprises 11 teams and ~200 team members. • Regular in-person hackathons have proven effective in promoting team science. • We describe a hackathon activity focused on open Translator clinical data sources. • Our ‘lessons learned’ have broad applicability across scientific domains. This special communication describes activities, products, and lessons learned from a recent hackathon that was funded by the National Center for Advancing Translational Sciences via the Biomedical Data Translator program (‘Translator’). Specifically, Translator team members self-organized and worked together to conceptualize and execute, over a five-day period, a multi-institutional clinical research study that aimed to examine, using open clinical data sources, relationships between sex, obesity, diabetes, and exposure to airborne fine particulate matter among patients with severe asthma. The goal was to develop a proof of concept that this new model of collaboration and data sharing could effectively produce meaningful scientific results and generate new scientific hypotheses. Three Translator Clinical Knowledge Sources, each of which provides open access (via Application Programming Interfaces) to data derived from the electronic health record systems of major academic institutions, served as the source of study data. Jupyter Python notebooks, shared in GitHub repositories, were used to call the knowledge sources and analyze and integrate the results. The results replicated established or suspected relationships between sex, obesity, diabetes, exposure to airborne fine particulate matter, and severe asthma. In addition, the results demonstrated specific differences across the three Translator Clinical Knowledge Sources, suggesting cohort- and/or environment-specific factors related to the services themselves or the catchment area from which each service derives patient data. Collectively, this special communication demonstrates the power and utility of intense, team-oriented hackathons and offers general technical, organizational, and scientific lessons learned.