PURPOSE:Adams-Oliver syndrome (AOS) is a genetically heterogeneous disorder with cardinal features of aplasia cutis congenita and terminal limb reduction defects. A minority of individuals with AOS develop potentially lethal pulmonary hypertension (PH) in infancy, a subgroup that has been refractory to genetic explanation. METHODS:We studied a cohort of individuals with AOS and no genetic diagnosis by genome and exome sequencing. We characterized rare, identified substitution variants in valosin-containing protein (VCP) in vitro using ATP hydrolysis, cryogenic-electron microscopy, thermal stability, and response to CB-5083, a VCP inhibitor. RESULTS:We report a new genetic etiology for AOS in 6 families with PH and 1 family without it. We show that AOS-related VCP variants are hypermorphic with respect to ATP hydrolysis and cause N-terminal domain hyperflexibility with impairment of interdomain coupling. Additionally, we find that CB-5083 inhibits the overactive ATP hydrolysis. Review of published cases of AOS with PH suggests that pulmonary vein stenosis is the most common mechanism. Clinical risk factors for PH in AOS include cutis marmorata telangiectatica congenita, prominent dilated subcutaneous veins and intrauterine growth restriction. CONCLUSION:We identify the prevalent genetic cause of pulmonary hypertension in AOS and highlight a potential therapeutic approach.
Background. Multidomain therapy has been proposed as a standard of care for Alzheimer's disease (AD). AD results from the interplay of multiple interacting dysfunctional biological systems. These systems can be categorized by domain, such as inflammation, cardiovascular health, proteostasis, or metabolism. Specific causes of AD differ between individuals, but each individual is likely to have causes stemming from multiple domains. Objectives. We sought to enumerate prospective randomized controlled trials (RCTs) for multidomain interventions for AD, and to statistically describe their inclusion criteria, trial design parameters (length, number of participants), and outcome measures. We sought to clarify gaps and opportunities in the research. Eligibility criteria. We include all cohort studies and RCTs for multidomain (also known as multimodal, multicomponent, multidimensional, or multisystem) therapy of any stage of AD. Results. There have been 22 studies (completed or reported as ongoing) of multidomain interventions for AD, including 18 RCTs. Of the 14 completed RCTs, 11 demonstrate benefit from their intervention in at least one arm. Conclusions. Multidomain therapy should be the standard of care for AD. Multidomain interventions (also known as treatments) should be employed widely, early, and first-line. Treatment or prevention is likely to be most effect at early, presymptomatic stages, but is worthwhile at all stages of disease. In order to influence multiple domains, multiple modes of therapy are likely necessary in all patients. Some individual modes, such as particular lifestyle interventions, may target multiple domains. Nevertheless, most patients will benefit from multiple modes of intervention (multimodal intervention) that together target multiple domains. Standard-of-care guidelines should explicitly include multidomain interventions. Future clinical trials must be designed to iteratively improve multidomain therapies. Payors should embrace reimbursement for effective multidomain intervention, including personalized coaching. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This work was supported by the Alzheimer's Translational Pillar of Providence St. Joseph Health. No authors or their institutions at any time received payment or services from a third party for any aspect of the submitted work. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes No new data were generated for this manuscript.
Background: Medical and lifestyle management are crucial for Alzheimer's disease (AD). Cerebral blood flow (CBF), vital for brain health, and influenced by modifiable risk factors, is reduced in AD and may become uncoupled from metabolism due to neurovascular dysfunction in later stages. Objective: This mid-trial analysis tested the hypothesis that a coached, multi-modal intervention (PREVENTION) improved ASL-MRI-measured CBF and diabetic risk (QUICKI) in patients with early AD. Methods: The control arm received recommendations and medical management for one year; the active arm additionally received coaching, exercise training, and supplementation. We hypothesized that those in (1) the active arm and (2) with higher intervention adherence would have improved post-trial QUICKI and CBF, particularly in regions relevant to exercise, cardiovascular, diabetic, and AD risk. Post-trial CBF was analyzed using a linear model including arm, baseline CBF, adherence, age, education, and depressive symptoms. Change in QUICKI was analyzed using mixed effects general linear models, including arm, adherence, time, and interactions between time and treatment group and time and adherence, controlling for age. Results: The active arm (n = 18) showed greater post-trial CBF in regions related to exercise, cardiovascular, diabetic, and AD risk, compared to control (n = 20), but did not differ in global CBF, QUICKI, or adherence. Higher adherence scores were associated with greater regional post-trial CBF and improvement in QUICKI, but not global CBF. Conclusions: In this small sample, we found evidence that a multi-modal intervention focused on medical management, exercise, and a carbohydrate-restricted diet improved diabetic risk and CBF in patients with AD.
Rare diseases affect approximately 300 million people worldwide, and over 95
Alzheimer’s disease (AD) leading to cognitive decline and dementia results from the interplay of multiple interacting dysfunctional biological systems. These systems can be categorized by domain, such as inflammation, cardiovascular health, proteostasis, or metabolism. Specific causes of AD differ between individuals, but each individual is likely to have causes stemming from multiple domains. Personalized multidomain therapy has been proposed as a standard of care for AD. We sought to enumerate and describe prospective randomized controlled trials (RCTs) for multidomain interventions for AD, and to extract their inclusion criteria, trial design parameters (length, number of participants), and outcome measures. We sought to clarify gaps and opportunities in research and clinical translation. We conducted a scoping review using the standardized PRISMA-ScR methodological framework. We include all cohort studies and RCTs for multidomain (also known as multimodal, multicomponent, multidimensional, or multisystem) therapy of any stage of AD, published for all dates through July 28, 2025. There have been 23 studies (completed or reported as ongoing) of multidomain interventions for AD, including 19 RCTs. Of the 15 completed RCTs, 12 demonstrate benefit from their intervention in at least one arm. Although these RCTs differ widely in their parameters, the majority support the use of multidomain therapy, and show effect sizes greater than reported for unimodal therapies, including pharmaceuticals. Multidomain therapy should be the standard of care for AD. Multidomain interventions (also known as treatments) should be employed widely, early, and first-line. Treatment or prevention is likely to be most effective at early, presymptomatic stages, but is worthwhile at all stages of disease. In order to influence multiple domains, multiple modes of therapy are likely necessary in all patients. Some individual modes, such as particular lifestyle interventions, may target multiple domains. Nevertheless, most patients will benefit from multiple modes of intervention (multimodal intervention) that together target multiple domains. Standard-of-care guidelines should explicitly include multidomain interventions. Future clinical trials must be designed to iteratively improve multidomain therapies. Payors should embrace reimbursement for effective multidomain intervention, including personalized coaching.
The growing availability of biomedical data offers vast potential to improve human health, but the complexity and lack of integration of these datasets often limit their utility. To address this, the Biomedical Data Translator Consortium has developed an open-source knowledge graph-based system-Translator-designed to integrate, harmonize, and make inferences over diverse biomedical data sources. We announce here Translator's initial public release and provide an overview of its architecture, standards, user interface, and core features. Translator employs a scalable, federated, knowledge graph framework for the integration of clinical, genomic, pharmacological, and other biomedical knowledge sources, enabling query retrieval, inference, and hypothesis generation. Translator's user interface is designed to support the exploration of knowledge relationships and the generation of insights, without requiring deep technical expertise and gradually revealing more detailed evidence, provenance, and confidence information, as needed by a given user. To demonstrate Translator's application and impact, we highlight features of the user interface in the context of three real-world use cases: suggesting potential therapeutics for patients with rare disease; explaining the mechanism of action of a pipeline drug; and screening and validating drug candidates in a model organism. We discuss strengths and limitations of reasoning within a largely federated system and the need for rich concept modeling and deep provenance tracking. Finally, we outline future directions for enhancing Translator's functionality and expanding its data sources. Translator represents a significant step forward in making complex biomedical knowledge more accessible and actionable, aiming to accelerate translational research and improve patient care.
The Coaching for Cognition in Alzheimer’s (COCOA) Trial was a prospective RCT testing a remotely coached multimodal lifestyle intervention for participants early on the Alzheimer’s disease spectrum. Intervention focused on diet, exercise, cognitive training, sleep, stress, and social engagement. Enrollment criteria targeted individuals with cognitive decline who were able to engage remotely with a professional coach. COCOA demonstrated cognitive and functional benefits. Dense omics data were collected on 53 individuals (≥ 58 years). We sought to identify blood analytes that mediated the effects of specific elements of the multimodal intervention on specific outcomes. Outcomes were assessed with the MCI Screen (MCIS) and the Functional Assessment Staging Tool (FAST). We combined these and other measures with proteomics and metabolomics data. We analyzed the resulting dataset of over 300,000 distinct molecular data points—reflecting over 1400 measures— assayed over a period of two years. We used MEGENA to hierarchically multiscale cluster analytes based on correlated responses and identified individual metabolites and functional clusters associated with each intervention and outcome. We analyzed individual time courses of key analyte mediators to illustrate personalized effects of interventions and individualized functional and cognitive outcomes. Distinct sets of correlated serum analytes (“communities”) convey effects to functional (FAST) outcome and to cognitive (MCIS) outcome. Distinct communities respond to different modalities of intervention. Participants followed different aspects of the multimodal recommendations to different extents, and the analytes in their blood also responded idiosyncratically; analyte trajectories in different individuals show distinct dynamics. We made personalized predictions of future inflections in outcome based on observed changes in key serum mediators. We validated results with data from the Precision Recommendations for Environmental Variables, Exercise, Nutrition and Training Interventions to Optimize Neurocognition (PREVENTION) Trial. Lifestyle interventions have profound effects on blood metabolites ( Figure 1 ). These in turn convey subtler specific effects to cognition and broad-based effects to function. Pathways that ameliorate the impact of AD via lifestyle interventions in some individuals include nitrogen subsystems, kidney function, and mitochondrial metabolism. These highlight the importance of clinical attention to overall health spanning multiple organ systems in individuals across the Alzheimer’s disease spectrum.
Medical management and lifestyle are potentially crucial interventions for Alzheimer’s disease (AD). In this study, we present the relationship between adherence of a personalized multi-modal intervention for AD on change in cerebral blood flow (CBF) after 12-months. The PREVENTION study is an ongoing randomized clinical trial (McEwen, 2001). Thirty-three participants with biomarker evidence of amyloidosis had completed the study at the time of the analysis ( Table 1 ). While both arms received personalized multi-modal lifestyle recommendations and four medical visits, the active arm also received dietary counseling, group physical and cognitive exercise, health coaching, and nutritional supplements free of charge. We examined the effects of the 1) the intervention and 2) adherence on CBF. We hypothesized that 1) the active arm and 2) higher intervention adherence would have improved CBF in regions related to level of physical activity (Kleinloog, 2019; Chapman, 2013) and those pertinent to AD. CBF was assessed using arterial spin labeling (ASL). Adherence was measured using the clinician rating scale (CRS), which uses a scale of 1-7 (Kemp, 1998). Participants were divided into two groups based on a cutoff of 5 (passive acceptance). One participant was excluded from this analysis due to missing CRS data. Effects were assessed using a two-tailed t-test. Treatment arms did not differ in any demographic measures at baseline or CRS. Preliminary findings indicate that regional blood flow declined over one year across the whole sample ( Table 2 ). However, individuals with higher adherence experienced increased blood flow in the fusiform gyrus and less blood flow reduction, compared to those with lower adherence, in the anterior cingulate and hippocampus. Findings were borderline significant in the fusiform and anterior cingulate, but not the hippocampus. In this small sample, we found evidence that higher adherence increased or attenuated decline in CBF in regions impacted by physical activity, one modality of the PREVENTION intervention. We did not see an effect in the hippocampus, possibly due to small sample size. We did not find an effect of treatment arm, potentially because both receive recommendations and medical management, and did not differ in adherence.
As large clinical and multiomics datasets and knowledge resources accumulate, they need to be transformed into computable and actionable information to support automated reasoning. These datasets range from laboratory experiment results to electronic health records (EHRs). Barriers to accessibility and sharing of such datasets include diversity of content, size and privacy. Effective transformation of data into information requires harmonization of stakeholder goals, implementation, enforcement of standards regarding quality and completeness, and availability of resources for maintenance and updates. Systems such as the Biomedical Data Translator leverage knowledge graphs (KGs), structured and machine learning readable knowledge representation, to encode knowledge extracted through inference. We focus here on the transformation of data from multiomics datasets and EHRs into compact knowledge, represented in a KG data structure. We demonstrate this data transformation in the context of the Translator ecosystem, including clinical trials, drug approvals, cancer, wellness, and EHR data. These transformations preserve individual privacy. We provide access to the five resulting KGs through the Translator framework. We show examples of biomedical research questions supported by our KGs, and discuss issues arising from extracting biomedical knowledge from multiomics data.
Background: A carbohydrate-restricted diet aimed at lowering insulin levels has the potential to slow Alzheimer’s disease (AD). Restricting carbohydrate consumption reduces insulin resistance, which could improve glucose uptake and neural health. A hallmark feature of AD is widespread cortical thinning; however, no study has demonstrated that lower net carbohydrate (nCHO) intake is linked to attenuated cortical atrophy in patients with AD and confirmed amyloidosis. Objective: We tested the hypothesis that individuals with AD and confirmed amyloid burden eating a carbohydrate-restricted diet have thicker cortex than those eating a moderate-to-high carbohydrate diet. Methods: A total of 31 patients (mean age 71.4±7.0 years) with AD and confirmed amyloid burden were divided into two groups based on a 130 g/day nCHO cutoff. Cortical thickness was estimated from T1-weighted MRI using FreeSurfer. Cortical surface analyses were corrected for multiple comparisons using cluster-wise probability. We assessed group differences using a two-tailed two-independent sample t-test. Linear regression analyses using nCHO as a continuous variable, accounting for confounders, were also conducted. Results: The lower nCHO group had significantly thicker cortex within somatomotor and visual networks. Linear regression analysis revealed that lower nCHO intake levels had a significant association with cortical thickness within the frontoparietal, cingulo-opercular, and visual networks. Conclusions: Restricting carbohydrates may be associated with reduced atrophy in patients with AD. Lowering nCHO to under 130 g/day would allow patients to follow the well-validated MIND diet while benefiting from lower insulin levels.
Huntington’s disease (HD) is a monogenic disorder that is caused by a CAG repeat expansion in the HTT gene. However, beyond the CAG repeat size other genes also contribute to variations in neurodegeneration of the cortex and striatum as well as the timing of disease onset1,2. The standard method to find genetic modifiers of HD has been the use of genome-wide association studies (GWAS) of large numbers of unrelated patients1,3-5. Previous efforts in this vein have identified single nucleotide variants (SNVs) significantly associated with pathways involved in DNA damage and handling that modify HD age of onset (AO)1,3-8. However, many of these associations have small effect sizes, and typically it is not known whether the SNVs identified with GWAS are the basis for the modifying effect. Here, to augment modifier GWAS, we set out to identify variants that may modify AO in HD by performing family-based studies. We performed whole genome sequencing in families with HD in which individuals with similar CAG expansions showed variation in AO (ranging from a 3- to 20-year difference). We examined the segregation of every variant in the genome and associated the occurrence of those variants with AO. Focusing on rare and uncommon variants, we used a priori knowledge to examine the proximity of our top variants to previously reported GWAS loci. Further, we developed an HD impact scoring system to rank each variant and highlight those most likely to be impactful in the context of influencing the pathology associated with the CAG repeat expansion mutation. Pathway enrichment analysis of these genes revealed numerous pathways previously implicated in HD, as well as novel pathways that may be important in disease onset. Finally, we showed that a putative AO modifier in the ovarian-tumor-domain-containing deubiquitinase 3 (OTUD3) gene correlated with an altered rate of degeneration in patient-derived neurons, and that knockdown of OTUD3 accelerated degeneration in a human cell model of HD, validating our approach. This family-based strategy creates a novel resource for the HD community and establishes a framework that could be applied to study genetic modifiers of many other rare familial diseases.
BACKGROUND:Comprehensive treatment of Alzheimer's disease and related dementias (ADRD) requires not only pharmacologic treatment but also management of existing medical conditions and lifestyle modifications including diet, cognitive training, and exercise. Personalized, multimodal therapies are needed to best prevent and treat Alzheimer's disease (AD).OBJECTIVE:The Coaching for Cognition in Alzheimer's (COCOA) trial was a prospective randomized controlled trial to test the hypothesis that a remotely coached multimodal lifestyle intervention would improve early-stage AD.METHODS:Participants with early-stage AD were randomized into two arms. Arm 1 (N = 24) received standard of care. Arm 2 (N = 31) additionally received telephonic personalized coaching for multiple lifestyle interventions. The primary outcome was a test of the hypothesis that the Memory Performance Index (MPI) change over time would be better in the intervention arm than in the control arm. The Functional Assessment Staging Test was assessed for a secondary outcome. COCOA collected psychometric, clinical, lifestyle, genomic, proteomic, metabolomic, and microbiome data at multiple timepoints (dynamic dense data) across two years for each participant.RESULTS:The intervention arm ameliorated 2.1 [1.0] MPI points (mean [SD], p = 0.016) compared to the control over the two-year intervention. No important adverse events or side effects were observed.CONCLUSION:Multimodal lifestyle interventions are effective for ameliorating cognitive decline and have a larger effect size than pharmacological interventions. Dietary changes and exercise are likely to be beneficial components of multimodal interventions in many individuals. Remote coaching is an effective intervention for early stage ADRD. Remote interventions were effective during the COVID pandemic.
Alzheimer’s disease (AD) is a complex neurodegenerative condition that requires a comprehensive treatment approach. In addition to pharmacologic treatment, managing existing medical conditions and incorporating lifestyle modifications, such as diet, cognitive training, and exercise, are potentially crucial components of an effective intervention. In this study, we present preliminary results of a personalized multi-modal intervention for AD. The Precision Recommendations to Optimize Neurocognition (PREVENTION) study is an ongoing 12-month randomized clinical trial. Fifty participants (mean age 71.9(SD = 7.2); 24 Female) with biomarker evidence of AD amyloidosis have been recruited. While both arms receive personalized data-driven, lifestyle recommendations designed to target multiple systemic pathways implicated in AD, the active arm also receives health coaching, dietary counseling, exercise training, cognitive stimulation, and nutritional supplements. Comprehensive clinical, cognitive, neuroimaging, and genetic data are collected at baseline, and post-intervention. Herein, we examined the effects of the intervention on global cognition (MoCA) and regional brain volumes using general linear mixed models and compare to literature values (or historical controls). Participants in the two groups (23 active; 27 control) did not differ significantly in demographic, cognitive, or brain imaging measures at baseline. Thirty-one participants (16 active; 15 control) have completed the study. Within-group 12-month changes in MoCA (active:-1.4(3.5); control:-1.7(3.6)) were not significant, did not differ significantly between groups and were borderline (p = .08) less than expected rate of decline. While both groups declined significantly in hippocampal volumes, percent changes in total gray matter, entorhinal cortex, and precuneus volumes were significantly less in the active arm (Table). In this preliminary analysis of an active multimodal lifestyle intervention paired with personalized health optimization, regional volume loss in certain key areas of interest in AD was significantly lower in the active arm compared to control; further, the decline in global cognition was not statistically significant and was attenuated for the entire cohort compared to literature values. These preliminary results in a relatively small cohort suggest the intervention may beneficially impact cognition, memory, and brain structure. Future analyses will focus on elucidating changes in the biological systems being targeted by the intervention to help uncover the underlying mechanisms.
The goal of this Research Topic is to shed light on the progress made in the past decade in the Computational Genomics field, to gauge its future challenges, and to provide a thorough overview of the field’s current status. We hope that this article Research Topic will inform, inspire and provide guidance to researchers in the field. Foundational Research Topic ranging from still unsolved evolutionary mechanisms at the genomic level and the challenges posed by human genomic variation are powerful drivers of current and future research in Computational Genomics. The challenge that genomics poses to computational theory is to be highlighted, ranging from the role of Artificial Intelligence and Deep Learning (DL) in this context to the likely impact of emerging Single Cell RNA Sequencing (scRNA-Seq) methodologies in transcriptomics. Steady progress in tools for automated managing of large heterogeneous genomic and biological data is likely to bring good dividends in the near future. Specific applications of computational genomics support cancer studies for tasks such as drug repositioning and finding the role of immune system genes in cancer. Also, computational genomics is a key helper to plant science in the effort to cope with the effects of climate changes in the long run, with global food security as a goal. Here is an overview of the issues presented in this Research Topic. The evolution of genomes and codon encodings is a source of key fundamental questions still needing an answer. In this area, Belinky et al. highlight major differences between prokaryotes and eukaryotes regarding the double substitutions of nucleotides in codon encodings. Dong et al. provide experimental evidence on the performance of recently developed DL methods compared to more traditional flavors of Machine Learning (ML) when applied to the prediction of risk in cancer. Using a very large cohort of patients and three cancer test types, they give interesting hints for further research on this Research Topic. The Human Genome Project (HGP) lasted from 1990 to 2003 and has brought about a scientific revolution in genomics of the type the philosopher Thomas S. Kuhn has described. As with every scientific revolution, its long-lasting value is that of posing new questions. Singh et al. argue that the Human Pangenome Project is the next logical step on the road opened by the HGP, allowing us to reach new heights in genomic research in the near future. OPEN ACCESS
MOTIVATION:With the rapidly growing volume of knowledge and data in biomedical databases, improved methods for knowledge-graph-based computational reasoning are needed in order to answer translational questions. Previous efforts to solve such challenging computational reasoning problems have contributed tools and approaches, but progress has been hindered by the lack of an expressive analysis workflow language for translational reasoning and by the lack of a reasoning engine-supporting that language-that federates semantically integrated knowledge-bases. RESULTS:We introduce ARAX, a new reasoning system for translational biomedicine that provides a web browser user interface and an application programming interface (API). ARAX enables users to encode translational biomedical questions and to integrate knowledge across sources to answer the user's query and facilitate exploration of results. For ARAX, we developed new approaches to query planning, knowledge-gathering, reasoning and result ranking and dynamically integrate knowledge providers for answering biomedical questions. To illustrate ARAX's application and utility in specific disease contexts, we present several use-case examples. AVAILABILITY AND IMPLEMENTATION:The source code and technical documentation for building the ARAX server-side software and its built-in knowledge database are freely available online (https://github.com/RTXteam/RTX). We provide a hosted ARAX service with a web browser interface at arax.rtx.ai and a web API endpoint at arax.rtx.ai/api/arax/v1.3/ui/. SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
Aging manifests as progressive deteriorations in homeostasis, requiring systems-level perspectives to investigate the gradual molecular dysregulation of underlying biological processes. Here, we report systemic changes in the molecular regulation of biological processes under multiple lifespan-extending interventions. Differential Rank Conservation (DIRAC) analyses of mouse liver proteomics and transcriptomics data show that mechanistically distinct lifespan-extending interventions (acarbose, 17α-estradiol, rapamycin, and calorie restriction) generally tighten the regulation of biological modules. These tightening patterns are similar across the interventions, particularly in processes such as fatty acid oxidation, immune response, and stress response. Differences in DIRAC patterns between proteins and transcripts highlight specific modules which may be tightened via augmented cap-independent translation. Moreover, the systemic shifts in fatty acid metabolism are supported through integrated analysis of liver transcriptomics data with a mouse genome-scale metabolic model. Our findings highlight the power of systems-level approaches for identifying and characterizing the biological processes involved in aging and longevity.
Knowledge graphs have become a common approach for knowledge representation. Yet, the application of graph methodology is elusive due to the sheer number and complexity of knowledge sources. In addition, semantic incompatibilities hinder efforts to harmonize and integrate across these diverse sources. As part of The Biomedical Translator Consortium, we have developed a knowledge graph-based question-answering system designed to augment human reasoning and accelerate translational scientific discovery: the Translator system. We have applied the Translator system to answer biomedical questions in the context of a broad array of diseases and syndromes, including Fanconi anemia, primary ciliary dyskinesia, multiple sclerosis, and others. A variety of collaborative approaches have been used to research and develop the Translator system. One recent approach involved the establishment of a monthly "Question-of-the-Month (QotM) Challenge" series. Herein, we describe the structure of the QotM Challenge; the six challenges that have been conducted to date on drug-induced liver injury, cannabidiol toxicity, coronavirus infection, diabetes, psoriatic arthritis, and ATP1A3-related phenotypes; the scientific insights that have been gleaned during the challenges; and the technical issues that were identified over the course of the challenges and that can now be addressed to foster further development of the prototype Translator system. We close with a discussion on Large Language Models such as ChatGPT and highlight differences between those models and the Translator system.
Background Biomedical translational science is increasingly using computational reasoning on repositories of structured knowledge (such as UMLS, SemMedDB, ChEMBL, Reactome, DrugBank, and SMPDB in order to facilitate discovery of new therapeutic targets and modalities. The NCATS Biomedical Data Translator project is working to federate autonomous reasoning agents and knowledge providers within a distributed system for answering translational questions. Within that project and the broader field, there is a need for a framework that can efficiently and reproducibly build an integrated, standards-compliant, and comprehensive biomedical knowledge graph that can be downloaded in standard serialized form or queried via a public application programming interface (API). Results To create a knowledge provider system within the Translator project, we have developed RTX-KG2, an open-source software system for building—and hosting a web API for querying—a biomedical knowledge graph that uses an Extract-Transform-Load approach to integrate 70 knowledge sources (including the aforementioned core six sources) into a knowledge graph with provenance information including (where available) citations. The semantic layer and schema for RTX-KG2 follow the standard Biolink model to maximize interoperability. RTX-KG2 is currently being used by multiple Translator reasoning agents, both in its downloadable form and via its SmartAPI-registered interface. Serializations of RTX-KG2 are available for download in both the pre-canonicalized form and in canonicalized form (in which synonyms are merged). The current canonicalized version (KG2.7.3) of RTX-KG2 contains 6.4M nodes and 39.3M edges with a hierarchy of 77 relationship types from Biolink. Conclusion RTX-KG2 is the first knowledge graph that integrates UMLS, SemMedDB, ChEMBL, DrugBank, Reactome, SMPDB, and 64 additional knowledge sources within a knowledge graph that conforms to the Biolink standard for its semantic layer and schema. RTX-KG2 is publicly available for querying via its API at arax.rtx.ai/api/rtxkg2/v1.2/openapi.json . The code to build RTX-KG2 is publicly available at github:RTXteam/RTX-KG2 .