Since publication of the FAIR Guiding Principles in 2016, the scientific community has increasingly sought to make experimental data findable, accessible, interoperable, and reusable. Operationalizing the FAIR principles in routine scientific workflows remains challenging without a standardized, workable infrastructure. With over 10,000 datasets from over 40 institutions, spanning more than 50 diverse assay types ranging from single-cell sequencing technologies to 2D and 3D spatial omics, the U.S. National Institutes of Health (NIH) Human Bio-Molecular Atlas Program (HuBMAP) consortium has been ideally situated to create a FAIR ecosystem. With the goal of achieving data "FAIRness," HuBMAP developed and implemented well-defined, community-endorsed metadata reporting standards across the research lifecycle. These reporting standards include detailed schemas, harmonized across a multitude of assays, that define the metadata associated with a dataset and the organization of the corresponding data files. These standards ensure documentation of the data collection process, of the data themselves, and of the manner in which the data are packaged for sharing, while remaining compliant with the Health Insurance Portability and Accountability Act (HIPAA). The use of these reporting standards, in tandem with technology to foster adherence, allows HuBMAP to fulfill its goal of generating FAIR data for open dissemination through its Data Portal and Human Reference Atlas. The procedures and simple workflow adopted by HuBMAP investigators serve as a model for other scientific communities aiming to maximize the value of varied datasets addressing a shared research question. The HuBMAP end-to-end, metadata-centered workflow has been replicated and enhanced by the NIH Cellular Senescence Network (SenNet) consortium and is readily available through open-source technology for others to utilize.
The NIH Common Fund Data Ecosystem (CFDE) integrates data resources from 18 NIH Common Fund programs for discovery and integrative analysis. These programs generate valuable but heterogeneous datasets that can be difficult to discover, access, and reuse. CFDE aims to provide a collaborative, community-built infrastructure that links and enriches Common Fund programs. We describe the evolution, structure, and core technologies of CFDE, including practical approaches that support submission, integration, visualization, and public release of multimodal data. Training programs and workforce initiatives lower barriers to adoption. CFDE has devised solutions to critical issues facing cross-program initiatives, including data scale and heterogeneity, dataset integration, and long-term sustainability. We demonstrate the utility of linking Common Fund resources through integrative tools and cross-dataset queries to yield insights that would otherwise be infeasible. Collectively, CFDE shows that a standards-driven, federated approach enhances and unifies cross-disciplinary resources, fostering collaboration and data-driven discovery.
Cellular senescence is a hallmark of aging and a driver of functional decline across tissues, yet its heterogeneity and context dependence have limited systematic study. The Common Fund’s Cellular Senescence Network (SenNet) Program addresses this challenge by generating multimodal, multi-tissue datasets that profile senescent cells across the human lifespan and complementary mouse models. The SenNet Data Portal ( https://data.sennetconsortium.org ) serves as the public gateway to these resources, providing open access to harmonized single-cell, spatial, imaging, transcriptomic, and proteomic data; senescence biomarker catalogs; and standardized protocols that can be used to comprehensively identify and characterize senescent cells in mouse and human tissue. As of January 2026, the portal hosts 1,753 publicly available human and mouse datasets across 15 organs using 6 general assay types. Experts from 13 Tissue Mapping Centers (TMCs) and 12 Technology Development and Application (TDAs) components contribute tissue data, analyze data, identify senescent biomarkers, and agree on panels for cross-tissue antibody harmonization. They also register human tissue data into the Human Reference Atlas (HRA) and develop user interfaces for the multiscale and multimodal exploration of this data. Built on a scalable hybrid cloud microservices architecture by the Consortium Organization and Data Coordinating Center (CODCC), the Portal enables data submission, management, integrated analysis, spatial context mapping, and cross-species senescence mapping critical for aging research. This paper presents user needs, the Portal’s architecture, data processing workflows, and senescence-focused analytical tools. The paper also presents usage scenarios illustrating applications in biomarker discovery, quality benchmarking, hypothesis generation, spatial analysis, cost-efficient profiling, and cell distance distribution analysis. Current limitations and planned extensions—including expanded spatial-omics releases and improved tools for senotype characterization—are discussed. SenNet protocols, code, and user interfaces are freely available on https://docs.sennetconsortium.org/apis .
The Human BioMolecular Atlas Program (HuBMAP) aims to construct a 3D Human Reference Atlas (HRA) of the healthy adult body. Experts from 20+ consortia collaborate to develop a Common Coordinate Framework (CCF), knowledge graphs and tools that describe the multiscale structure of the human body (from organs and tissues down to cells, genes and biomarkers) and to use the HRA to characterize changes that occur with aging, disease and other perturbations. HRA v.2.0 covers 4,499 unique anatomical structures, 1,195 cell types and 2,089 biomarkers (such as genes, proteins and lipids) from 33 ASCT+B tables and 65 3D Reference Objects linked to ontologies. New experimental data can be mapped into the HRA using (1) cell type annotation tools (for example, Azimuth), (2) validated antibody panels or (3) by registering tissue data spatially. This paper describes HRA user stories, terminology, data formats, ontology validation, unified analysis workflows, user interfaces, instructional materials, application programming interfaces, flexible hybrid cloud infrastructure and previews atlas usage applications.
BACKGROUND:Social Determinants of Health impact health outcomes. Area Deprivation Index (ADI) is used to risk-adjust for neighborhood affluence/deprivation but guidance on choosing deprivation cutoffs is lacking. We hypothesize that different ADI cutoffs are required for different insurance types. METHODS:National Surgical Quality Improvement Program data 2013-2019 merged with electronic health records from three academic healthcare systems. Desirability of Outcome Ranking (DOOR) assessed the association of ADI cutoffs for different insurance types, adjusted for operative stress, frailty, and case status (elective, urgent, emergent). Secondary analyses assessed the association of ADI with case status. RESULTS:Patients with Private insurance living in areas with ADI>85 had higher/worse DOOR outcomes, which lost significance after adjusting for case status. Medicare cases with ADI>75 exhibited higher/worse DOOR outcomes even after adjusting for case status. ADI was not associated with outcomes in the Medicaid and Uninsured groups. High ADI was associated with increased odds of urgent and emergent cases for the Private and Medicare but not Medicaid or Uninsured groups. CONCLUSIONS:ADI is a useful metric to identify at-risk patients and can be used for risk adjustment. Health systems must understand their population demographics and use their data to determine ADI cutoffs. Patients in deprived neighborhoods have higher odds of urgent and emergent surgeries, despite having Private insurance or Medicare, suggesting that delays/barriers to primary and preventive care may be a major driver of worse outcomes. While insurance coverage is important, healthcare policies supporting reductions in urgent/emergent cases could have the largest impact on improving outcomes.
The Homo sapiens Chromosomal Location Ontology (HSCLO) is designed to facilitate the integration of human genomic features into biomedical knowledge graphs from releases GRCh37 and GRCh38 at multiple resolutions. HSCLO comprises two distinct versions, HSCLO37 and HSCLO38, each tailored to its respective human genome release. This ontology supports the efficient integration and analysis of human genomic data across scales ranging from entire chromosomes to individual base pairs, thereby enhancing data retrieval and interoperability within large-scale biomedical datasets. Unlike existing ontologies that primarily focus on genomic feature identification or annotation, HSCLO is specifically engineered to optimize the interoperability and scalability of genomic data within biomedical knowledge graphs. The utility and performance of HSCLO are demonstrated through a case study involving the integration of high-resolution chromatin interaction data, which reveals significant improvements in query efficiency and data linkage. HSCLO represents a valuable resource for advancing research in disease genetics, personalized medicine, and other domains that require complex genomic data integration.
INTRODUCTION There is need to detect and intervene in pre-clinical phases of Alzheimer’s disease (AD). Electronic health records (EHRs) may help predict AD using machine learning methods. METHODS We identified EHRs for 19,473 cases with AD and 111,922 controls. Records spanned 10 or more years prior to AD diagnosis. We trained a random forest model (employing 5-fold cross-validation with 2,499 features) to predict AD 10 years prior to its onset using a 75/25% train/test split and then computed permuted feature importance. RESULTS We achieved an area under the ROC curve of 0.80. Feature importance identified factors associated with AD, including age, sex, race, ethnicity, BMI, cardiovascular diseases, inflammation, pain, sleep and mood disorders, trauma, other neurodegenerative disorders, diuretics, colon-related disorders and procedures, seizures, and vitamin B12. DISCUSSION This is the first EHR-based model to predict AD 10 years prior to onset, which could help predict AD and inform prevention/early intervention.
Once considered a tissue culture-specific phenomenon, cellular senescence has now been linked to various biological processes with both beneficial and detrimental roles in humans, rodents and other species. Much of our understanding of senescent cell biology still originates from tissue culture studies, where each cell in the culture is driven to an irreversible cell cycle arrest. By contrast, in tissues, these cells are relatively rare and difficult to characterize, and it is now established that fully differentiated, postmitotic cells can also acquire a senescence phenotype. The SenNet Biomarkers Working Group was formed to provide recommendations for the use of cellular senescence markers to identify and characterize senescent cells in tissues. Here, we provide recommendations for detecting senescent cells in different tissues based on a comprehensive analysis of existing literature reporting senescence markers in 14 tissues in mice and humans. We discuss some of the recent advances in detecting and characterizing cellular senescence, including molecular senescence signatures and morphological features, and the use of circulating markers. We aim for this work to be a valuable resource for both seasoned investigators in senescence-related studies and newcomers to the field. Senescent cells have complex and important roles in cancer and ageing, but they are quite rare and difficult to characterize in tissues in vivo. In this Expert Recommendation, the SenNet Biomarkers Working Group discusses recent advances in detecting and characterizing cellular senescence and provides recommendations for senescence markers in 14 human and mouse tissues.
Importance:Insurance coverage expansion has been proposed as a solution to improving health disparities, but insurance expansion alone may be insufficient to alleviate care access barriers. Objective:To assess the association of Area Deprivation Index (ADI) with postsurgical textbook outcomes (TO) and presentation acuity for individuals with private insurance or Medicare. Design, Setting, and Participants:This cohort study used data from the National Surgical Quality Improvement Program (2013-2019) merged with electronic health record data from 3 academic health care systems. Data were analyzed from June 2022 to August 2023. Exposure:Living in a neighborhood with an ADI greater than 85. Main Outcomes and Measures:TO, defined as absence of unplanned reoperations, Clavien-Dindo grade 4 complications, mortality, emergency department visits/observation stays, and readmissions, and presentation acuity, defined as having preoperative acute serious conditions (PASC) and urgent or emergent cases. Results:Among a cohort of 29 924 patients, the mean (SD) age was 60.6 (15.6) years; 16 424 (54.9%) were female, and 13 500 (45.1) were male. A total of 14 306 patients had private insurance and 15 618 had Medicare. Patients in highly deprived neighborhoods (5536 patients [18.5%]), with an ADI greater than 85, had lower/worse odds of TO in both the private insurance group (adjusted odds ratio [aOR], 0.87; 95% CI, 0.76-0.99; P = .04) and Medicare group (aOR, 0.90; 95% CI, 0.82-1.00; P = .04) and higher odds of PASC and urgent or emergent cases. The association of ADIs greater than 85 with TO lost significance after adjusting for PASC and urgent/emergent cases. Differences in the probability of TO between the lowest-risk (ADI ≤85, no PASC, and elective surgery) and highest-risk (ADI >85, PASC, and urgent/emergent surgery) scenarios stratified by frailty were highest for very frail patients (Risk Analysis Index ≥40) with differences of 40.2% and 43.1% for those with private insurance and Medicare, respectively. Conclusions and Relevance:This study found that patients living in highly deprived neighborhoods had lower/worse odds of TO and higher presentation acuity despite having private insurance or Medicare. These findings suggest that insurance coverage expansion alone is insufficient to overcome health care disparities, possibly due to persistent barriers to preventive care and other complex causes of health inequities.
PURPOSE In the United States, a comprehensive national breast cancer registry (CR) does not exist. Thus, care and coverage decisions are based on data from population subsets, other countries, or models. We report a prototype real-world research data mart to assess mortality, morbidity, and costs for breast cancer diagnosis and treatment. METHODS With institutional review board approval and Health Insurance Portability and Accountability Act (HIPPA) compliance, a multidisciplinary clinical and research data warehouse (RDW) expert group curated demographic, risk, imaging, pathology, treatment, and outcome data from the electronic health records (EHR), radiology (RIS), and CR for patients having breast imaging and/or a diagnosis of breast cancer in our institution from January 1, 2004, to December 31, 2020. Domains were defined by prebuilt views to extract data denormalized according to requirements from the existing RDW using an export, transform, load pattern. Data dictionaries were included. Structured query language was used for data cleaning. RESULTS Five-hundred eighty-nine elements (EHR 311, RIS 211, and CR 67) were mapped to 27 domains; all, except one containing CR elements, had cancer and noncancer cohort views, resulting in a total of 53 views (average 12 elements/view; range, 4-67). EHR and RIS queries returned 497,218 patients with 2,967,364 imaging examinations and associated visit details. Cancer biology, treatment, and outcome details for 15,619 breast cancer cases were imported from the CR of our primary breast care facility for this prototype mart. CONCLUSION Institutional real-world data marts enable comprehensive understanding of care outcomes within an organization. As clinical data sources become increasingly structured, such marts may be an important source for future interinstitution analysis and potentially an opportunity to create robust real-world results that could be used to support evidence-based national policy and care decisions for breast cancer.
Over the past decade, there has been substantial growth in both the quantity and complexity of available biomedical data. In order to more efficiently harness this extensive data and alleviate challenges associated with integration of multi-omics data, we developed Petagraph, a biomedical knowledge graph that encompasses over 32 million nodes and 118 million relationships. Petagraph leverages more than 180 ontologies and standards in the Unified Biomedical Knowledge Graph (UBKG) to embed millions of quantitative genomics data points. Petagraph provides a cohesive data environment that enables users to efficiently analyze, annotate, and discern relationships within and across complex multi-omics datasets supported by UBKG’s annotation scaffold. We demonstrate how queries on Petagraph can generate meaningful results across various research contexts and use cases.
The Human BioMolecular Atlas Program (HuBMAP) aims to construct a reference 3D structural, cellular, and molecular atlas of the healthy adult human body. The HuBMAP Data Portal (https://portal.hubmapconsortium.org) serves experimental datasets and supports data processing, search, filtering, and visualization. The Human Reference Atlas (HRA) Portal (https://humanatlas.io) provides open access to atlas data, code, procedures, and instructional materials. Experts from more than 20 consortia are collaborating to construct the HRA's Common Coordinate Framework (CCF), knowledge graphs, and tools that describe the multiscale structure of the human body (from organs and tissues down to cells, genes, and biomarkers) and to use the HRA to understand changes that occur at each of these levels with aging, disease, and other perturbations. The 6th release of the HRA v2.0 covers 36 organs with 4,499 unique anatomical structures, 1,195 cell types, and 2,089 biomarkers (e.g., genes, proteins, lipids) linked to ontologies and 2D/3D reference objects. New experimental data can be mapped into the HRA using (1) three cell type annotation tools (e.g., Azimuth) or (2) validated antibody panels (OMAPs), or (3) by registering tissue data spatially. This paper describes the HRA user stories, terminology, data formats, ontology validation, unified analysis workflows, user interfaces, instructional materials, application programming interface (APIs), flexible hybrid cloud infrastructure, and previews atlas usage applications.
Translational research requires data at multiple scales of biological organization. Advancements in sequencing and multi-omics technologies have increased the availability of these data, but researchers face significant integration challenges. Knowledge graphs (KGs) are used to model complex phenomena, and methods exist to construct them automatically. However, tackling complex biomedical integration problems requires flexibility in the way knowledge is modeled. Moreover, existing KG construction methods provide robust tooling at the cost of fixed or limited choices among knowledge representation models. PheKnowLator (Phenotype Knowledge Translator) is a semantic ecosystem for automating the FAIR (Findable, Accessible, Interoperable, and Reusable) construction of ontologically grounded KGs with fully customizable knowledge representation. The ecosystem includes KG construction resources (e.g., data preparation APIs), analysis tools (e.g., SPARQL endpoint resources and abstraction algorithms), and benchmarks (e.g., prebuilt KGs). We evaluated the ecosystem by systematically comparing it to existing open-source KG construction methods and by analyzing its computational performance when used to construct 12 different large-scale KGs. With flexible knowledge representation, PheKnowLator enables fully customizable KGs without compromising performance or usability.
Objective: Develop an ordinal Desirability of Outcome Ranking (DOOR) for surgical outcomes to examine complex associations of Social Determinants of Health.Background: Studies focused on single or binary composite outcomes may not detect health disparities.Methods: Three health care system cohort study using NSQIP (2013-2019) linked with EHR and risk-adjusted for frailty, preoperative acute serious conditions (PASC), case status and operative stress assessing associations of multilevel Social Determinants of Health of race/ethnicity, insurance type (Private 13,957; Medicare 15,198; Medicaid 2835; Uninsured 2963) and Area Deprivation Index (ADI) on DOOR and the binary Textbook Outcomes (TO).Results: Patients living in highly deprived neighborhoods (ADI>85) had higher odds of PASC [adjusted odds ratio (aOR)=1.13, CI=1.02-1.25, P<0.001] and urgent/emergent cases (aOR=1.23, CI=1.16-1.31, P<0.001). Increased odds of higher/less desirable DOOR scores were associated with patients identifying as Black versus White and on Medicare, Medicaid or Uninsured versus Private insurance. Patients with ADI>85 had lower odds of TO (aOR=0.91, CI=0.85-0.97, P=0.006) until adjusting for insurance. In contrast, patients with ADI>85 had increased odds of higher DOOR (aOR=1.07, CI=1.01-1.14, P<0.021) after adjusting for insurance but similar odds after adjusting for PASC and urgent/emergent cases.Conclusions: DOOR revealed complex interactions between race/ethnicity, insurance type and neighborhood deprivation. ADI>85 was associated with higher odds of worse DOOR outcomes while TO failed to capture the effect of ADI. Our results suggest that presentation acuity is a critical determinant of worse outcomes in patients in highly deprived neighborhoods and without insurance. Including risk adjustment for living in deprived neighborhoods and urgent/emergent surgeries could improve the accuracy of quality metrics.
Multiplexed antibody-based imaging enables the detailed characterization of molecular and cellular organization in tissues. Advances in the field now allow high-parameter data collection (>60 targets); however, considerable expertise and capital are needed to construct the antibody panels employed by these methods. Organ mapping antibody panels are community-validated resources that save time and money, increase reproducibility, accelerate discovery and support the construction of a Human Reference Atlas.
Rehabilitation research focuses on determining the components of a treatment intervention, the mechanism of how these components lead to recovery and rehabilitation, and ultimately the optimal intervention strategies to maximize patients' physical, psychologic, and social functioning. Traditional randomized clinical trials that study and establish new interventions face challenges, such as high cost and time commitment. Observational studies that use existing clinical data to observe the effect of an intervention have shown several advantages over RCTs. Electronic Health Records (EHRs) have become an increasingly important resource for conducting observational studies. To support these studies, we developed a clinical research datamart, called ReDWINE (Rehabilitation Datamart With Informatics iNfrastructure for rEsearch), that transforms the rehabilitation-related EHR data collected from the UPMC health care system to the Observational Health Data Sciences and Informatics (OHDSI) Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) to facilitate rehabilitation research. The standardized EHR data stored in ReDWINE will further reduce the time and effort required by investigators to pool, harmonize, clean, and analyze data from multiple sources, leading to more robust and comprehensive research findings. ReDWINE also includes deployment of data visualization and data analytics tools to facilitate cohort definition and clinical data analysis. These include among others the Open Health Natural Language Processing (OHNLP) toolkit, a high-throughput NLP pipeline, to provide text analytical capabilities at scale in ReDWINE. Using this comprehensive representation of patient data in ReDWINE for rehabilitation research will facilitate real-world evidence for health interventions and outcomes.
Nonalcoholic fatty liver disease (NAFLD) and nonalcoholic steatohepatitis (NASH) are highly prevalent but underdiagnosed. We used an electronic health record data network to test a population-level risk stratification strategy using noninvasive tests (NITs) of liver fibrosis. Data were obtained from PCORnet® sites in the East, Midwest, Southwest, and Southeast United States from patients aged ≥ 18 with or without ICD-10-CM diagnosis codes for NAFLD, NASH, and NASH-cirrhosis between 9/1/2017 and 8/31/2020. Average and standard deviations (SD) for Fibrosis-4 index (FIB-4), NAFLD fibrosis score (NFS), and Hepatic Steatosis Index (HSI) were estimated by site for each patient cohort. Sample-wide estimates were calculated as weighted averages across study sites. Of 11,875,959 patients, 0.8
Objective Rare disease research requires data sharing networks to power translational studies. We describe novel use of Research Electronic Data Capture (REDCap), a web application for managing clinical data, by the National Mesothelioma Virtual Bank, a federated biospecimen, and data sharing network. Materials and Methods National Mesothelioma Virtual Bank (NMVB) uses REDCap to integrate honest broker activities, enabling biospecimen and associated clinical data provisioning to investigators. A Web Portal Query tool was developed to source and visualize REDCap data in interactive, faceted search, enabling cohort discovery by public users. An AWS Lambda function behind an API calculates the counts visually presented, while protecting record level data. The user-friendly interface, quick responsiveness, automatic generation from REDCap, and flexibility to new data, was engineered to sustain the NMVB research community. Results NMVB implementations enabled a network of 8 research institutions with over 2000 mesothelioma cases, including clinical annotations and biospecimens, and public users' cohort discovery and summary statistics. NMVB usage and impact is demonstrated by high website visits (>150 unique queries per month), resource use requests (>50 letter of interests), and citations (>900) to papers published using NMVB resources. Discussion NMVB's REDCap implementation and query tool is a framework for implementing federated and integrated rare disease biobanks and registries. Advantages of this framework include being low-cost, modular, scalable, and efficient. Future advances to NVMB's implementations will include incorporation of -omics data and development of downstream analysis tools to advance mesothelioma and rare disease research. Conclusion NVMB presents a framework for integrating biobanks and patient registries to enable translational research for rare diseases.
Objective:Assess associations of Social Determinants of Health (SDoH) using Area Deprivation Index (ADI), race/ethnicity and insurance type with Textbook Outcomes (TO). Summary Background Data:Individual- and contextual-level SDoH affect health outcomes, but only one SDoH level is usually included. Methods:Three healthcare system cohort study using National Surgical Quality Improvement Program (2013-2019) linked with ADI risk-adjusted for frailty, case status and operative stress examining TO/TO components (unplanned reoperations, complications, mortality, Emergency Department/Observation Stays and readmissions). Results:Cohort (34,251 cases) mean age 58.3 [SD=16.0], 54.8% females, 14.1% Hispanics, 11.6% Non-Hispanic Blacks, 21.6% with ADI>85, and 81.8% TO. Racial and ethnic minorities, non-Private insurance, and ADI>85 patients had increased odds of urgent/emergent surgeries (aORs range: 1.17-2.83, all P<.001). Non-Hispanic Black patients, ADI>85 and non-Private insurances had lower TO odds (aORs range: 0.55-0.93, all P<.04), but ADI>85 lost significance after including case status. Urgent/emergent versus elective had lower TO odds (aOR=0.51, P<.001). ADI>85 patients had higher complication and mortality odds. Estimated reduction in TO probability was 9.9% (CI=7.2%-12.6%) for urgent/emergent cases, 7.0% (CI=4.6%-9.3%) for Medicaid, and 1.6% (CI=0.2%-3.0%) for non-Hispanic Black patients. TO probability difference for lowest-risk (White-Private-ADI≤85-elective) to highest-risk (Black-Medicaid-ADI>85-urgent/emergent) was 29.8% for very frail patients. Conclusion:Multi-level SDoH had independent effects on TO, predominately affecting outcomes through increased rates/odds of urgent/emergent surgeries driving complications and worse outcomes. Lowest-risk versus highest-risk scenarios demonstrated the magnitude of intersecting SDoH variables. Combination of insurance type and ADI should be used to identify high-risk patients to redesign care pathways to improve outcomes. Risk adjustment including contextual neighborhood deprivation and patient-level SDoH could reduce unintended consequences of value-based programs.