Since publication of the FAIR Guiding Principles in 2016, the scientific community has increasingly sought to make experimental data findable, accessible, interoperable, and reusable. Operationalizing the FAIR principles in routine scientific workflows remains challenging without a standardized, workable infrastructure. With over 10,000 datasets from over 40 institutions, spanning more than 50 diverse assay types ranging from single-cell sequencing technologies to 2D and 3D spatial omics, the U.S. National Institutes of Health (NIH) Human Bio-Molecular Atlas Program (HuBMAP) consortium has been ideally situated to create a FAIR ecosystem. With the goal of achieving data "FAIRness," HuBMAP developed and implemented well-defined, community-endorsed metadata reporting standards across the research lifecycle. These reporting standards include detailed schemas, harmonized across a multitude of assays, that define the metadata associated with a dataset and the organization of the corresponding data files. These standards ensure documentation of the data collection process, of the data themselves, and of the manner in which the data are packaged for sharing, while remaining compliant with the Health Insurance Portability and Accountability Act (HIPAA). The use of these reporting standards, in tandem with technology to foster adherence, allows HuBMAP to fulfill its goal of generating FAIR data for open dissemination through its Data Portal and Human Reference Atlas. The procedures and simple workflow adopted by HuBMAP investigators serve as a model for other scientific communities aiming to maximize the value of varied datasets addressing a shared research question. The HuBMAP end-to-end, metadata-centered workflow has been replicated and enhanced by the NIH Cellular Senescence Network (SenNet) consortium and is readily available through open-source technology for others to utilize.
The NIH Common Fund Data Ecosystem (CFDE) integrates data resources from 18 NIH Common Fund programs for discovery and integrative analysis. These programs generate valuable but heterogeneous datasets that can be difficult to discover, access, and reuse. CFDE aims to provide a collaborative, community-built infrastructure that links and enriches Common Fund programs. We describe the evolution, structure, and core technologies of CFDE, including practical approaches that support submission, integration, visualization, and public release of multimodal data. Training programs and workforce initiatives lower barriers to adoption. CFDE has devised solutions to critical issues facing cross-program initiatives, including data scale and heterogeneity, dataset integration, and long-term sustainability. We demonstrate the utility of linking Common Fund resources through integrative tools and cross-dataset queries to yield insights that would otherwise be infeasible. Collectively, CFDE shows that a standards-driven, federated approach enhances and unifies cross-disciplinary resources, fostering collaboration and data-driven discovery.
Cellular senescence is a hallmark of aging and a driver of functional decline across tissues, yet its heterogeneity and context dependence have limited systematic study. The Common Fund’s Cellular Senescence Network (SenNet) Program addresses this challenge by generating multimodal, multi-tissue datasets that profile senescent cells across the human lifespan and complementary mouse models. The SenNet Data Portal ( https://data.sennetconsortium.org ) serves as the public gateway to these resources, providing open access to harmonized single-cell, spatial, imaging, transcriptomic, and proteomic data; senescence biomarker catalogs; and standardized protocols that can be used to comprehensively identify and characterize senescent cells in mouse and human tissue. As of January 2026, the portal hosts 1,753 publicly available human and mouse datasets across 15 organs using 6 general assay types. Experts from 13 Tissue Mapping Centers (TMCs) and 12 Technology Development and Application (TDAs) components contribute tissue data, analyze data, identify senescent biomarkers, and agree on panels for cross-tissue antibody harmonization. They also register human tissue data into the Human Reference Atlas (HRA) and develop user interfaces for the multiscale and multimodal exploration of this data. Built on a scalable hybrid cloud microservices architecture by the Consortium Organization and Data Coordinating Center (CODCC), the Portal enables data submission, management, integrated analysis, spatial context mapping, and cross-species senescence mapping critical for aging research. This paper presents user needs, the Portal’s architecture, data processing workflows, and senescence-focused analytical tools. The paper also presents usage scenarios illustrating applications in biomarker discovery, quality benchmarking, hypothesis generation, spatial analysis, cost-efficient profiling, and cell distance distribution analysis. Current limitations and planned extensions—including expanded spatial-omics releases and improved tools for senotype characterization—are discussed. SenNet protocols, code, and user interfaces are freely available on https://docs.sennetconsortium.org/apis .
The Human BioMolecular Atlas Program (HuBMAP) aims to construct a 3D Human Reference Atlas (HRA) of the healthy adult body. Experts from 20+ consortia collaborate to develop a Common Coordinate Framework (CCF), knowledge graphs and tools that describe the multiscale structure of the human body (from organs and tissues down to cells, genes and biomarkers) and to use the HRA to characterize changes that occur with aging, disease and other perturbations. HRA v.2.0 covers 4,499 unique anatomical structures, 1,195 cell types and 2,089 biomarkers (such as genes, proteins and lipids) from 33 ASCT+B tables and 65 3D Reference Objects linked to ontologies. New experimental data can be mapped into the HRA using (1) cell type annotation tools (for example, Azimuth), (2) validated antibody panels or (3) by registering tissue data spatially. This paper describes HRA user stories, terminology, data formats, ontology validation, unified analysis workflows, user interfaces, instructional materials, application programming interfaces, flexible hybrid cloud infrastructure and previews atlas usage applications.
The Playbook Workflow Builder (PWB) is a web-based platform to dynamically construct and execute bioinformatics workflows by utilizing a growing network of input datasets, semantically annotated API endpoints, and data visualization tools contributed by an ecosystem of collaborators. Via a user-friendly user interface, workflows can be constructed from contributed building-blocks without technical expertise. The output of each step of the workflow is added into reports containing textual descriptions, figures, tables, and references. To construct workflows, users can click on cards that represent each step in a workflow, or construct workflows via a chat interface that is assisted by a large language model (LLM). Completed workflows are compatible with Common Workflow Language (CWL) and can be published as research publications, slideshows, and posters. To demonstrate how the PWB generates meaningful hypotheses that draw knowledge from across multiple resources, we present several use cases. For example, one of these use cases prioritizes drug targets for individual cancer patients using data from the NIH Common Fund programs GTEx, LINCS, Metabolomics, GlyGen, and ExRNA. The workflows created with PWB can be repurposed to tackle similar use cases using different inputs. The PWB platform is available from: https://playbook-workflow-builder.cloud/.
Many biomedical research projects produce large-scale datasets that may serve as resources for the research community for hypothesis generation, facilitating diverse use cases. Towards the goal of developing infrastructure to support the findability, accessibility, interoperability, and reusability (FAIR) of biomedical digital objects and maximally extracting knowledge from data, complex queries that span across data and tools from multiple resources are currently not easily possible. By utilizing existing FAIR application programming interfaces (APIs) that serve knowledge from many repositories and bioinformatics tools, different types of complex queries and workflows can be created by using these APIs together. The Playbook Workflow Builder (PWB) is a web-based platform that facilitates interactive construction of workflows by enabling users to utilize an ever-growing network of input datasets, semantically annotated API endpoints, and data visualization tools contributed by an ecosystem. Via a user-friendly web-based user interface (UI), workflows can be constructed from contributed building-blocks without technical expertise. The output of each step of the workflows are provided in reports containing textual descriptions, as well as interactive and downloadable figures and tables. To demonstrate the ability of the PWB to generate meaningful hypotheses that draw knowledge from across multiple resources, we present several use cases. For example, one of these use cases sieves novel targets for individual cancer patients using data from the GTEx, LINCS, Metabolomics, GlyGen, and the ExRNA Communication Consortium (ERCC) Common Fund (CF) Data Coordination Centers (DCCs). The workflows created with the PWB can be published and repurposed to tackle similar use cases using different inputs. The PWB platform is available from: . ### Competing Interest Statement The authors have declared no competing interest.
The Human BioMolecular Atlas Program (HuBMAP) aims to construct a reference 3D structural, cellular, and molecular atlas of the healthy adult human body. The HuBMAP Data Portal (https://portal.hubmapconsortium.org) serves experimental datasets and supports data processing, search, filtering, and visualization. The Human Reference Atlas (HRA) Portal (https://humanatlas.io) provides open access to atlas data, code, procedures, and instructional materials. Experts from more than 20 consortia are collaborating to construct the HRA's Common Coordinate Framework (CCF), knowledge graphs, and tools that describe the multiscale structure of the human body (from organs and tissues down to cells, genes, and biomarkers) and to use the HRA to understand changes that occur at each of these levels with aging, disease, and other perturbations. The 6th release of the HRA v2.0 covers 36 organs with 4,499 unique anatomical structures, 1,195 cell types, and 2,089 biomarkers (e.g., genes, proteins, lipids) linked to ontologies and 2D/3D reference objects. New experimental data can be mapped into the HRA using (1) three cell type annotation tools (e.g., Azimuth) or (2) validated antibody panels (OMAPs), or (3) by registering tissue data spatially. This paper describes the HRA user stories, terminology, data formats, ontology validation, unified analysis workflows, user interfaces, instructional materials, application programming interface (APIs), flexible hybrid cloud infrastructure, and previews atlas usage applications.
This paper uses accounting concepts—particularly the concept of Return on Investment (ROI)—to reveal the quantitative value of scientific research pertaining to a major US cyberinfrastructure project (XSEDE—the eXtreme Science and Engineering Discovery Environment). XSEDE provides operational and support services for advanced information technology systems, cloud systems, and supercomputers supporting non-classified US research, with an average budget for XSEDE of US$20M+ per year over the period studied (2014–2021). To assess the financial effectiveness of these services, we calculated a proxy for ROI, and converted quantitative measures of XSEDE service delivery into financial values using costs for service from the US marketplace. We calculated two estimates of ROI: a Conservative Estimate, functioning as a lower bound and using publicly available data for a lower valuation of XSEDE services; and a Best Available Estimate, functioning as a more accurate estimate, but using some unpublished valuation data. Using the largest dataset assembled for analysis of ROI for a cyberinfrastructure project, we found a Conservative Estimate of ROI of 1.87, and a Best Available Estimate of ROI of 3.24. Through accounting methods, we show that XSEDE services offer excellent value to the US government, that the services offered uniquely by XSEDE (that is, not otherwise available for purchase) were the most valuable to the facilitation of US research activities, and that accounting-based concepts hold great value for understanding the mechanisms of scientific research generally.
This paper explores the financial effectiveness of a national advanced computing support organization within the United States (US) called the eXtreme Science and Engineering Discovery Environment (XSEDE). XSEDE was funded by the National Science Foundation (NSF) in 2011 to manage delivery of advanced computing support to researchers in the US working on non-classified research. In this paper, we describe the methodologies employed to calculate the return on investment (ROI) for governmental expenditures on XSEDE and present a lower bound on the US government’s ROI for XSEDE from 2014 to 2020. For each year of the XSEDE project considered, XSEDE delivered measurable value to the US that exceeded the cost incurred by the Federal Government to fund XSEDE. That is, the US Federal Government’s ROI for XSEDE is at least 1 each year. Over the course of the study period, the ROI for XSEDE rose from 0.99 to 1.78. This increase was due partly to our ability to assign a value to more and more of XSEDE’s services over time and partly to the value of certain XSEDE services increasing over time. From 2014 to 2020, XSEDE offered an ROI of more than $1.5 in value for every $1.0 invested by the US Federal Government. Because our estimations were very conservative, this figure represents the lower bound of the value created by XSEDE. The most important part of ”returns” created by XSEDE are the actual outcomes it enables in terms of education, enabling new discoveries, and supporting the creation of new inventions that improve quality of life. In future work we will use newly developed accounting methodologies to begin assessing the value of the outcomes of XSEDE.
Ticks from the genus Rhipicephalus have enormous global economic impact as ectoparasites of cattle. Rhipicephalus microplus and Rhipicephalus annulatus are known to harbor infectious pathogens such as Babesia bovis, Babesia bigemina, and Anaplasma marginale. Having reference quality genomes of these ticks would advance research to identify druggable targets for chemical entities with acaricidal activity and refine anti-tick vaccine approaches. We sequenced and assembled the genomes of R. microplus and R. annulatus, using Pacific Biosciences and HiSeq 4000 technologies on very high molecular weight genomic DNA. We used 22 and 29 SMRT cells on the Pacific Biosciences Sequel for R. microplus and R. annulatus, respectively, and 3 lanes of the Illumina HiSeq 4000 platform for each tick. The PacBio sequence yields for R. microplus and R. annulatus were 21.0 and 27.9 million subreads, respectively, which were assembled with Canu v. 1.7. The final Canu assemblies consisted of 92,167 and 57,796 contigs with an average contig length of 39,249 and 69,055 bp for R. microplus and R. annulatus, respectively. Annotated genome quality was assessed by BUSCO analysis to provide quantitative measures for each assembled genome. Over 82% and 92% of the 1066 member BUSCO gene set was found in the assembled genomes of R. microplus and R. annulatus, respectively. For R. microplus, only 189 of the 1066 BUSCO genes were missing and only 140 were present in a fragmented condition. For R. annulatus, only 75 of the BUSCO genes were missing and only 109 were present in a fragmented condition. The raw sequencing reads and the assembled contigs/scaffolds are archived at the National Center for Biotechnology Information.
The longhorned tick, Haemaphysalis longicornis, feeds upon a wide range of bird and mammalian hosts. Mammalian hosts include cattle, deer, sheep, goats, humans, and horses. This tick is known to transmit a number of pathogens causing tick-borne diseases, and was the vector of a recent serious outbreak of oriental theileriosis in New Zealand. A New Zealand-USA consortium was established to sequence, assemble, and annotate the genome of this tick, using ticks obtained from New Zealand's North Island. In New Zealand, the tick is considered exclusively parthenogenetic and this trait was deemed useful for genome assembly. Very high molecular weight genomic DNA was sequenced on the Illumina HiSeq4000 and the long-read Pac Bio Sequel platforms. Twenty-eight SMRT cells produced a total of 21.3 million reads which were assembled with Canu on a reserved supercomputer node with access to 12 TB of RAM, running continuously for over 24 days. The final assembly dataset consisted of 34,211 contigs with an average contig length of 215,205 bp. The quality of the annotated genome was assessed by BUSCO analysis, an approach that provides quantitative measures for the quality of an assembled genome. Over 95% of the BUSCO gene set was found in the assembled genome. Only 48 of the 1066 BUSCO genes were missing and only 9 were present in a fragmented condition. The raw sequencing reads and the assembled contigs/scaffolds are archived at the National Center for Biotechnology Information.
We have demonstrated the synergy of a wide-area SLASH2 file system [1] with remote bioinformatics workflows between Extreme Science and Engineering Discovery Environment [2] sites using the Galaxy Project's web-based platform [3] for reproducible data analysis. Wide-area Galaxy workflows were enabled by establishing a geographically-distributed SLASH2 instance between the Greenfield [4] system at Pittsburgh Supercomputing Center [5] and virtual machines incorporating storage within the Corral [6] file system at the Texas Advanced Computing Center [7]. Analysis tasks submitted through a single Galaxy instance seamlessly leverage data available from either site. In this paper, we explore the advantages of SLASH2 for enabling workflows from Galaxy Main [8].
One of the challenges to adoption of HPC is the disjunction between those who need it and those who know it. Biology (specifically, genomics) is a growing field for computational use, but the typical biologist does not have an established informatics background. The National Center for Genome Analysis Support (NCGAS) aids users in getting past the initial shock of the command line and guides them toward savvy cluster use. NCGAS is initiating a push to become domain champions alongside Oklahoma State's Brian Cougar. Our position at IU gives us a close relationship with XSEDE and we already fulfill a role in pushing users toward XSEDE resources when our local clusters are ill-suited to the job. We currently act as liaison between biologists and Jetstream, IU and TACC's research computing cloud. Typical issues include: Software installation; Software usage - what parameters do I choose, and how do I interpret the results; Batch job submission; Understanding how queues and job handlers work; Data movement, Spinning up VMs on Jetstream We will discuss how we have structured our support, and illustrate our impact on XSEDE resources.
Reconstructing the genome and transcriptome for a new or extant species are essential steps in expanding our understanding of the organism’s active RNA landscape and gene regulatory dynamics, as well as for developing therapeutic targets to fight disease. The advancement of sequencing technologies has paved the way to generate high-quality draft transcriptomes. With many possible approaches available to accomplish this task, there is a need for a closer investigation of the factors that influence the quality of the results. We carried out an extensive survey of variety of elements that are important in transcriptome assembly. We utilized the human RNA-Seq data from the Sequencing Quality Control Consortium (SEQC) as a well-characterized and comprehensive resource with an available, well-studied human reference genome. Our results indicate that the quality of the library construction significantly impacts the quality of the assembly. Higher coverage of the genome is not as important as the quality of the input RNA-Seq data. Thus, once a certain coverage is attained, the quality of the assembly is mainly dependent on the base-calling accuracy of the input sequencing reads; and it is important to avoid saturating the assembler with extra coverage.
In metagenome analysis, computational methods for assembly, taxonomic profiling and binning are key components facilitating downstream biological data interpretation. However, a lack of consensus about benchmarking datasets and evaluation metrics complicates proper performance assessment. The Critical Assessment of Metagenome Interpretation (CAMI) challenge has engaged the global developer community to benchmark their programs on datasets of unprecedented complexity and realism. Benchmark metagenomes were generated from ~700 newly sequenced microorganisms and ~600 novel viruses and plasmids, including genomes with varying degrees of relatedness to each other and to publicly available ones and representing common experimental setups. Across all datasets, assembly and genome binning programs performed well for species represented by individual genomes, while performance was substantially affected by the presence of related strains. Taxonomic profiling and binning programs were proficient at high taxonomic ranks, with a notable performance decrease below the family level. Parameter settings substantially impacted performances, underscoring the importance of program reproducibility. While highlighting current challenges in computational metagenomics, the CAMI results provide a roadmap for software selection to answer specific research questions.
It is crucial to understand the performance of transcriptome assemblies to improve current practices. Investigating the factors that affect a transcriptome assembly is very important and is the primary goal of our project. To that end, we designed a multi-step pipeline consisting of variety of pre-processing and quality control steps. XSEDE allocations enabled us to achieve the computational demands of the project. The high memory Blacklight and Green field systems at Pittsburgh Supercomputing Center were essential to accomplish multiple steps of this project. This paper presents the computational aspects of our comprehensive transcriptome assembly and validation study.
BACKGROUND:The Cancer Genome Atlas Project (TCGA) is a National Cancer Institute effort to profile at least 500 cases of 20 different tumor types using genomic platforms and to make these data, both raw and processed, available to all researchers. TCGA data are currently over 1.2 Petabyte in size and include whole genome sequence (WGS), whole exome sequence, methylation, RNA expression, proteomic, and clinical datasets. Publicly accessible TCGA data are released through public portals, but many challenges exist in navigating and using data obtained from these sites. We developed TCGA Expedition to support the research community focused on computational methods for cancer research. Data obtained, versioned, and archived using TCGA Expedition supports command line access at high-performance computing facilities as well as some functionality with third party tools. For a subset of TCGA data collected at University of Pittsburgh, we also re-associate TCGA data with de-identified data from the electronic health records. Here we describe the software as well as the architecture of our repository, methods for loading of TCGA data to multiple platforms, and security and regulatory controls that conform to federal best practices.RESULTS:TCGA Expedition software consists of a set of scripts written in Bash, Python and Java that download, extract, harmonize, version and store all TCGA data and metadata. The software generates a versioned, participant- and sample-centered, local TCGA data directory with metadata structures that directly reference the local data files as well as the original data files. The software supports flexible searches of the data via a web portal, user-centric data tracking tools, and data provenance tools. Using this software, we created a collaborative repository, the Pittsburgh Genome Resource Repository (PGRR) that enabled investigators at our institution to work with all TCGA data formats, and to interrogate these data with analysis pipelines, and associated tools. WGS data are especially challenging for individual investigators to use, due to issues with downloading, storage, and processing; having locally accessible WGS BAM files has proven invaluable.CONCLUSION:Our open-source, freely available TCGA Expedition software can be used to create a local collaborative infrastructure for acquiring, managing, and analyzing TCGA data and other large public datasets.
Jonathan C. Silverstein合作论文数NorthShore University HealthSystem5