Abstract Purpose: The Human Tumor Atlas Network (HTAN) sets out to map the cellular and molecular architecture of human tumors over space and time to advance precision oncology. The HTAN Data Coordinating Center (DCC) underpins this mission by standardizing, integrating, and distributing multimodal data; ensuring legacy and impact. Methods: The DCC developed a scalable cloud ecosystem (Synapse.org, Google BigQuery, custom data portal) for data ingestion, governance, validation, and dissemination. A community-driven process, aligned with NCI standards, produced a consistent metadata schema that ensures interoperability across rich clinical and biospecimen information and all data modalities. The HTAN data portal provides a single landing page for users of HTAN data featuring filter-based search, visualization of imaging and single-cell datasets and detailed documentation. Data are disseminated via a tiered model, with imaging and dbGaP-controlled access sequencing data available from NCI Cancer Research Data Commons General Commons and open-access processed data on Synapse.org. Current work transitions to a modular LinkML data model, adds an AI-assisted curation interface, and includes a streamlined medallion architecture in BigQuery and portal enhancements. Results: The first five years of HTAN (release v7.0) produced 334 TB of multimodal data (0.23M files) from 2,372 cases and 11,378 biospecimens, spanning >60 disease types and >25 assays. HTAN supports >3,400 unique global users a month. The most frequently accessed data includes scRNA-seq (1075 downloading users), multiplex imaging (197), and spatial transcriptomics (419). Analysis of dbGaP requests to-date (N=127) shows broad academic (75%) and industry use with research themes focused on genome instability (115 requests), immune evasion (52), and metastasis (49), often with multi-omic and AI-driven approaches. Conclusions: HTAN has established a globally utilized, harmonized foundation for spatially resolved, multimodal cancer research. This enduring platform for community-driven discovery is built on robust data management and transparent access models. The next phase deepens integration of spatial and clinical data, expanding AI-readiness, and strengthening support for reuse and reproducibility. We invite researchers to engage with this resource. ChatGPT 5, Gemini 2.5 Pro, & Claude Sonnet 4.5 were used to summarize usage metrics, conduct thematic analysis of dbGaP applications, and initial abstract drafting. All content was evaluated and approved by the authors. Citation Format: Adam Taylor, Ashley Clayton, Aditi Gopalan, Milen Nikolov, Thomas Yu, David Gibbs, Yamina Katariya, Dar'ya Pozhidayeva, Ino de Bruijn, Seluck Onur Sumer, Kristen Anton, Jennifer Altreuter, Alex Lash, Ethan Cerami, Nikolaus Schultz, Vesteinn Thorsson. Reading the map: An invitation to the resources of the Human Tumor Atlas Network [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 4169.
Cancer research increasingly relies on large-scale, multimodal datasets that capture the complexity of tumor ecosystems across diverse patients, cancer types, and disease stages. The Human Tumor Atlas Network (HTAN) generates such data, including single-cell transcriptomics, proteomics, and multiplexed imaging. However, the volume and heterogeneity of the data present challenges for researchers seeking to integrate, explore, and analyze these datasets at scale. To this end, HTAN developed a cloud-based infrastructure that transforms clinical and assay metadata into aggregate Google BigQuery tables, hosted through the Institute for Systems Biology Cancer Gateway in the Cloud (ISB-CGC). This infrastructure introduces two key innovations: (1) a provenance-based HTAN ID table that simplifies cohort construction and cross-assay integration, and (2) the novel adaptation of BigQuery's geospatial functions for use in spatial biology, enabling neighborhood and correlation analysis of tumor microenvironments. We demonstrate these capabilities through R and Python notebooks that highlight use cases such as identifying precancer and organ-specific sample cohorts, integrating multimodal datasets, and analyzing single-cell and spatial data. By lowering technical and computational barriers, this infrastructure provides a cost-effective and intuitive entry point for researchers, highlighting the potential of cloud-based platforms to accelerate cancer discoveries.
The exploration of genotypic variants impacting phenotypes is a cornerstone in genetics research. The emergence of vast collections containing deeply genotyped and phenotyped families has made it possible to pursue the search for variants associated with complex diseases. However, managing these large-scale data sets requires specialized computational tools to organize and analyze the extensive data. Genotypes and Phenotypes in Families (GPF) is an open-source platform that manages genotypes and phenotypes derived from collections of families. GPF allows interactive exploration of genetic variants, enrichment analysis for de novo mutations, phenotype/genotype association tools, and secure data sharing. GPF is used to disseminate two family collection data sets, SSC and SPARK, for the study of autism, built by the Simons Foundation. The GPF instance at the Simons Foundation (GPF-SFARI) provides protected access to comprehensive genotypic and phenotypic data for SSC and SPARK. GPF-SFARI also provides public access to an extensive collection of de novo mutations from individuals with autism and related disorders and to gene-level statistics of the protected data sets characterizing the genes' roles in autism. However, GPF is versatile and can manage genotypic data from other small or large family collections. Here, we highlight the primary features of GPF within the context of GPF-SFARI.
Data from the first phase of the Human Tumor Atlas Network (HTAN) are now available, comprising 8,425 biospecimens from 2,042 research participants profiled with more than 20 molecular assays. The data were generated to study the evolution from precancerous to advanced disease. The HTAN Data Coordinating Center (DCC) has enabled their dissemination and effective reuse. We describe the diverse datasets, how to access them, data standards, underlying infrastructure and governance approaches, and our methods to sustain community engagement. HTAN data can be accessed through the HTAN Portal, explored in visualization tools-including CellxGene, Minerva and cBioPortal-and analyzed in the cloud through the NCI Cancer Research Data Commons. Infrastructure was developed to enable data ingestion and dissemination through the Synapse platform. The HTAN DCC's flexible and modular approach to sharing complex cancer research data offers valuable insights to other data-coordination efforts and researchers looking to leverage HTAN data.
The exploration of genotypic variants impacting phenotypes is a cornerstone in genetics research. The emergence of vast collections containing deeply genotyped and phenotyped families has made it possible to pursue the search for variants associated with complex diseases. However, managing these large-scale datasets requires specialized computational tools tailored to organize and analyze the extensive data. GPF (Genotypes and Phenotypes in Families) is an open-source platform ( https://github.com/iossifovlab/gpf ) that manages genotypes and phenotypes derived from collections of families. The GPF interface allows interactive exploration of genetic variants, enrichment analysis for de novo mutations, and phenotype/genotype association tools. In addition, GPF allows researchers to share their data securely with the broader scientific community. GPF is used to disseminate two large-scale family collection datasets (SSC, SPARK) for the study of autism funded by the SFARI foundation. However, GPF is versatile and can manage genotypic data from other small or large family collections. Our GPF-SFARI GPF instance ( https://gpf.sfari.org/ ) provides protected access to comprehensive genotypic and phenotypic data for the SSC and SPARK. In addition, GPF-SFARI provides public access to an extensive collection of de novo mutations identified in individuals with autism and related disorders and to gene-level statistics of the protected datasets characterizing the genes’ roles in autism. Here, we highlight the primary features of GPF within the context of GPF-SFARI.
Supplementary Figures 1-3, Tables 1-9, Methods from Genomic and Biological Characterization of Exon 4 KRAS Mutations in Human Cancer
The volume of genomics and health data is growing rapidly, driven by sequencing for both research and clinical use. However, under current practices, the data is fragmented into many distinct datasets, and researchers must go through a separate application process for each dataset. This is time-consuming both for the researchers and the data stewards, and it reduces the velocity of research and new discoveries that could improve human health. We propose to simplify this process, by introducing a standard Library Card that identifies and authenticates researchers across all participating datasets. Each researcher would only need to apply once to establish their bona fides as a qualified researcher, and could then use the Library Card to access a wide range of datasets that use a compatible data access policy and authentication protocol.
In individuals with autism spectrum disorder (ASD), de novo mutations have previously been shown to be significantly correlated with lower IQ but not with the core characteristics of ASD: deficits in social communication and interaction and restricted interests and repetitive patterns of behavior. We extend these findings by demonstrating in the Simons Simplex Collection that damaging de novo mutations in ASD individuals are also significantly and convincingly correlated with measures of impaired motor skills. This correlation is not explained by a correlation between IQ and motor skills. We find that IQ and motor skills are distinctly associated with damaging mutations and, in particular, that motor skills are a more sensitive indicator of mutational severity than is IQ, as judged by mutational type and target gene. We use this finding to propose a combined classification of phenotypic severity: mild (little impairment of either), moderate (impairment mainly to motor skills), and severe (impairment of both IQ and motor skills).
In individuals with Autism Spectrum Disorder (ASD), de novo mutations have previously been shown to be significantly correlated with lower IQ, but not with the core characteristics of ASD: deficits in social communication and interaction, and restricted interests and repetitive patterns of behavior. We extend these findings by demonstrating in the Simons Simplex Collection that damaging de novo mutations in ASD individuals are also significantly and convincingly correlated with measures of impaired motor skills. This correlation is not explained by a correlation between IQ and motor skills. We find that IQ and motor skills are distinctly associated with damaging mutations and, in particular, that motor skills are a more sensitive indicator of mutational severity, as judged by the type and its gene target. We use this finding to propose a combined classification of phenotypic severity: mild (little impairment of both), moderate (impairment mainly to motor skills) and severe (impairment of both).
Introduction: Many patients with lung cancers cannot receive platinum-containing regimens owing to comorbid medical conditions. We designed the PPB (paclitaxel, pemetrexed, and bevacizumab) regimen to maintain or improve outcomes while averting the unique toxicities of platinum-based chemotherapies.Methods: We enrolled patients with untreated, advanced lung adenocarcinomas with measurable disease and no contraindications to bevacizumab. Participants received paclitaxel, 90 mg/m(2), pemetrexed, 500 mg/m(2), and bevacizumab, 10 mg/kg, every 14 days for 6 months and continued to receive pemetrexed and bevacizumab every 14 days until progression or unacceptable toxicity.Results: Of the 44 patients treated, 50% were women; the median age was 61 years and 89% had a Karnofsky performance status of at least 80%. We genotyped 38 patients with the following results: Kirsten rat sarcoma viral oncogene homolog gene (KRAS), 16; anaplastic lymphoma receptor tyrosine kinase gene (ALK), three; B-Raf proto-oncogene, serine/threonine kinase gene (BRAF) V600E, two; erb-b2 receptor tyrosine kinase 2 gene (HER2)/phosphatidylinositol-4,5-bisphosphate 3-kinase catalytic subunit alpha gene (PIK3CA), one; epidermal growth factor receptor gene (EGFR) exon 20 insertion, one; and driver 15, none. A total of 23 patients achieved a PR (52%, 95% confidence interval: 37-68), including seven of 16 with KRAS-mutant tumors. The overall survival rate at 2 years was 43% with a median of 17 months (95% confidence interval: 10-29). Grade 3/4 treatment-related toxicities included elevated alanine transaminase level (16%), fatigue (16%), leukopenia (9%), anemia (7%), elevated aspartate transaminase level (7%), edema (5%), and pleural effusions (5%). Two patients died of respiratory failure without disease progression.Conclusions: The PPB regimen produced a high response rate in patients with lung adenocarcinomas regardless of mutational status. Survival and toxicities were comparable to those in the phase II reports testing platinum-containing doublets with bevacizumab. These results justify use of the PPB regimen in fit patients in whom three-drug regimens including bevacizumab are appropriate. (C) 2016 International Association for the Study of Lung Cancer. Published by Elsevier Inc. All rights reserved.
Autism spectrum disorder (ASD) is a complex neurodevelopmental disorder with a strong genetic basis. Yet, only a small fraction of potentially causal genes-about 65 genes out of an estimated several hundred-are known with strong genetic evidence from sequencing studies. We developed a complementary machine-learning approach based on a human brain-specific gene network to present a genome-wide prediction of autism risk genes, including hundreds of candidates for which there is minimal or no prior genetic evidence. Our approach was validated in a large independent case-control sequencing study. Leveraging these genome-wide predictions and the brain-specific network, we demonstrated that the large set of ASD genes converges on a smaller number of key pathways and developmental stages of the brain. Finally, we identified likely pathogenic genes within frequent autism associated copy-number variants and proposed genes and pathways that are likely mediators of ASD across multiple copy-number variants. All predictions and functional insights are available at http://asd.princeton.edu.
Autism spectrum disorder (ASD) is a range of major neurodevelopmental disabilities with a strong genetic basis. Yet, owing to extensive genetic heterogeneity, multiple modes of inheritance and limited study sizes, sequencing and quantitative genetics approaches have had limited success in characterizing the complex genetics of ASD. Currently, only a small fraction of potentially causal genes—about 65 genes out of an estimated severalhundred—are known based on strong genetic evidence. Hence, there isa critical need for complementary approaches to further characterize the genetic basis of ASD, enabling development of better screening and therapeutics. Here, we use a machine-learning approach based on a human brain-specific functional gene interaction network to present a genome-wide prediction of autism-associated genes, including hundreds of candidate genes for which there is minimal or no prior genetic evidence. Our approach is validated in an independent case-control sequencing study of approximately 2,500families. Leveraging these genome-wide predictions and the brain-specificnetwork, we demonstrate that the large set of ASD genes converges on a smaller number of key cellular pathways and specific developmental stages of the brain. Specifically, integration with spatiotemporal transcriptome expression data implicates early fetal and midfetal stages of the developing human brain in ASD etiology. Likewise, analysis of the connectivity of topautism genes in the brain-specific interaction network reveals the breadthof autism-associated functional modules, processes, and pathways in the brain. Finally, we identify likely pathogenic genes within the most frequent autism-associated copy-number-variants (CNVs) and propose genes and pathways that are likely mediators of autism across multiple CNVs. All the predictions, interactions, and functional insights from this work are available to biomedical researchers at asd.princeton.edu .
Existing methods for interpreting protein variation focus on annotating mutation pathogenicity rather than detailed interpretation of variant deleteriousness and frequently use only sequence-based or structure-based information. We present VIPUR, a computational framework that seamlessly integrates sequence analysis and structural modelling (using the Rosetta protein modelling suite) to identify and interpret deleterious protein variants. To train VIPUR, we collected 9477 protein variants with known effects on protein function from multiple organisms and curated structural models for each variant from crystal structures and homology models. VIPUR can be applied to mutations in any organism's proteome with improved generalized accuracy (AUROC .83) and interpretability (AUPR .87) compared to other methods. We demonstrate that VIPUR's predictions of deleteriousness match the biological phenotypes in ClinVar and provide a clear ranking of prediction confidence. We use VIPUR to interpret known mutations associated with inflammation and diabetes, demonstrating the structural diversity of disrupted functional sites and improved interpretation of mutations associated with human diseases. Lastly, we demonstrate VIPUR's ability to highlight candidate variants associated with human diseases by applying VIPUR to de novo variants associated with autism spectrum disorders.
PURPOSE:In clinical trials, traditional monitoring methods, paper documentation, and outdated collection systems lead to inaccuracies of study information and inefficiencies in the process. Integrated electronic systems offer an opportunity to collect data in real time.PATIENTS AND METHODS:We created a computer software system to collect 13 patient-reported symptomatic adverse events and patient-reported Karnofsky performance status, semi-automated RECIST measurements, and laboratory data, and we made this information available to investigators in real time at the point of care during a phase II lung cancer trial. We assessed data completeness within 48 hours of each visit. Clinician satisfaction was measured.RESULTS:Forty-four patients were enrolled, for 721 total visits. At each visit, patient-reported outcomes (PROs) reflecting toxicity and disease-related symptoms were completed using a dedicated wireless laptop. All PROs were distributed in batch throughout the system within 24 hours of the visit, and abnormal laboratory data were available for review within a median of 6 hours from the time of sample collection. Manual attribution of laboratory toxicities took a median of 1 day from the time they were accessible online. Semi-automated RECIST measurements were available to clinicians online within a median of 2 days from the time of imaging. All clinicians and 88% of data managers felt there was greater accuracy using this system.CONCLUSION:Existing data management systems can be harnessed to enable real-time collection and review of clinical information during trials. This approach facilitates reporting of information closer to the time of events, and improves efficiency, and the ability to make earlier clinical decisions.
e19647 Background: In clinical trials, delayed and disorganized data reporting due to traditional toxicity symptom monitoring methods, paper documentation and out-dated collection systems lead to inaccuracies of critical study information and inefficiencies in the process. Electronic systems offer an opportunity to collect this information in real-time from various sources. The feasibility of such an approach is unknown. Methods: We created a computer software system to collect PROs of symptomatic toxicities, automated RECIST-based tumor response, and lab data, and made this information available to investigators in real-time at the point of care during a phase II lung cancer trial. We assessed data completeness within 48hrs of each visit. Clinician satisfaction was measured. Results: We enrolled 44 patients from 1/23/09 to 09/20/11. There were a total of 725 visits, with mean and median being 13 (range, 1-64) and 10, respectively. At each visit, patients completed self-reports using a dedicated wireless laptop 99.6% of the time; only 3 reports were not completed due to lack of a laptop or technical issues with the institutional Intranet. All PROs were completely distributed in batch throughout the system within 24hrs of the visit, and similarly lab data were available for review with a mean of 26hrs from the time the laboratory received the specimen. Manual attribution to lab toxicities in the system took a median of 1 day. Automated RECIST tumor measurements were available to clinicians online with a median of 2 days from the time of imaging; only 8% obtained >7 days after scans were performed. Throughout the trial, there was improvement in data acquisition times (i.e., a “learning curve”). 89% (16/18) of clinicians and research study assistants felt there was greater accuracy in collection of PROs, radiographic responses and trial data. Conclusions: Existing data management systems can be harnessed to enable real-time collection and review of clinical information during trials. This approach facilitates reporting of information closer to the time of events and may improve accuracy and efficiency, as well as ability to make earlier clinical decisions.
Abstract CA125 antigen is found to be elevated in 85% of patients diagnosed with advanced epithelial ovarian cancer. CA125 has a minimal role as a screening modality. The gene which encodes for the CA125 glycoprotein, MUC16, has been previously identified and sequenced. MUC16 has a molecular size of 22,000 bps and consists of cytoplasmic domain, transmembrane region and 9 external domain tandem repeats. The large molecular weight of this gene has made it difficult to study and track. N terminal signal peptide and transcription factor or factors that triggers MUC16 in ovarian cancer is not well demonstrated. The proximal 114 amino acids of MUC16 can transform NIH/3T3 cells and increase soft agar formation and matrigel invasion. We have designed several constructs from the region about 1000 nucleotides upstream and 300 nucleotides downstream of the MUC16 start Methionine with secretory pMetridia Luciferase Reporter vector (Clonetech, CA). Four constructs of varying lengths had the MUC16 UTR TSS site whereas two did not. Putative binding sites for transcription factors like ER, STAT1, STAT3, CRE-BP1 (ATF2) and NFKB are also present. These constructs were then transfected individually into CA125 negative ovarian cell line, SK-Ov-3, and also into CA125 positive ovarian cell lines, SK-Ov-8 and CAOV3. All cell lines were stably selected with G418 and were used within 10 passages. Western blots confirm the presence of ERalpha, STAT1 and STAT3. Stably transfected cell lines were cultured in low serum conditions with or without 17β-Estradiol or IL-1β, were used to see the effect of transcription. Clinical CA125 levels were also documented for each construct. Secretory pMetridia Luciferase signals were detected in the supernant of the culture cells. There is no difference between untreated or 17β-Estradiol or IL-1β treated cells. However, gel retardation studies with OVCAR3 and SKOV8 wild type cell lines suggested CREB/ATF (CREBZF) transcription factor involvement and confirmatory reporter gene studies are planned. Currently, we are assessing the putative transcription factor binding sites within the 1300 nucleotides region from which we derived the constructs. We conclude that these studies will lead us to better understand the biology of MUC16 gene expression and give us an insight about the role of transcription factors. Citation Format: {Authors}. {Abstract title} [abstract]. In: Proceedings of the 102nd Annual Meeting of the American Association for Cancer Research; 2011 Apr 2-6; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2011;71(8 Suppl):Abstract nr 2033. doi:10.1158/1538-7445.AM2011-2033
Abstract We describe an autosomal-dominant syndrome characterized by multiple non-pigmented, exophytic melanocytic nevi and an increased susceptibility for melanoma, caused by germline mutations in the histone deubiquitinase BAP1. To identify the causative alterations, we performed comprehensive genomic analyses in two unrelated families with numerous dermal nevi composed largely of large, epithelioid melanocytes with abundant amphophilic cytoplasm and large, pleomorphic, vesicular nuclei with prominent nucleoli. Both families each had one proband with uveal melanoma, and three probands in one family had cutaneous melanoma. Array-based comparative genomic hybridization (aCGH) revealed losses of parts of or the entire chromosome 3 in 11 of 22 neoplasms studied. Genotypic analyses revealed that the deletions invariably affected the chromosome from the unaffected parent. Genome partitioning of the minimally deleted region on chromosome 3p21 followed by massively parallel sequencing revealed two different inactivating germline mutations of the BAP1 tumor suppressor gene that in both families segregated with the phenotype. In almost all tumors the remaining wild type BAP1 allele was eliminated by deletion, separate inactivating mutations, or loss of heterozygosity. 35 of 40 nevi (88%) showed mutations in BRAF, while the uveal melanomas had mutations in GNAQ. Our data identify BAP1 as a highly penetrant susceptibility gene for melanocytic neoplasia. Somatic BAP1 mutations have recently been reported in uveal melanoma and linked to the metastatic phenotype. Our observation of frequent bi-allelic inactivation of BAP1 in nevi indicates that the role of BAP1 in melanocytic neoplasia is more complex, and may differ depending on other factors such as the type of melanocyte (uveal or cutaneous) and the co-existing oncogenic mutation. Citation Format: {Authors}. {Abstract title} [abstract]. In: Proceedings of the 102nd Annual Meeting of the American Association for Cancer Research; 2011 Apr 2-6; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2011;71(8 Suppl):Abstract nr LB-125. doi:10.1158/1538-7445.AM2011-LB-125
Annotation of prostate cancer genomes provides a foundation for discoveries that can impact disease understanding and treatment. Concordant assessment of DNA copy number, mRNA expression, and focused exon resequencing in 218 prostate cancer tumors identified the nuclear receptor coactivator NCOA2 as an oncogene in ∼11% of tumors. Additionally, the androgen-driven TMPRSS2-ERG fusion was associated with a previously unrecognized, prostate-specific deletion at chromosome 3p14 that implicates FOXP1, RYBP, and SHQ1 as potential cooperative tumor suppressors. DNA copy-number data from primary tumors revealed that copy-number alterations robustly define clusters of low- and high-risk disease beyond that achieved by Gleason score. The genomic and clinical outcome data from these patients are now made available as a public resource.
Abstract Mutations in RAS proteins occur widely in human cancer. Prompted by the confirmation of KRAS mutation as a predictive biomarker of response to epidermal growth factor receptor (EGFR)–targeted therapies, limited clinical testing for RAS pathway mutations has recently been adopted. We performed a multiplatform genomic analysis to characterize, in a nonbiased manner, the biological, biochemical, and prognostic significance of Ras pathway alterations in colorectal tumors and other solid tumor malignancies. Mutations in exon 4 of KRAS were found to occur commonly and to predict for a more favorable clinical outcome in patients with colorectal cancer. Exon 4 KRAS mutations, all of which were identified at amino acid residues K117 and A146, were associated with lower levels of GTP-bound RAS in isogenic models. These same mutations were also often accompanied by conversion to homozygosity and increased gene copy number, in human tumors and tumor cell lines. Models harboring exon 4 KRAS mutations exhibited mitogen-activated protein/extracellular signal-regulated kinase kinase dependence and resistance to EGFR-targeted agents. Our findings suggest that RAS mutation is not a binary variable in tumors, and that the diversity in mutant alleles and variability in gene copy number may also contribute to the heterogeneity of clinical outcomes observed in cancer patients. These results also provide a rationale for broader KRAS testing beyond the most common hotspot alleles in exons 2 and 3. Cancer Res; 70(14); 5901–11. ©2010 AACR.