Genomic data sharing increases our understanding of factors that influence health and diseases by enabling additional research questions from secondary data users, increasing statistical power through combining multiple data sources, facilitating reproducibility and validation of research results, and supporting innovation with the development of research tools and methodologies. Research using large-scale genomic and other-omics data is no longer limited by storage space or slow downloading speed for individual datasets. Innovations like cloud infrastructure enable computing over many datasets at multiple locations at once. These developments have increased the need for faster, more efficient processes for data access and sharing. To support scientific exploration and meet the demand for analysis of biomedical data, the National Cancer Institute (NCI) has created two large collections of the broad-use studies within the database of Genotypes and Phenotypes (dbGaP). These collections comply with the consent of the study (data use limitations) for how secondary access to studies are determined. NCI’s Collection of Datasets for General Research Use comprises 284 studies with individual-level data sets, and permits approved users to explore broad research interest, including methods and tool development. NCI’s Collection of Datasets for Health, Medical, and Biomedical (HMB) Research Purposes is comprised of 65 studies of individual-level that are permitted for research interests specific to any health, medical, or biomedical research only. The HMB collection could be used for methods and tool development; research interests involving ancestry/populations studies must be dependent on a health/medical condition. Through these collections, investigators will have the potential to add access to 349 datasets to their approved research projects or submit a new project request for access to the broad-use collections. As new genomic studies are registered by NCI for release through dbGaP, they will be automatically added to these collections. NCI anticipates through the implementation of these broad-use collections a requestor could have access to approximately 70% of NCI’s studies in controlled-access repositories.The streamlined access to broad-use datasets expedites data sharing and potentially accelerates the discovery process. This approach reduces redundancies for obtaining controlled-access data. Citation Format: Freddie L. Pruitt, Michael Feolo, Subhashini Jagu, Jaime Guidry Auvil. NCI’s broad-use collections: Accelerating discovery process by improving access to individual-level genomics and other -omics data [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2023; Part 2 (Clinical Trials and Late-Breaking Research); 2023 Apr 14-19; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2023;83(8_Suppl):Abstract nr LB073.
Identifying relevant studies and harmonizing datasets are major hurdles for data reuse. Common Data Elements (CDEs) can help identify comparable study datasets and reduce the burden of retrospective data harmonization, but they have not been required, historically. The collaborative team at PhenX and dbGaP developed an approach to use PhenX variables as a set of CDEs to link phenotypic data and identify comparable studies in dbGaP. Variables were identified as either comparable or related, based on the data collection mode used to harmonize data across mapped datasets. We further added a CDE data field in the dbGaP data submission packet to indicate use of PhenX and annotate linkages in the future. Some 13,653 dbGaP variables from 521 studies were linked through PhenX variable mapping. These variable linkages have been made accessible for browsing and searching in the repository through dbGaP CDE-faceted search filter and the PhenX variable search tool. New features in dbGaP and PhenX enable investigators to identify variable linkages among dbGaP studies and reveal opportunities for cross-study analysis.
To create a scientific resource of expression quantitative trail loci (eQTL), we conducted a genome-wide association study (GWAS) using genotypes obtained from whole genome sequencing (WGS) of DNA and gene expression levels from RNA sequencing (RNA-seq) of whole blood in 2622 participants in Framingham Heart Study. We identified 6,778,286 cis -eQTL variant-gene transcript (eGene) pairs at p < 5 × 10 –8 (2,855,111 unique cis -eQTL variants and 15,982 unique eGenes) and 1,469,754 trans -eQTL variant-eGene pairs at p < 1e−12 (526,056 unique trans -eQTL variants and 7233 unique eGenes). In addition, 442,379 cis -eQTL variants were associated with expression of 1518 long non-protein coding RNAs (lncRNAs). Gene Ontology (GO) analyses revealed that the top GO terms for cis- eGenes are enriched for immune functions (FDR < 0.05). The cis -eQTL variants are enriched for SNPs reported to be associated with 815 traits in prior GWAS, including cardiovascular disease risk factors. As proof of concept, we used this eQTL resource in conjunction with genetic variants from public GWAS databases in causal inference testing (e.g., COVID-19 severity). After Bonferroni correction, Mendelian randomization analyses identified putative causal associations of 60 eGenes with systolic blood pressure, 13 genes with coronary artery disease, and seven genes with COVID-19 severity. This study created a comprehensive eQTL resource via BioData Catalyst that will be made available to the scientific community. This will advance understanding of the genetic architecture of gene expression underlying a wide range of diseases.
Inferring subject ancestry using genetic data is an important step in genetic association studies, required for dealing with population stratification. It has become more challenging to infer subject ancestry quickly and accurately since large amounts of genotype data, collected from millions of subjects by thousands of studies using different methods, are accessible to researchers from repositories such as the database of Genotypes and Phenotypes (dbGaP) at the National Center for Biotechnology Information (NCBI). Study-reported populations submitted to dbGaP are often not harmonized across studies or may be missing. Widely-used methods for ancestry prediction assume that most markers are genotyped in all subjects, but this assumption is unrealistic if one wants to combine studies that used different genotyping platforms. To provide ancestry inference and visualization across studies, we developed a new method, GRAF-pop, of ancestry prediction that is robust to missing genotypes and allows researchers to visualize predicted population structure in color and in three dimensions. When genotypes are dense, GRAF-pop is comparable in quality and running time to existing ancestry inference methods EIGENSTRAT, FastPCA, and FlashPCA2, all of which rely on principal components analysis (PCA). When genotypes are not dense, GRAF-pop gives much better ancestry predictions than the PCA-based methods. GRAF-pop employs basic geometric and probabilistic methods; the visualized ancestry predictions have a natural geometric interpretation, which is lacking in PCA-based methods. Since February 2018, GRAF-pop has been successfully incorporated into the dbGaP quality control process to identify inconsistencies between study-reported and computationally predicted populations and to provide harmonized population values in all new dbGaP submissions amenable to population prediction, based on marker genotypes. Plots, produced by GRAF-pop, of summary population predictions are available on dbGaP study pages, and the software, is available at https://www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/Software.cgi.
Genome-wide association studies (GWAS) usually rely on the assumption that different samples are not from closely related individuals. Detection of duplicates and close relatives becomes more difficult both statistically and computationally when one wants to combine datasets that may have been genotyped on different platforms. The dbGaP repository at the National Center of Biotechnology Information (NCBI) contains datasets from hundreds of studies with over one million samples. There are many duplicates and closely related individuals both within and across studies from different submitters. Relationships between studies cannot always be identified by the submitters of individual datasets. To aid in curation of dbGaP, we developed a rapid statistical method called Genetic Relationship and Fingerprinting (GRAF) to detect duplicates and closely related samples, even when the sets of genotyped markers differ and the DNA strand orientations are unknown. GRAF extracts genotypes of 10,000 informative and independent SNPs from genotype datasets obtained using different methods, and implements quick algorithms that enable it to find all of the duplicate pairs from more than 880,000 samples within and across dbGaP studies in less than two hours. In addition, GRAF uses two statistical metrics called All Genotype Mismatch Rate (AGMR) and Homozygous Genotype Mismatch Rate (HGMR) to determine subject relationships directly from the observed genotypes, without estimating probabilities of identity by descent (IBD), or kinship coefficients, and compares the predicted relationships with those reported in the pedigree files. We implemented GRAF in a freely available C++ program of the same name. In this paper, we describe the methods in GRAF and validate the usage of GRAF on samples from the dbGaP repository. Other scientists can use GRAF on their own samples and in combination with samples downloaded from dbGaP.
Background Identification of single nucleotide polymorphisms (SNPs) associated with gene expression levels, known as expression quantitative trait loci (eQTLs), may improve understanding of the functional role of phenotype-associated SNPs in genome-wide association studies (GWAS). The small sample sizes of some previous eQTL studies have limited their statistical power. We conducted an eQTL investigation of microarray-based gene and exon expression levels in whole blood in a cohort of 5257 individuals, exceeding the single cohort size of previous studies by more than a factor of 2. Results We detected over 19,000 independent lead cis -eQTLs and over 6000 independent lead trans -eQTLs, targeting over 10,000 gene targets (eGenes), with a false discovery rate (FDR) < 5%. Of previously published significant GWAS SNPs, 48% are identified to be significant eQTLs in our study. Some trans -eQTLs point toward novel mechanistic explanations for the association of the SNP with the GWAS-related phenotype. We also identify 59 distinct blocks or clusters of trans -eQTLs, each targeting the expression of sets of six to 229 distinct trans -eGenes. Ten of these sets of target genes are significantly enriched for microRNA targets (FDR < 5%). Many of these clusters are associated in GWAS with multiple phenotypes. Conclusions These findings provide insights into the molecular regulatory patterns involved in human physiology and pathophysiology. We illustrate the value of our eQTL database in the context of a recent GWAS meta-analysis of coronary artery disease and provide a list of targeted eGenes for 21 of 58 GWAS loci.
The National Center for Biotechnology Information (NCBI) provides a large suite of online resources for biological information and data, including the GenBank((R)) nucleic acid sequence database and the PubMed database of citations and abstracts for published life science journals. The Entrez system provides search and retrieval operations for most of these data from 37 distinct databases. The E-utilities serve as the programming interface for the Entrez system. Augmenting many of the Web applications are custom implementations of the BLAST program optimized to search specialized data sets. New resources released in the past year include iCn3D, MutaBind, and the Antimicrobial Resistance Gene Reference Database; and resources that were updated in the past year include My Bibliography, SciENcv, the Pathogen Detection Project, Assembly, Genome, the Genome Data Viewer, BLAST and PubChem. All of these resources can be accessed through the NCBI home page at www.ncbi.nlm.nih.gov.
The Alzheimer's Disease Sequencing Project (ADSP) was announced in February 2012 by National Institutes of Health (NIH) to sequence the genomes of a large number of well-characterized individuals in order to identify a broad range of AD risk and protective gene variants, with the ultimate goal of facilitating the identification of new pathways for therapeutic approaches and prevention. To better facilitate the community's access the ADSP data, the NIA Genetics of Alzheimer's Disease Data Storage Site (NIAGADS) collaborated with NIH Database of Genotypes and Phenotypes (dbGaP) and Sequencing Read Archive (SRA) and developed the web-based ADSP Data Portal. Investigators with data access approval from NIH can log in to the portal website via NIH iTrust authentication. The portal features a dynamic HTML user interface and allows the user review study design, data quality metrics, and select sequencing data to download with various filtering criteria such as cohort, diagnosis, APOE genotype, and sequencing depth. The user then checks out selected data in a shopping cart file. This cart file allows the user to retrieve data for the selected samples from dbGaP and SRA using the SRA Tools software. As of January 2014, whole genome sequencing for 410 individuals from multiplex families are available to qualified investigators. All ADSP whole-exome and whole-genome sequencing data will be complete in summer 2014. The ADSP website and data portal can be reached at https://www.niagads.org/adsp or http://www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs000572.v1.p1.
The 1000 Genomes Project aims to provide a deep characterization of human genome sequence variation by sequencing at a level that should allow the genome-wide detection of most variants with frequencies as low as 1%. However, in the major histocompatibility complex (MHC), only the top 10 most frequent haplotypes are in the 1% frequency range whereas thousands of haplotypes are present at lower frequencies. Given the limitation of both the coverage and the read length of the sequences generated by the 1000 Genomes Project, the highly variable positions that define HLA alleles may be difficult to identify. We used classical Sanger sequencing techniques to type the HLA-A, HLA-B, HLA-C, HLA-DRB1 and HLA-DQB1 genes in the available 1000 Genomes samples and combined the results with the 103,310 variants in the MHC region genotyped by the 1000 Genomes Project. Using pairwise identity-by-descent distances between individuals and principal component analysis, we established the relationship between ancestry and genetic diversity in the MHC region. As expected, both the MHC variants and the HLA phenotype can identify the major ancestry lineage, informed mainly by the most frequent HLA haplotypes. To some extent, regions of the genome with similar genetic or similar recombination rate have similar properties. An MHC-centric analysis underlines departures between the ancestral background of the MHC and the genome-wide picture. Our analysis of linkage disequilibrium (LD) decay in these samples suggests that overestimation of pairwise LD occurs due to a limited sampling of the MHC diversity. This collection of HLA-specific MHC variants, available on the dbMHC portal, is a valuable resource for future analyses of the role of MHC in population and disease studies.
NCBI maintains information about genes primarily in two contexts. One context is defined by public sequence information, such as annotation of RefSeqs (see RefSeq chapter) or linking with records in the International Nucleotide Sequence Database Consortium or INSDC (see Genome Reference Consortium chapter). Connection of sequence information to a GeneID or a UniGene cluster identifier is critical to any analysis of gene expression. The second context for defining a gene is by mapped phenotype. GeneIDs are not assigned to mapped loci for all taxa, but when they are, the expectation is that the genes will eventually be connected to sequence as the molecular basis for the phenotype is defined.
Nature Reviews Genetics 12, 730–736 (2011) In the above article, the incorrect link was provided for GWAS Central. The correct link should have been http://www.gwascentral.org. In the Further Information Box, the link to http://gwas.nih.gov was incorrectly described as 'GWAS Central (includes policy)'.
Tissue AntigensVolume 75, Issue 3 p. 199-200 The Babel Tower revisited: SNPs – Indels – CNVs. Confusion in naming sequence variant always rises from ashes P. A. Gourraud, Corresponding Author P. A. Gourraud Department of Neurology, University of California, San Francisco, CA, USAPierre-Antoine GourraudDepartment of NeurologyUniversity of California513 Parnassus AvenueSan Francisco, California 94143USATel: +1 415 476 3136Fax: +1 415 476 5229e-mail: [email protected]Search for more papers by this authorM. Feolo, M. Feolo National Center for Biotechnology Information, Bethesda, MD, USASearch for more papers by this author P. A. Gourraud, Corresponding Author P. A. Gourraud Department of Neurology, University of California, San Francisco, CA, USAPierre-Antoine GourraudDepartment of NeurologyUniversity of California513 Parnassus AvenueSan Francisco, California 94143USATel: +1 415 476 3136Fax: +1 415 476 5229e-mail: [email protected]Search for more papers by this authorM. Feolo, M. Feolo National Center for Biotechnology Information, Bethesda, MD, USASearch for more papers by this author First published: 05 February 2010 https://doi.org/10.1111/j.1399-0039.2009.01424.xCitations: 2Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onFacebookTwitterLinkedInRedditWechat No abstract is available for this article.Citing Literature Volume75, Issue3March 2010Pages 199-200 RelatedInformation
This chapter provides an introduction to the major, freely available, Internet-accessible databases in human and medical genetics used by healthcare providers in the diagnosis, management, and genetic Counseling of persons with inherited disorders and their families, as well as by researchers for gene discovery, recording allelic variants, and cataloging genotype-phenotype relationships. Databases discussed include: GeneTests (view: www.genetests.org); Online Mendelian Inheritance in Man (view: www.ncbi.nlm.nih.gov/Omim); locus specific databases (LSDBs) identified at the Human Genome Variation Society (HGVS) web site (http://www.HGVS.org/dblist.html); DatabasE of Chromosome Imbalance and Phenotype in Humans using Ensembl Resources (view: http://decipher.sanger.ac.uk); Entrez Gene (view: ncbi.nlm.nih.gov/gene); dbGap: Database of Genotype and Phenotype (view: ncbi.nlm.nih.gov/dbgap); and the Human Gene Mutation Database HGMD (R) (view: http://www.hgmd.org).
We describe a novel approach to genetic association analyses with proteins sub-divided into biologically relevant smaller sequence features (SFs), and their variant types (VTs). SFVT analyses are particularly informative for study of highly polymorphic proteins such as the human leukocyte antigen (HLA), given the nature of its genetic variation: the high level of polymorphism, the pattern of amino acid variability, and that most HLA variation occurs at functionally important sites, as well as its known role in organ transplant rejection, autoimmune disease development and response to infection. Further, combinations of variable amino acid sites shared by several HLA alleles (shared epitopes) are most likely better descriptors of the actual causative genetic variants. In a cohort of systemic sclerosis patients/controls, SFVT analysis shows that a combination of SFs implicating specific amino acid residues in peptide binding pockets 4 and 7 of HLA-DRB1 explains much of the molecular determinant of risk.