Array genotyping is a cost-effective and widely used tool that enables assessment of up to millions of genetic markers in hundreds of thousands of individuals. Genotyping array data are typically highly accurate but sensitive to mixing of DNA samples from multiple individuals before or during genotyping. Contaminated samples can lead to genotyping errors and consequently cause false positive signals or reduce power of association analyses. Here, we propose a new method to identify contaminated samples and the sources of contamination within a genotyping batch. Through analysis of array intensity and genotype data from intentionally mixed samples and 22,366 samples of the Michigan Genomics Initiative, an ongoing biobank-based study, we show that our method can reliably estimate contamination. We also show that identifying sources of contamination can implicate problematic sample processing steps and guide process improvements. Compared to existing methods, our approach can estimate the proportion of contaminating DNA more accurately, eliminate the need for external databases of allele frequencies, and provide contamination estimates that are more robust to the ancestral origin of the contaminating sample.
BACKGROUND It can be useful to know the transgene insertion site in transgenic mice for a variety of reasons, but determining the insertion site generally is a time consuming, expensive, and laborious task. METHODS A simple method is presented to determine transgene insertion sites that combines the enrichment of a sequencing library by polymerase chain reaction (PCR) for sequences containing the transgene, followed by next-generation sequencing of the enriched library. This method was applied to determine the site of integration of the thyroid peroxidase promoter-Cre recombinase mouse transgene that is commonly used to create thyroid-specific gene deletions. RESULTS The insertion site was found to be between bp 12,372,316 and 12,372,324 on mouse chromosome 9, with the nearest characterized genes being Cntn5 and Jrkl, ∼1.5 and 0.9 Mbp from the transgene, respectively. One advantage of knowing a transgene insertion site is that it facilitates distinguishing hemizygous from homozygous transgenic mice. Although this can be accomplished by real-time quantitative PCR, the expected Ct difference is only one cycle, which is challenging to assess accurately. Therefore, the transgene insertion site information was used to develop a 3-primer qualitative PCR assay that readily distinguishes wild type, hemizygous, and homozygous TPO-Cre mice based upon size differences of the wild type and transgenic allele PCR products. CONCLUSIONS Identification of the genomic insertion site of the thyroid peroxidase promoter-Cre mouse transgene should facilitate the use of these mice in studies of thyroid biology.
This article includes supplemental data. Please visit http://www.fasebj.org to obtain this information.Multiple recent publications on RNA sequencing (RNA-seq) have demonstrated the power of next-generation sequencing technologies in whole-transcriptome analysis. Vendor-specific protocols used for RNA library construction often require at least 100 ng total RNA. However, under certain conditions, much less RNA is available for library construction. In these cases, effective transcriptome profiling requires amplification of subnanogram amounts of RNA. Several commercial RNA amplification kits are available for amplification prior to library construction for next-generation sequencing, but these kits have not been comprehensively field evaluated for accuracy and performance of RNA-seq for picogram amounts of RNA. To address this, 4 types of amplification kits were tested with 3 different concentrations, from 5 ng to 50 pg, of a commercially available RNA. Kits were tested at multiple sites to assess reproducibility and ease of use. The human total reference RNA used was spiked with a control pool of RNA molecules in order to further evaluate quantitative recovery of input material. Additional control data sets were generated from libraries constructed following polyA selection or ribosomal depletion using established kits and protocols. cDNA was collected from the different sites, and libraries were synthesized at a single site using established protocols. Sequencing runs were carried out on the Illumina platform. Numerous metrics were compared among the kits and dilutions used. Overall, no single kit appeared to meet all the challenges of small input material. However, it is encouraging that excellent data can be recovered with even the 50 pg input total RNA.
ABSTRACT Clostridium difficile is the most commonly identified pathogen among health care-associated infections in the United States. There is a need for accurate and low-cost typing tools that produce comparable data across studies (i.e., portable data) to help characterize isolates during epidemiologic investigations of C. difficile outbreaks and sporadic cases of disease. The most popular C. difficile-typing technique is PCR ribotyping, and we previously developed methods using fluorescent PCR primers and amplicon sizing on a Sanger-style sequencer to generate fluorescent PCR ribotyping data. This technique has been used to characterize tens of thousands of C. difficile isolates from cases of disease. Here, we present validation of a protocol for the cost-effective generation of fluorescent PCR ribotyping data. A key component of this protocol is the ability to accurately identify PCR ribotypes against an online database (http://walklab.rcg.montana.edu) at no cost. We present results from a blinded multicenter study to address data portability across four different laboratories and three different sequencing centers. Our standardized protocol and centralized database for typing of C. difficile pathogens will increase comparability between studies so that important epidemiologic linkages between cases of disease and patterns of emergence can be rapidly identified.
We report sequencing-based whole-genome association analyses to evaluate the impact of rare and founder variants on stature in 6,307 individuals on the island of Sardinia. We identify two variants with large effects. One variant, which introduces a stop codon in the GHR gene, is relatively frequent in Sardinia (0.87% versus <0.01% elsewhere) and in the homozygous state causes Laron syndrome involving short stature. We find that this variant reduces height in heterozygotes by an average of 4.2 cm (-0.64 s.d.). The other variant, in the imprinted KCNQ1 gene (minor allele frequency (MAF) = 7.7% in Sardinia versus <1% elsewhere) reduces height by an average of 1.83 cm (-0.31 s.d.) when maternally inherited. Additionally, polygenic scores indicate that known height-decreasing alleles are at systematically higher frequencies in Sardinians than would be expected by genetic drift. The findings are consistent with selection for shorter stature in Sardinia and a suggestive human example of the proposed 'island effect' reducing the size of large mammals.
Gonçalo Abecasis, Francesco Cucca, David Schlessinger, Serena Sanna and colleagues report ~17.6 million genetic variants from whole-genome sequencing of 2,120 Sardinians. They assess the impact of these variants on circulating lipid levels and five inflammatory biomarkers. We report ∼17.6 million genetic variants from whole-genome sequencing of 2,120 Sardinians; 22% are absent from previous sequencing-based compilations and are enriched for predicted functional consequences. Furthermore, ∼76,000 variants common in our sample (frequency >5%) are rare elsewhere (<0.5% in the 1000 Genomes Project). We assessed the impact of these variants on circulating lipid levels and five inflammatory biomarkers. We observe 14 signals, including 2 major new loci, for lipid levels and 19 signals, including 2 new loci, for inflammatory markers. The new associations would have been missed in analyses based on 1000 Genomes Project data, underlining the advantages of large-scale sequencing in this founder population.
Jonathan N. V. Martinson, Susan Broadaway, Egan Lohman, Christina Johnson, M. Jahangir Alam, Mohammed Khaleduzzaman, Kevin W. Garey, Jessica Schlackman, Vincent B. Young, Kavitha Santhosh, Krishna Rao, Robert H. Lyons, Jr., Seth T. Walk Department of Microbiology and Immunology, Montana State University, Bozeman, Montana, USA; College of Pharmacy, University of Houston, Houston, Texas, USA; Department of Medicine, Division of Infectious Diseases, University of Pittsburgh, Pittsburgh, Pennsylvania, USA; Department of Internal Medicine, Division of Infectious Diseases, University of Michigan, Ann Arbor, Michigan, USA; Department of Microbiology and Immunology, University of Michigan, Ann Arbor, Michigan, USA; Division of Infectious Diseases, Veterans Affairs Ann Arbor Healthcare System, Ann Arbor, Michigan, USA; Department of Biological Chemistry, University of Michigan, Ann Arbor, Michigan, USA
Background—Heterozygous mutations in sarcomere genes in hypertrophic cardiomyopathy (HCM) are proposed to exert their effect through gain of function for missense mutations or loss of function for truncating mutations. However, allelic expression from individual mutations has not been sufficiently characterized to support this exclusive distinction in human HCM. Methods and Results—Sarcomere transcript and protein levels were analyzed in septal myectomy and transplant specimens from 46 genotyped HCM patients with or without sarcomere gene mutations and 10 control hearts. For truncating mutations in MYBPC3, the average ratio of mutant:wild-type transcripts was ≈1:5, in contrast to ≈1:1 for all sarcomere missense mutations, confirming that nonsense transcripts are uniquely unstable. However, total MYBPC3 mRNA was significantly increased by 9-fold in HCM samples with MYBPC3 mutations compared with control hearts and with HCM samples without sarcomere gene mutations. Full-length MYBPC3 protein content was not different between MYBPC3 mutant HCM and control samples, and no truncated proteins were detected. By absolute quantification of abundance with multiple reaction monitoring, stoichiometric ratios of mutant sarcomere proteins relative to wild type were strikingly variable in a mutation-specific manner, with the fraction of mutant protein ranging from 30% to 84%. Conclusions—These results challenge the concept that haploinsufficiency is a unifying mechanism for HCM caused by MYBPC3 truncating mutations. The range of allelic imbalance for several missense sarcomere mutations suggests that certain mutant proteins may be more or less stable or incorporate more or less efficiently into the sarcomere than wild-type proteins. These mutation-specific properties may distinctly influence disease phenotypes.
The utility of genotype imputation in genome-wide association studies is increasing as progressively larger reference panels are improved and expanded through whole-genome sequencing. Developing general guidelines for optimally cost-effective imputation, however, requires evaluation of performance issues that include the relative utility of study-specific compared with general/multipopulation reference panels; genotyping with various array scaffolds; effects of different ethnic backgrounds; and assessment of ranges of allele frequencies. Here we compared the effectiveness of study-specific reference panels to the commonly used 1000 Genomes Project (1000G) reference panels in the isolated Sardinian population and in cohorts of European ancestry including samples from Minnesota (USA). We also examined different combinations of genome-wide and custom arrays for baseline genotypes. In Sardinians, the study-specific reference panel provided better coverage and genotype imputation accuracy than the 1000G panels and other large European panels. In fact, even gene-centered custom arrays (interrogating ~200 000 variants) provided highly informative content across the entire genome. Gain in accuracy was also observed for Minnesotans using the study-specific reference panel, although the increase was smaller than in Sardinians, especially for rare variants. Notably, a combined panel including both study-specific and 1000G reference panels improved imputation accuracy only in the Minnesota sample, and only at rare sites. Finally, we found that when imputation is performed with a study-specific reference panel, cutoffs different from the standard thresholds of MACH-Rsq and IMPUTE-INFO metrics should be used to efficiently filter badly imputed rare variants. This study thus provides general guidelines for researchers planning large-scale genetic studies.
We present 21 microsatellite loci developed for Florida salt marsh voles (Microtus pennsylvanicus dukecampbelli). Microsatellites were identified from single molecule real time sequencing (Pacific Biosciences). We screened 30 loci and identified 21 loci as suitable for genotyping. We screened 17 individuals from Long Cabbage Key, and 3 individuals from an unnamed island. There was no significant departure from Hardy–Weinberg equilibrium or linkage equilibrium. Fifteen of the 21 loci were variable, with overall observed heterozygosity averaging 0.39, and a mean number of alleles of 3.14. Linkage disequilibrium estimate of N e was 10.7 (95 % CI 6.1–20.1). These markers will be useful for conservation genetics studies of this endangered species.
The semidominant Danforth's short tail (Sd) mutation arose spontaneously in the 1920s. The homozygous Sd phenotype includes severe malformations of the axial skeleton with an absent tail, kidney agenesis, anal atresia, and persistent cloaca. The Sd mutant phenotype mirrors features seen in human caudal malformation syndromes including urorectal septum malformation, caudal regression, VACTERL association, and persistent cloaca. The Sd mutation was previously mapped to a 0.9 cM region on mouse chromosome 2qA3. We performed Sanger sequencing of exons and intron/exon boundaries mapping to the Sd critical region and did not identify any mutations. We then performed DNA enrichment/capture followed by next-generation sequencing (NGS) of the critical genomic region. Standard bioinformatic analysis of paired-end sequence data did not reveal any causative mutations. Interrogation of reads that had been discarded because only a single end mapped correctly to the Sd locus identified an early transposon (ETn) retroviral insertion at the Sd locus, located 12.5 kb upstream of the Ptf1a gene. We show that Ptf1a expression is significantly upregulated in Sd mutant embryos at E9.5. The identification of the Sd mutation will lead to improved understanding of the developmental pathways that are misregulated in human caudal malformation syndromes.
Examining Y The evolution of human populations has long been studied with unique sequences from the nonrecombining, male-specific Y chromosome (see the Perspective by Cann ). Poznik et al. (p. 562 ) examined 9.9 Mb of the Y chromosome from 69 men from nine globally divergent populations—identifying population and individual specific sequence variants that elucidate the evolution of the Y chromosome. Sequencing of maternally inherited mitochondrial DNA allowed comparison between the relative rates of evolution, which suggested that the coalescence, or origin, of the human Y chromosome and mitochondria both occurred approximately 120 thousand years ago. Francalacci et al. (p. 565 ) investigated the sequence divergence of 1204 Y chromosomes that were sampled within the isolated and genetically informative Sardinian population. The sequence analyses, along with archaeological records, were used to calibrate and increase the resolution of the human phylogenetic tree.
, 565 (2013); 341 Science et al. Paolo Francalacci European Y-Chromosome Phylogeny Low-Pass DNA Sequencing of 1200 Sardinians Reconstructs This copy is for your personal, non-commercial use only. clicking here. colleagues, clients, or customers by , you can order high-quality copies for your If you wish to distribute this article to others here. following the guidelines can be obtained by Permission to republish or repurpose articles or portions of articles ): December 1, 2013 www.sciencemag.org (this information is current as of The following resources related to this article are available online at http://www.sciencemag.org/content/341/6145/565.full.html version of this article at: including high-resolution figures, can be found in the online Updated information and services, http://www.sciencemag.org/content/suppl/2013/08/01/341.6145.565.DC1.html can be found at: Supporting Online Material http://www.sciencemag.org/content/341/6145/565.full.html#related found at: can be related to this article A list of selected additional articles on the Science Web sites http://www.sciencemag.org/content/341/6145/565.full.html#ref-list-1 , 12 of which can be accessed free: cites 25 articles This article http://www.sciencemag.org/content/341/6145/565.full.html#related-urls 1 articles hosted by HighWire Press; see: cited by This article has been http://www.sciencemag.org/cgi/collection/genetics Genetics subject collections: This article appears in the following
The complex network of specialized cells and molecules in the immune system has evolved to defend against pathogens, but inadvertent immune system attacks on “self” result in autoimmune disease. Both genetic regulation of immune cell levels and their relationships with autoimmunity are largely undetermined. Here, we report genetic contributions to quantitative levels of 95 cell types encompassing 272 immune traits, in a cohort of 1,629 individuals from four clustered Sardinian villages. We first estimated trait heritability, showing that it can be substantial, accounting for up to 87% of the variance (mean 41%). Next, by assessing ∼8.2 million variants that we identified and confirmed in an extended set of 2,870 individuals, 23 independent variants at 13 loci associated with at least one trait. Notably, variants at three loci (HLA, IL2RA, and SH2B3/ATXN2) overlap with known autoimmune disease associations. These results connect specific cellular phenotypes to specific genetic variants, helping to explicate their involvement in disease.
As part of the DNA Sequencing Research Group of the Association of Biomolecular Resource Facilities, we have tested the reproducibility of the Roche/454 GS-FLX Titanium System at five core facilities. Experience with the Roche/454 system ranged from <10 to >340 sequencing runs performed. All participating sites were supplied with an aliquot of a common DNA preparation and were requested to conduct sequencing at a common loading condition. The evaluation of sequencing yield and accuracy metrics was assessed at a single site. The study was conducted using a laboratory strain of the Dutch elm disease fungus Ophiostoma novo-ulmi strain H327, an ascomycete, vegetatively haploid fungus with an estimated genome size of 30-50 Mb. We show that the Titanium System is reproducible, with some variation detected in loading conditions, sequencing yield, and homopolymer length accuracy. We demonstrate that reads shorter than the theoretical minimum length are of lower overall quality and not simply truncated reads. The O. novo-ulmi H327 genome assembly is 31.8 Mb and is comprised of eight chromosome-length linear scaffolds, a circular mitochondrial conti of 66.4 kb, and a putative 4.2-kb linear plasmid. We estimate that the nuclear genome encodes 8613 protein coding genes, and the mitochondrion encodes 15 genes and 26 tRNAs.
Steady-state gene expression is a coordination of synthesis and decay of RNA through epigenetic regulation, transcription factors, micro RNAs (miRNAs), and RNA-binding proteins. Here, we present bromouride labeling and sequencing (Bru-Seq) and bromouridine pulse-chase and sequencing (BruChase-Seq) to assess genome-wide changes to RNA synthesis and stability in human fibroblasts at homeostasis and after exposure to the proinflammatory tumor necrosis factor (TNF). The inflammatory response in human cells involves rapid and dramatic changes in gene expression, and the Bru-Seq and BruChase-Seq techniques revealed a coordinated and complex regulation of gene expression both at the transcriptional and posttranscriptional levels. The combinatory analysis of both RNA synthesis and stability using Bru-Seq and BruChase-Seq allows for a much deeper understanding of mechanisms of gene regulation than afforded by the analysis of steady-state total RNA and should be useful in many biological settings.
Isolating high-priority segments of genomes greatly enhances the efficiency of next-generation sequencing (NGS) by allowing researchers to focus on their regions of interest. For the 2010-11 DNA Sequencing Research Group (DSRG) study, we compared outcomes from two leading companies, Agilent Technologies (Santa Clara, CA, USA) and Roche NimbleGen (Madison, WI, USA), which offer custom-targeted genomic enrichment methods. Both companies were provided with the same genomic sample and challenged to capture identical genomic locations for DNA NGS. The target region totaled 3.5 Mb and included 31 individual genes and a 2-Mb contiguous interval. Each company was asked to design its best assay, perform the capture in replicates, and return the captured material to the DSRG-participating laboratories. Sequencing was performed in two different laboratories on Genome Analyzer IIx systems (Illumina, San Diego, CA, USA). Sequencing data were analyzed for sensitivity, specificity, and coverage of the desired regions. The success of the enrichment was highly dependent on the design of the capture probes. Overall, coverage variability was higher for the Agilent samples. As variant discovery is the ultimate goal for a typical targeted sequencing project, we compared samples for their ability to sequence single-nucleotide polymorphisms (SNPs) as a test of the ability to capture both chromosomes from the sample. In the targeted regions, we detected 2546 SNPs with the NimbleGen samples and 2071 with Agilent's. When limited to the regions that both companies included as baits, the number of SNPs was ∼1000 for each, with Agilent and NimbleGen finding a small number of unique SNPs not found by the other.
Multiple recent publications on RNA-Seq have demonstrated the power of next generation sequencing technologies in whole transcriptome analysis. The vendor specific protocols used for RNA library construction typically require at least 100ng of total RNA. However, under certain conditions such as single cells, stem cells, difficult to isolate cell types, or fractionated cancer cells, only a small amount of material is available. In these cases, effective transcriptome profiling requires amplification of subnanogram amounts of RNA. Several RNA amplification kits are available for amplification prior to library construction and next generation sequencing but these kits have not been comprehensively field evaluated for accuracy and performance of RNA-Seq for picogram amounts of RNA. This study conducted by the DNA Sequencing Research Group (DSRG) focuses on the evaluation of amplification kits for RNA-Seq. Four commercial amplification kits were chosen: Ovation v2 (NuGEN Technologies), SMARTer (Clontech), Seqplex (Sigma Aldrich), and Super-AMP (Miltenyi Biotech). Starting material was 5ng, 500pg and 50pg of human total reference RNA (Clontech) spiked with Ambion ERCC control mix (Life Technologies) following the manufacturer's protocol. Each kit was tested at 3 different sites to assess reproducibility. Total RNA and ERCC RNA spike-in control mixes from the same lots were sent to 12 ABRF lab sites for amplification and cDNA generation. Ideally, this would have resulted in 36 different amplified samples, 3 from each input RNA. Libraries were constructed at one site from the amplified cDNAs using the TruSeq RNA library preparation kit on the Tecan Freedom EVO Liquid Handling Robot. As an unamplified control, ribosomal depletion and PolyA selection were performed separately using 5ng, 100ng and 1ug of total RNA prior to library construction. All libraries were pooled and sequenced using the Illumina HiSeq platform. An overview of the study and the results will be presented.
Objective To identify disease-causing mutations within coding regions of 11 known NPHP genes (NPHP1-NPHP11) in a cohort of 192 patients diagnosed with a nephronophthisis-associated ciliopathy, at low cost. Methods Mutation analysis was carried out using PCR-based 48.48 Access Array microfluidic technology (Fluidigm) with consecutive next-generation sequencing. We applied a 10-fold primer multiplexing approach allowing PCR-based amplification of 475 amplicons (251 exons) for 48 DNA samples simultaneously. After four rounds of amplification followed by indexing all of 192 patient-derived products with different barcodes in a subsequent PCR, 2×100 paired-end sequencing was performed on one lane of a HiSeq2000 instrument (Illumina). Bioinformatics analysis was performed using ‘CLC Genomics Workbench’ software. Potential mutations were confirmed by Sanger sequencing and shown to segregate. Results Bioinformatics analysis revealed sufficient coverage of 30×for 168/192 (87.5%) DNA samples (median 449×) and of 234 out of 251 targeted coding exons (sensitivity: 93.2%). For proof-of-principle, we analysed 20 known mutations and identified 18 of them in the correct zygosity state (90%). Likewise, we identified pathogenic mutations in 34/192 patients (18%) and discovered 23 novel mutations in the genes NPHP3 (7), NPHP4 (3), IQCB1 (4), CEP290 (7), RPGRIP1L (1), and TMEM67 (1). Additionally, we found 40 different single heterozygous missense variants of unknown significance. Conclusions We conclude that the combined approach of array-based multiplexed PCR-amplification on a Fluidigm Access Array platform followed by next-generation sequencing is highly cost-efficient and strongly facilitates diagnostic mutation analysis in broadly heterogeneous Mendelian disorders.