The goal of the Collaborative Cross (CC) project was to generate and distribute over 1000 independent mouse recombinant inbred strains derived from eight inbred founders. With inbreeding nearly complete, we estimated the extinction rate among CC lines at a remarkable 95%, which is substantially higher than in the derivation of other mouse recombinant inbred populations. Here, we report genome-wide allele frequencies in 347 extinct CC lines. Contrary to expectations, autosomes had equal allelic contributions from the eight founders, but chromosome X had significantly lower allelic contributions from the two inbred founders with underrepresented subspecific origins (PWK/PhJ and CAST/EiJ). By comparing extinct CC lines to living CC strains, we conclude that a complex genetic architecture is driving extinction, and selection pressures are different on the autosomes and chromosome X Male infertility played a large role in extinction as 47% of extinct lines had males that were infertile. Males from extinct lines had high variability in reproductive organ size, low sperm counts, low sperm motility, and a high rate of vacuolization of seminiferous tubules. We performed QTL mapping and identified nine genomic regions associated with male fertility and reproductive phenotypes. Many of the allelic effects in the QTL were driven by the two founders with underrepresented subspecific origins, including a QTL on chromosome X for infertility that was driven by the PWK/PhJ haplotype. We also performed the first example of cross validation using complementary CC resources to verify the effect of sperm curvilinear velocity from the PWK/PhJ haplotype on chromosome 2 in an independent population across multiple generations. While selection typically constrains the examination of reproductive traits toward the more fertile alleles, the CC extinct lines provided a unique opportunity to study the genetic architecture of fertility in a widely genetically variable population. We hypothesize that incompatibilities between alleles with different subspecific origins is a key driver of infertility. These results help clarify the factors that drove strain extinction in the CC, reveal the genetic regions associated with poor fertility in the CC, and serve as a resource to further study mammalian infertility.
Genotyping microarrays are an important resource for genetic mapping, population genetics, and monitoring of the genetic integrity of laboratory stocks. We have developed the third generation of the Mouse Universal Genotyping Array (MUGA) series, GigaMUGA, a 143,259-probe Illumina Infinium II array for the house mouse (Mus musculus). The bulk of the content of GigaMUGA is optimized for genetic mapping in the Collaborative Cross and Diversity Outbred populations, and for substrain-level identification of laboratory mice. In addition to 141,090 single nucleotide polymorphism probes, GigaMUGA contains 2006 probes for copy number concentrated in structurally polymorphic regions of the mouse genome. The performance of the array is characterized in a set of 500 high-quality reference samples spanning laboratory inbred strains, recombinant inbred lines, outbred stocks, and wild-caught mice. GigaMUGA is highly informative across a wide range of genetically diverse samples, from laboratory substrains to other Mus species. In addition to describing the content and performance of the array, we provide detailed probe-level annotation and recommendations for quality control.
Nat. Genet. 47, 353–360 (2015); published online 2 March 2015; corrected after print 16 April 2015 In the version of this article initially published, an accession number was not provided for RNA-seq data sets. The RNA-seq data sets that passed quality control are available at the Sequence Read Archive (SRA) under accession SRP056236.
RNA-seq technology enables large-scale studies of allele-specific expression (ASE), or the expression difference between maternal and paternal alleles. Here, we study ASE in animals for which parental RNA-seq data are available. While most methods for determining ASE rely on read alignment, read alignment either leads to reference bias or requires knowledge of genomic variants in each parental strain. When RNA-seq data are available for both parental strains of a hybrid animal, it is possible to infer ASE with minimal reference bias and without knowledge of parental genomic variants. Our approach first uses parental RNA-seq reads to discover maternal and paternal versions of transcript sequences. Using these alternative transcript sequences as features, we estimate abundance levels of transcripts in the hybrid animal using a modified lasso linear regression model.We tested our methods on synthetic data from the mouse transcriptome and compared our results with those of Trinity, a state-of-the-art de novo RNA-seq assembler. Our methods achieved high sensitivity and specificity in both identifying expressed transcripts and transcripts exhibiting ASE. We also ran our methods on real RNA-seq mouse data from two F1 samples with wild-derived parental strains and were able to validate known genes exhibiting ASE, as well as confirm the expected maternal contribution ratios in all genes and genes on the X chromosome.
Numerous microarray genotype-calling methods rely on fitting a parametric model to clusters derived from the hybridization intensities of training data. However, in most cases we are uncertain about the expected sample distribution and the resulting parametric model tends to be inaccurate if the assumptions of the data distribution are not met. Moreover, many methods assume four genotypes (reference allele, alternate allele, heterozygous allele, or no call) and use a common parametric model that applies to all probes. We demonstrate that conversion of probe intensities to discrete genotypes and applying the same model to all probes results in information loss and even incorrect genotype calls due to genomic variations within probes [6]. We make no assumption about the data distribution. We represent cluster distribution using a non-parametric model which is consistent with the data. The model can be easily evaluated and it provides genotype calls for a given sample's marker using a table lookup. Furthermore, our algorithms have no prior assumptions concerning the number of genotype calls. We apply the algorithms to each probe separately whereas others infer a common set of clusters that apply to all probes. We demonstrate our methods on Collaborative Cross (CC) genetic reference mice population and all samples are genotyped using a 78,000-marker genotyping array on Illumina platform. Our algorithm exhibits high concordance with Illumina genotype calls and achieves > 98% call rates on all CC samples. Code for the described algorithms is available by request from the authors.
Many methods have been developed for mapping quantitative trait loci (QTLs) using microarrays. Traditional methods for QTL mapping rely on the assumption that biallelic genotype calls represent the complete genetic variation at a marker. In reality, the process of converting microarray intensities to discrete genotype calls results in the loss of marker information on other variations involving the marker sequence, such as nearby SNPs, deletions, or copy numbers. We have developed a novel approach to QTL mapping that directly uses microarray marker intensities. Our method scans for marker windows where the intensity distances between sample pairs are correlated with the quantitative phenotype differences. The presence of such markers indicates that samples which are genetically close together in the region also share similar phenotype values, suggesting the presence of a QTL. The significance of putative QTLs is then assessed through permutation testing. By directly incorporating genotype intensities, our method eliminates intermediate processes such as genotype calling or ancestry inference that may introduce uncertainty or data loss. We tested our method on synthetic phenotype data of mice genotyped with the 78K-marker MegaMUGA array, and our results compared favorably to those of R/qtl, a well-establishe QTL mapping package. In addition, we used our method to map the binary albino trait in inbred and backcrossed mice to the tyrosinase ( Tyr ) gene on chromosome 7, and we also verified several QTLs found to affect colitis-related traits from a previous mouse study.
In this paper, we contrast the resolution and accuracy of determining recombination boundaries using genotyping arrays compared to high-throughput sequencing. In addition, we consider the impacts of sequence coverage and genetic diversity on localizing recombination boundaries. We developed a hidden Markov model for estimating recombination breakpoints based on variant observations seen in the read coverage spanning uniformly sized genomic windows. Our model includes 36 states representing all combinations of 8 genomes, and estimates a founder mosaic that is consistent with the variants observed in the aligned sequences. At HMM transition locations we consider the most likely founder-pair and refine the recombination breakpoints down to an interval spanning two informative variants. We compare this solution to alternate solutions based on microarrays that we have estimated. At 30x coverage the recombination mapping accuracy far exceeds the resolution attainable by any microarray. Even at coverages of 1x and below we are generally able to estimate recombination breakpoints with comparable accuracy.
We wish to study allele-specific expression in diploid organisms, specifically in F1 animals with inbred parental strains. Current methods for analyzing allele-specific expression rely on read alignment, which leads to reference bias unless there is prior knowledge of all genomic variants in the parental strains. However, in the case where RNA-seq data is available for both parental strains, we do not need prior knowledge of parental genomic variants. Our approach first uses parental RNA-seq reads to create maternal and paternal versions of transcript sequences, then estimates allele-specific expression levels in the F1 animal for each transcript. Using the parental versions of all candidate transcripts as features, we use a modified lasso penalized linear regression model for estimating abundance levels of expressed transcripts in the F1 animal. We tested our methods on synthetic data from the mouse transcriptome and compared our results with those of Trinity, a state-of-the-art de novo RNA-seq assembler. Our methods achieved much higher sensitivity and specificity in both identifying expressed transcripts and transcripts exhibiting allele-specific expression. We were also able to separately predict relative expression levels from paternal and maternal strains with more accuracy.
High-density genotyping arrays that measure hybridization of genomic DNA fragments to allele-specific oligonucleotide probes are widely used to genotype single nucleotide polymorphisms (SNPs) in genetic studies, including human genome-wide association studies. Hybridization intensities are converted to genotype calls by clustering algorithms that assign each sample to a genotype class at each SNP. Data for SNP probes that do not conform to the expected pattern of clustering are often discarded, contributing to ascertainment bias and resulting in lost information - as much as 50% in a recent genome-wide association study in dogs.
Numerous methods exist for inferring the ancestry mosaic of an admixed individual based on its genotypes and those of its ancestors. These methods rely on bialleic SNPs obtained from genotype calling algorithms, which classify each marker as belonging to one of four states (reference allele, alternate allele, heterozygous, or no call) based on probe hybridization intensity signals. We demonstrate that this conversion of probe intensities to discrete genotypes can lead to a loss of information and introduce errors via incorrect genotype calls. We propose a method that directly infers ancestry from probe intensities by minimizing the intensity difference between a target individual and one or more of its ancestors. We demonstrate our method on mice from the developing Collaborative Cross (CC) genetic reference population, which are admixtures of a common set of eight ancestors. Our samples were genotyped using a 7.8K-marker Illumina Infinium platform called the Mouse Universal Genotyping Array (MUGA). We compare our reconstructions with a standard genotype-based method and validate our results using DNA sequencing data. Our algorithm is able to use information not captured by genotype calls and avoid errors due to incorrect calls.
(4,4-Dimethoxy-butyl)-dimethyl-amine,the structure ofwhich was identified by IR,MS and 1HNMR,was synthesized from tetrahydrofuran via a series of reactions including ring opening,oxidation,aldolization and substitution.Research shows that in this process conditions,simple and easy to control,the reaction process is polluteness and friendly to the enviroment,that the yield is high.
Wei Wang (王薇)合作论文数Department of Computer Science, University of California at Los Angeles;Department of Computational Medicine, University of California at Los Angeles;Scalable Analytics Institute, University of California at Los Angeles1