Epigenetic mechanisms such as genomic imprinting demonstrate that molecular inheritance can deviate from typical Mendelian patterns. Despite this, the intergenerational inheritance of DNA methylation remains poorly understood. Here we developed a genome-wide approach to study epigenetic inheritance in mice using long-read nanopore sequencing. Using this approach in both liver and muscle, we found that ~93% of autosomal epigenetic inheritance patterns followed Mendel's laws, primarily driven by cis-acting methylation quantitative trait loci. However, we also identified extensive non-Mendelian inheritance, including emergent epigenetic inheritance patterns, widespread sex-specific DNA methylation patterns localized to the liver, and five seemingly new autosomal and X-linked imprinted genes. Notably, we also report an example of naturally occurring intergenerational paramutation, confirmed over strain-specific transposable elements within Capn11 and highly likely at Vps37c. Overall, an unexpectedly high ~7% of autosomal epigenetic inheritance patterns identified were non-Mendelian, highlighting the importance of epigenetic information in the analysis of inherited traits and disorders.
The drawing process is crucial to understanding the final result of a drawing. There has been a long history of understanding human drawing; what kinds of strokes people use and where they are placed. An area of interest in Artificial Intelligence is developing systems that simulate human behavior in drawing. However, there has been little work done to understand the order of strokes in the drawing process. Without sufficient understanding of natural drawing order, it is difficult to build models that can generate natural drawing processes. In this paper, we present a study comparing multiple types of stroke orders to confirm findings from previous work and demonstrate that multiple orderings of the same set of strokes can be perceived as human-drawn and different stroke order types achieve different perceived naturalness depending on the type of image prompt.
Transposable Elements (TEs) are DNA subsequences that have historically copied themselves throughout a genome. Apart from constituting a large fraction of all eukaryotic genomes, TEs are a significant source of genetic variation and are directly responsible for many diseases. TEs are also one of the most difficult genomic regions to analyze. A typical approach for identifying TE insertions (TEi) involves the detection of split-reads, which requires checking if each read can be split into TE and non-TE parts. Identification of the TE part depends on a model for each distinct TE class, and these classes vary significantly both within and between species. Previous methods for detecting segregating TEis depend on template libraries and their computational cost increases with the number of templates. Here we propose a novel template-free method for identifying the split-reads containing TEi boundaries called Frontier. We leverage the pervasiveness of TE sequences to identify candidate reads that might include the boundary of an insertion. We then apply machine learning methods to further classify whether the read includes actual TE-like sequence. For each predicted TEi boundary we apply a second classifier to infer the corresponding TE type (LINE, SINE, ALU, ERV/LTR). Both classifiers achieve high precision (> .9), recall (> .8) and F1 score (> .8) when applied to real data. The resulting trained model, can detect and classify about 50 million frontier reads in less than an hour. Frontier codes are available at github https://github.com/Anwica/Frontier.
Inexpensive and fast genome sequencing has yielded multiple genome assemblies that, taken together, can be considered as a single pangenome model. However, applying conventional alignment-based sequence analysis to the assemblies of a pangenome is computationally expensive and largely redundant. Here, we present an alignment-free method that analyzes the relationship of any new sample relative to a given pangenome model using selected k-mer queries. We select a representative set of k-mers from the pangenome as probes and determine their frequencies in the raw short-read sequence data. The selection of probes is designed to cover every base of the pangenome, maximize sharing, and identify informative probes that discriminate between haplotypes. The k-mer frequencies are determined using an FM-index built over the raw sequence data of the new sample. Prior to the k-mer search, the probes are reordered to maximize the shared suffixes between succesive k-mers, thus reducing the overall run time compared to executing each search independently. We aggregate the forward and reverse k-mer probe counts, save them in the appropriate rows of a count matrix and remap them back to their locations in the pangenome. The resulting probe database serves as a valuable resource for representing population-scale sequence variations based on the pangenome model.
Non-photorealistic rendering (NPR) and image processing algorithms are widely assumed as a proxy for drawing. However, this assumption is not well assessed due to the difficulty in collecting and registering freehand drawings. Alternatively, tracings are easier to collect and register, but there is no quantitative evaluation of tracing as a proxy for freehand drawing. In this paper, we compare tracing, freehand drawing, and computer-generated drawing approximation (CGDA) to understand their similarities and differences. We collected a dataset of 1,498 tracings and freehand drawings by 110 participants for 100 image prompts. Our drawings are registered to the prompts and include vector-based timestamped strokes collected via stylus input. Comparing tracing and freehand drawing, we found a high degree of similarity in stroke placement and types of strokes used over time. We show that tracing can serve as a viable proxy for freehand drawing because of similar correlations between spatio-temporal stroke features and labeled stroke types. Comparing hand-drawn content and current CGDA output, we found that 60% of drawn pixels corresponded to computer-generated pixels on average. The overlap tended to be commonly drawn content, but people's artistic choices and temporal tendencies remained largely uncaptured. We present an initial analysis to inform new CGDA algorithms and drawing applications, and provide the dataset for use by the community.
The laboratory mouse is the most widely used animal model for biomedical research, due in part to its well-annotated genome, wealth of genetic resources, and the ability to precisely manipulate its genome. Despite the importance of genetics for mouse research, genetic quality control (QC) is not standardized, in part due to the lack of cost-effective, informative, and robust platforms. Genotyping arrays are standard tools for mouse research and remain an attractive alternative even in the era of high-throughput whole-genome sequencing. Here, we describe the content and performance of a new iteration of the Mouse Universal Genotyping Array (MUGA), MiniMUGA, an array-based genetic QC platform with over 11,000 probes. In addition to robust discrimination between most classical and wild-derived laboratory strains, MiniMUGA was designed to contain features not available in other platforms: (1) chromosomal sex determination, (2) discrimination between substrains from multiple commercial vendors, (3) diagnostic SNPs for popular laboratory strains, (4) detection of constructs used in genetically engineered mice, and (5) an easy-to-interpret report summarizing these results. In-depth annotation of all probes should facilitate custom analyses by individual researchers. To determine the performance of MiniMUGA, we genotyped 6899 samples from a wide variety of genetic backgrounds. The performance of MiniMUGA compares favorably with three previous iterations of the MUGA family of arrays, both in discrimination capabilities and robustness. We have generated publicly available consensus genotypes for 241 inbred strains including classical, wild-derived, and recombinant inbred lines. Here, we also report the detection of a substantial number of XO and XXY individuals across a variety of sample types, new markers that expand the utility of reduced complexity crosses to genetic backgrounds other than C57BL/6, and the robust detection of 17 genetic constructs. We provide preliminary evidence that the array can be used to identify both partial sex chromosome duplication and mosaicism, and that diagnostic SNPs can be used to determine how long inbred mice have been bred independently from the relevant main stock. We conclude that MiniMUGA is a valuable platform for genetic QC, and an important new tool to increase the rigor and reproducibility of mouse research.
The azoxymethane carcinogen model of non-familial colorectal cancer has been used in mice to identify six new susceptibility loci and confirm 18 of 24 previous detected susceptibility loci. Using a population-based approach, the genetic architecture of colon cancer... The azoxymethane model of colorectal cancer (CRC) was used to gain insights into the genetic heterogeneity of nonfamilial CRC. We observed significant differences in susceptibility parameters across 40 mouse inbred strains, with 6 new and 18 of 24 previously identified mouse CRC modifier alleles detected using genome-wide association analysis. Tumor incidence varied in F1 as well as intercrosses and backcrosses between resistant and susceptible strains. Analysis of inheritance patterns indicates that resistance to CRC development is inherited as a dominant characteristic genome-wide, and that susceptibility appears to occur in individuals lacking a large-effect, or sufficient numbers of small-effect, polygenic resistance alleles. Our results suggest a new polygenic model for inheritance of nonfamilial CRC, and that genetic studies in humans aimed at identifying individuals with elevated susceptibility should be pursued through the lens of absence of dominant resistance alleles rather than for the presence of susceptibility alleles.
Reproductive success in the eight founder strains of the Collaborative Cross (CC) was measured using a diallel-mating scheme. Over a 48-month period we generated 4,448 litters, and provided 24,782 weaned pups for use in 16 different published experiments. We identified factors that affect the average litter size in a cross by estimating the overall contribution of parent-of-origin, heterosis, inbred, and epistatic effects using a Bayesian zero-truncated overdispersed Poisson mixed model. The phenotypic variance of litter size has a substantial contribution (82%) from unexplained and environmental sources, but no detectable effect of seasonality. Most of the explained variance was due to additive effects (9.2%) and parental sex (maternal vs. paternal strain; 5.8%), with epistasis accounting for 3.4%. Within the parental effects, the effect of the dam’s strain explained more than the sire’s strain (13.2% vs. 1.8%), and the dam’s strain effects account for 74.2% of total variation explained. Dams from strains C57BL/6J and NOD/ShiLtJ increased the expected litter size by a mean of 1.66 and 1.79 pups, whereas dams from strains WSB/EiJ, PWK/PhJ, and CAST/EiJ reduced expected litter size by a mean of 1.51, 0.81, and 0.90 pups. Finally, there was no strong evidence for strain-specific effects on sex ratio distortion. Overall, these results demonstrate that strains vary substantially in their reproductive ability depending on their genetic background, and that litter size is largely determined by dam’s strain rather than sire’s strain effects, as expected. This analysis adds to our understanding of factors that influence litter size in mammals, and also helps to explain breeding successes and failures in the extinct lines and surviving CC strains.
A large fraction of mammalian genome consists of transposable elements (TEs). These elements are segments of DNA that either move or are copied from one place in the genome to another. Such movements can cause deleterious mutations and drive chromosome evolution. Existing approaches search for TE insertion (TEi) by aligning millions of mostly irrelevant short reads to either a reference genome or a TE sequence library. Here we present a new local genome assembly based pipeline, called ELITE, for identifying and characterizing TEi. ELITE uses an msBWT-based data structure to store and index all the reads from a high-throughput sequencing dataset and leverages a sampled FM-index to detect TEi efficiently. In comparison with two existing tools, ELITE is faster and has a higher precision and recall rate in predicting TEi. ELITE also works on real data, which we validated using PCR assays of surrounding genomic context. Additional features of ELITE include finding zygosity status of a predicted TEi, discovering unannotated TEs that are distantly related to the target one, and providing a summary of TEi sharing pattern within a population.
The mouse reference is one of the most widely used and accurately assembled mammalian genomes, and is the foundation for a wide range of bioinformatics and genetics tools. However, it represents the genomic organization of a single inbred mouse strain. Recently, inexpensive and fast genome sequencing has enabled the assembly of other common mouse strains at a quality approaching that of the reference. However, using these alternative assemblies in standard genomics analysis pipelines presents significant challenges. It has been suggested that a pangenome reference assembly, which incorporates multiple genomes into a single representation, are the path forward, but there are few standards for, or instances of practical pangenome representations suitable for large eukaryotic genomes. We present a pragmatic graph-based pangenome representation as a genomic resource for the widely-used recombinant-inbred mouse genetic reference population known as the Collaborative Cross (CC) and its eight founder genomes. Our pangenome representation leverages existing standards for genomic sequence representations with backward-compatible extensions to describe graph topology and genome-specific annotations along paths. It packs 83 mouse genomes (8 founders + 75 CC strains) into a single graph representation that captures important notions relating genomes such as identity-by-descent and highly variable genomic regions. The introduction of special anchor nodes with sequence content provides a valid coordinate framework that divides large eukaryotic genomes into homologous segments and addresses most of the graph-based position reference issues. Parallel edges between anchors place variants within a context that facilitates orthogonal genome comparison and visualization. Furthermore, our graph structure allows annotations to be placed in multiple genomic contexts and simplifies their maintenance as the assembly improves. The CC reference pangenome provides an open framework for new tool chain development and analysis.
Two key features of recombinant inbred panels are well-characterized genomes and reproducibility. Here we report on the sequenced genomes of six additional Collaborative Cross (CC) strains and on inbreeding progress of 72 CC strains. We have previously reported on the sequences of 69 CC strains that were publicly available, bringing the total of CC strains with whole genome sequence up to 75. The sequencing of these six CC strains updates the efforts toward inbreeding undertaken by the UNC Systems Genetics Core. The timing reflects our competing mandates to release to the public as many CC strains as possible while achieving an acceptable level of inbreeding. The new six strains have a higher than average founder contribution from non-domesticus strains than the previously released CC strains. Five of the six strains also have high residual heterozygosity (>14%), which may be related to non-domesticus founder contributions. Finally, we report on updated estimates on residual heterozygosity across the entire CC population using a novel, simple and cost effective genotyping platform on three mice from each strain. We observe a reduction in residual heterozygosity across all previously released CC strains. We discuss the optimal use of different genetic resources available for the CC population.
BACKGROUND:Long read sequencing is changing the landscape of genomic research, especially de novo assembly. Despite the high error rate inherent to long read technologies, increased read lengths dramatically improve the continuity and accuracy of genome assemblies. However, the cost and throughput of these technologies limits their application to complex genomes. One solution is to decrease the cost and time to assemble novel genomes by leveraging "hybrid" assemblies that use long reads for scaffolding and short reads for accuracy.RESULTS:We describe a novel method leveraging a multi-string Burrows-Wheeler Transform with auxiliary FM-index to correct errors in long read sequences using a set of complementary short reads. We demonstrate that our method efficiently produces significantly more high quality corrected sequence than existing hybrid error-correction methods. We also show that our method produces more contiguous assemblies, in many cases, than existing state-of-the-art hybrid and long-read only de novo assembly methods.CONCLUSION:Our method accurately corrects long read sequence data using complementary short reads. We demonstrate higher total throughput of corrected long reads and a corresponding increase in contiguity of the resulting de novo assemblies. Improved throughput and computational efficiency than existing methods will help better economically utilize emerging long read sequencing technologies.
Inspired by curriculum learning, we present a two-tiered approach for CNN training. First, we learn an intermediate representation, called a pixel labeling. The pixel labels capture low-level details and textures within objects, providing a more complete semantic description that is used as input to a subsequent image classifier. The two learning tasks can be considered together as a single deep network. Extensive experiments show that our architecture, which includes an intermediate layer substantially outperforms fine-tuned CNN models trained without an intermediate target, even when the two networks have an identical overall topology and numbers of parameters. We demonstrate our approach on a histology image classification task.
Before genotyping microarrays can be used, calling algorithms must first be calibrated with a control set. Calling algorithms that evaluate hybridization intensity data on the basis of individual markers are better able to compensate for sequence specific variations. However, they require that the control set includes samples sufficient to exercise every marker in all of its allelic states. Minimizing the size of the control set is an important cost-saving measure for the design and production of custom, population-specific microarrays. As the size of the population that the array must discriminate between increases, the naïve choice of a control set grows quadratically. We show that a control set that is linear in the population size always exists, but the problem of finding such a linear-sized set grows combinatorially. We examine the problem of finding an optimally-sized control set, in particular for arrays designed to discriminate among an inbred population and their crosses. We derive tight, in the sense of being attainable, lower and upper bounds on the solution size. Further, we show that the problem is equivalent to the set cover problem. We make use of greedy approximate methods to the set cover problem along with our established bounds to create a branch-and-bound framework. We demonstrate our methods on simulated data, available microarrays, and one microarray being developed.
The Collaborative Cross (CC) is a multiparent panel of recombinant inbred (RI) mouse strains derived from eight founder laboratory strains. RI panels are popular because of their long-term genetic stability, which enhances reproducibility and integration of data collected across time and conditions. Characterization of their genomes can be a community effort, reducing the burden on individual users. Here we present the genomes of the CC strains using two complementary approaches as a resource to improve power and interpretation of genetic experiments. Our study also provides a cautionary tale regarding the limitations imposed by such basic biological processes as mutation and selection. A distinct advantage of inbred panels is that genotyping only needs to be performed on the panel, not on each individual mouse. The initial CC genome data were haplotype reconstructions based on dense genotyping of the most recent common ancestors (MRCAs) of each strain followed by imputation from the genome sequence of the corresponding founder inbred strain. The MRCA resource captured segregating regions in strains that were not fully inbred, but it had limited resolution in the transition regions between founder haplotypes, and there was uncertainty about founder assignment in regions of limited diversity. Here we report the whole genome sequence of 69 CC strains generated by paired-end short reads at 30× coverage of a single male per strain. Sequencing leads to a substantial improvement in the fine structure and completeness of the genomes of the CC. Both MRCAs and sequenced samples show a significant reduction in the genome-wide haplotype frequencies from two wild-derived strains, CAST/EiJ and PWK/PhJ. In addition, analysis of the evolution of the patterns of heterozygosity indicates that selection against three wild-derived founder strains played a significant role in shaping the genomes of the CC. The sequencing resource provides the first description of tens of thousands of new genetic variants introduced by mutation and drift in the CC genomes. We estimate that new SNP mutations are accumulating in each CC strain at a rate of 2.4 ± 0.4 per gigabase per generation. The fixation of new mutations by genetic drift has introduced thousands of new variants into the CC strains. The majority of these mutations are novel compared to currently sequenced laboratory stocks and wild mice, and some are predicted to alter gene function. Approximately one-third of the CC inbred strains have acquired large deletions (>10 kb) many of which overlap known coding genes and functional elements. The sequence of these mice is a critical resource to CC users, increases threefold the number of mouse inbred strain genomes available publicly, and provides insight into the effect of mutation and drift on common resources.
The ability to accurately monitor alterations in sperm motility is paramount to understanding multiple genetic and biochemical perturbations impacting normal fertilization. Computer-aided sperm analysis (CASA) of human sperm typically reports motile percentage and kinematic parameters at the population level, and uses kinematic gating methods to identify subpopulations such as progressive or hyperactivated sperm. The goal of this study was to develop an automated method that classifies all patterns of human sperm motility during in vitro capacitation following the removal of seminal plasma. We visually classified CASA tracks of 2817 sperm from 18 individuals and used a support vector machine-based decision tree to compute four hyperplanes that separate five classes based on their kinematic parameters. We then developed a web-based program, CASAnova, which applies these equations sequentially to assign a single classification to each motile sperm. Vigorous sperm are classified as progressive, intermediate, or hyperactivated, and nonvigorous sperm as slow or weakly motile. This program correctly classifies sperm motility into one of five classes with an overall accuracy of 89.9%. Application of CASAnova to capacitating sperm populations showed a shift from predominantly linear patterns of motility at initial time points to more vigorous patterns, including hyperactivated motility, as capacitation proceeds. Both intermediate and hyperactivated motility patterns were largely eliminated when sperm were incubated in noncapacitating medium, demonstrating the sensitivity of this method. The five CASAnova classifications are distinctive and reflect kinetic parameters of washed human sperm, providing an accurate, quantitative, and high-throughput method for monitoring alterations in motility.
Gene duplication and loss are major sources of genetic polymorphism in populations, and are important forces shaping the evolution of genome content and organization. We have reconstructed the origin and history of a 127-kbp segmental duplication, R2d, in the house mouse (Mus musculus). R2d contains a single protein-coding gene, Cwc22. De novo assembly of both the ancestral (R2d1) and the derived (R2d2) copies reveals that they have been subject to nonallelic gene conversion events spanning tens of kilobases. R2d2 is also a hotspot for structural variation: its diploid copy number ranges from zero in the mouse reference genome to >80 in wild mice sampled from around the globe. Hemizygosity for high copy-number alleles of R2d2 is associated in cis with meiotic drive; suppression of meiotic crossovers; and copy-number instability, with a mutation rate in excess of 1 per 100 transmissions in some laboratory populations. Our results provide a striking example of allelic diversity generated by duplication and demonstrate the value of de novo assembly in a phylogenetic context for understanding the mutational processes affecting duplicate genes.
Wild-derived mouse inbred strains are becoming increasingly popular for complex traits analysis, evolutionary studies, and systems genetics. Here, we report the whole-genome sequencing of two wild-derived mouse inbred strains, LEWES/EiJ and ZALENDE/EiJ, of Mus musculus domesticus origin. These two inbred strains were selected based on their geographic origin, karyotype, and use in ongoing research. We generated 14× and 18× coverage sequence, respectively, and discovered over 1.1 million novel variants, most of which are private to one of these strains. This report expands the number of wild-derived inbred genomes in the Mus genus from six to eight. The sequence variation can be accessed via an online query tool; variant calls (VCF format) and alignments (BAM format) are available for download from a dedicated ftp site. Finally, the sequencing data have also been stored in a lossless, compressed, and indexed format using the multi-string Burrows-Wheeler transform. All data can be used without restriction.
s from the 41st American Society of Andrology Annual Meeting
Genotyping microarrays are an important resource for genetic mapping, population genetics, and monitoring of the genetic integrity of laboratory stocks. We have developed the third generation of the Mouse Universal Genotyping Array (MUGA) series, GigaMUGA, a 143,259-probe Illumina Infinium II array for the house mouse (Mus musculus). The bulk of the content of GigaMUGA is optimized for genetic mapping in the Collaborative Cross and Diversity Outbred populations, and for substrain-level identification of laboratory mice. In addition to 141,090 single nucleotide polymorphism probes, GigaMUGA contains 2006 probes for copy number concentrated in structurally polymorphic regions of the mouse genome. The performance of the array is characterized in a set of 500 high-quality reference samples spanning laboratory inbred strains, recombinant inbred lines, outbred stocks, and wild-caught mice. GigaMUGA is highly informative across a wide range of genetically diverse samples, from laboratory substrains to other Mus species. In addition to describing the content and performance of the array, we provide detailed probe-level annotation and recommendations for quality control.
Wei Wang (王薇)合作论文数Department of Computer Science, University of California at Los Angeles;Department of Computational Medicine, University of California at Los Angeles;Scalable Analytics Institute, University of California at Los Angeles24