Cytosolic transfer RNAs (tRNAs), which are encoded as hundreds of genes in nuclear genomes, experience exceptionally high mutation rates and have been hypothesized to confer substantial mutational load in natural populations. Although this phenomenon appears universal across multicellular eukaryotes, a comprehensive characterization of standing variation in tRNA repertoires is still lacking in any system. Here, we resolve within-species allelic variation in nuclear-encoded tRNAs in three nematode species: Caenorhabditis elegans, C. briggsae, and C. tropicalis. We show that these genes carry signatures of high rates of historical transcription-associated mutagenesis and of purifying selection, resulting in allelic variation that includes pervasive instances of within-gene mismatches between the amino acid recognized by the tRNA backbone and that indicated by the anticodon. Furthermore, patterns of tRNA genomic organization and variation differ markedly from those of protein-coding regions. Individual genomes harbor distinct complements of tRNA genes with predicted functional differences, an observation that coincides with recent evidence that variation in tRNA expression and regulation contributes to human disease. Our findings offer an entry point for identifying the microevolutionary processes that act on tRNA repertoires and, in turn, connecting those processes to the macroevolutionary patterns that have more frequently been the focus of study.
Cytosolic transfer RNAs, which are encoded as hundreds of genes in eukaryotic nuclear genomes, experience exceptionally high rates of mutation and have been hypothesized to carry significant mutational load in natural populations. Comparisons across species show pervasive losses, gains, and functional remolding at tRNA loci, indicating rapid evolution over intermediate timescales, but the underlying molecular changes happening over short timescales remain poorly understood. A central gap in our understanding is the nature of deleterious tRNA mutations, how they affect fitness, and the extent to which tRNA pools vary functionally across genomes. Furthermore, new technologies for capturing tRNA expression and modification emphasize the importance of tRNA activity to organismal fitness, as tRNA dysregulation is associated with disease burden in humans, but this research area has not yet included possible consequences of mutational variation. To infer the dynamics of tRNA evolution over short timescales, and to integrate our understanding of fitness consequences of tRNA regulation with mutation load, we first require a comprehensive view of standing variation in cytosolic tRNAs within populations. In this study, we resolve within-species allelic variation in nuclear-encoded tRNAs in three nematode species: Caenorhabditis elegans , C. briggsae , and C. tropicalis . We show that these tRNA repertoires have been shaped by transcription-associated mutagenesis and selection and exhibit extensive allelic variation, including isotype switching between tRNA backbones and anticodons, and genomic organization and patterns of variation that differ markedly from those of protein-coding regions. Individual genomes harbor remarkably different pools of tRNA genes predicting a wide range tRNA functional complements within a species. ### Competing Interest Statement The authors have declared no competing interest. U.S. National Science Foundation, https://ror.org/021nxhr62, EDGE 2319796, NSF-Simons Southeast Center for Mathematics and Biology (NSF DMS-1764406, Simons Foundation/SFARI 594594)
An outstanding question in the evolution of gene expression is the composition of the underlying regulatory architecture and the processes that shape it. Mutations affecting a gene's expression may reside locally in cis or distally in trans; the accumulation of these changes, their interactions, and their modes of inheritance influence how traits are expressed and how they evolve. Here, we interrogated gene expression variation in Caenorhabditis elegans, including the first allele-specific expression analysis in this system, capturing effects in cis and in trans that govern gene expression differences between the reference strain N2 and 7 wild strains. We observed extensive compensatory regulation, in which opposite effects in cis and trans at individual genes mitigate expression differences among strains, and that genes with expression differences exhibit strain specificity. As the genomic distance increased between N2 and each wild strain, the number of genes with expression differences also increased. We also report for the first time that expression-variable genes are lower expressed on average than genes without expression differences, a trend that may extend to humans and Drosophila melanogaster and may reflect the selection constraints that govern the universal anticorrelation between gene expression and rate of protein evolution. Together, these and other observed trends support the conclusion that many C. elegans genes are under stabilizing selection for expression level, but we also highlight outliers that may be biologically significant. To provide community access to our data, we introduce an easily accessible, interactive web application for gene-based queries: https://wildworm.biosci.gatech.edu/ase/.
Variation in gene expression is a feature of all living systems and has recently been characterized extensively among wild strains of the model organism Caenorhabditis elegans. To enable researchers to query gene expression and gene expression variation at any gene of interest, we have created a user-friendly web application that shares RNA-seq transcription data for 208 wild C. elegans strains generated by the Caenorhabditis Natural Diversity Resource (CaeNDR). Here, we describe the features of the web application and the details of the data and data processing underlying it. We hope that this website, wildworm.biosci.gatech.edu/cendrexp/ , will help C. elegans researchers better understand their favorite genes and strains.
AbstractAn outstanding question in the evolution of gene expression is the relative influence of neutral processes versus natural selection, including adaptive change driven by directional selection as well as stabilizing selection, which may include compensatory dynamics. These forces shape patterns of gene expression variation within and between species, including the regulatory mechanisms governing expression incisandtrans. In this study, we interrogate intraspecific gene expression variation among seven wildC. elegansstrains, with varying degrees of genomic divergence from the reference strain N2, leveraging this system’s unique advantages to comprehensively evaluate gene expression evolution. By capturing allele-specific and between-strain changes in expression, we characterize the regulatory architecture and inheritance mode of gene expression variation withinC. elegansand assess their relationship to nucleotide diversity, genome evolutionary history, gene essentiality, and other biological factors. We conclude that stabilizing selection is a dominant influence in maintaining expression phenotypes within the species, and the discovery that genes with higher overall expression tend to exhibit fewer expression differences supports this conclusion, as do widespread instances ofcisdifferences compensated intrans. Moreover, analyses of human expression data replicate our finding that higher expression genes have less variable expression. We also observe evidence for directional selection driving expression divergence, and that expression divergence accelerates with increasing genomic divergence. To provide community access to the data from this first analysis of allele-specific expression inC. elegans, we introduce an interactive web application, where users can submit gene-specific queries to view expression, regulatory pattern, inheritance mode, and other information:https://wildworm.biosci.gatech.edu/ase/.
The discovery that experimental delivery of dsRNA can induce gene silencing at target genes revolutionized genetics research, by both uncovering essential biological processes and creating new tools for developmental geneticists. However, the efficacy of exogenous RNA interference (RNAi) varies dramatically within the Caenorhabditis elegans natural population, raising questions about our understanding of RNAi in the lab relative to its activity and significance in nature. Here, we investigate why some wild strains fail to mount a robust RNAi response to germline targets. We observe diversity in mechanism: in some strains, the response is stochastic, either on or off among individuals, while in others, the response is consistent but delayed. Increased activity of the Argonaute PPW-1, which is required for germline RNAi in the laboratory strain N2, rescues the response in some strains but dampens it further in others. Among wild strains, genes known to mediate RNAi exhibited very high expression variation relative to other genes in the genome as well as allelic divergence and strain-specific instances of pseudogenization at the sequence level. Our results demonstrate functional diversification in the small RNA pathways in C. elegans and suggest that RNAi processes are evolving rapidly and dynamically in nature.
Though natural systems harbor genetic and phenotypic variation, research in model organisms is often restricted to a reference strain. Focusing on a reference strain yields a great depth of knowledge but potentially at the cost of breadth of understanding. Furthermore, tools developed in the reference context may introduce bias when applied to other strains, posing challenges to defining the scope of variation within model systems. Here, we evaluate how genetic differences among 5 wild Caenorhabditis elegans strains affect gene expression and its quantification, in general and after induction of the RNA interference (RNAi) response. Across strains, 34% of genes were differentially expressed in the control condition, including 411 genes that were not expressed at all in at least 1 strain; 49 of these were unexpressed in reference strain N2. Reference genome mapping bias caused limited concern: despite hyperdiverse hotspots throughout the genome, 92% of variably expressed genes were robust to mapping issues. The transcriptional response to RNAi was highly strain- and target-gene-specific and did not correlate with RNAi efficiency, as the 2 RNAi-insensitive strains showed more differentially expressed genes following RNAi treatment than the RNAi-sensitive reference strain. We conclude that gene expression, generally and in response to RNAi, differs across C. elegans strains such that the choice of strain may meaningfully influence scientific inferences. Finally, we introduce a resource for querying gene expression variation in this dataset at https://wildworm.biosci.gatech.edu/rnai/.
Here we provide a protocol for generating RNA from synchronized C. elegans F1s from crosses of multiple wild strains to a common reference strain (for example, for allele specific expression analysis). This protocol, which comprises 9 sequential days (6 active days) followed by RNA extraction, is optimized for synchronization of parental and offspring worms, retention of embryos, and timing. Others could easily adapt this protocol for crossing designs other than the specified multiple wild strains x common reference. This protocol accommodates 8 concurrent crosses and one extra parent with 3-4 active worm pickers on Days 3 and 6, and should accommodate increased or decreased numbers of crosses with commensurate personnel adjustments.
This dataset holds all non-GEO-hosted supplemental data files for manuscript "Beyond the reference: gene expression variation and transcriptional response to RNAi in C. elegans". Please see the linked preprint/publication for full details. The PDF _guide_to_datafiles.pdf gives details on the format and content of each of the included files.
A universal feature of living systems is that natural variation in genotype underpins variation in phenotype. Yet, research in model organisms is often constrained to a single genetic background, the reference strain. Further, genomic studies that do evaluate wild strains typically rely on the reference strain genome for read alignment, leading to the possibility of biased inferences based on incomplete or inaccurate mapping; the extent of reference bias can be difficult to quantify. As an intermediary between genome and organismal traits, gene expression is well positioned to describe natural variability across genotypes generally and in the context of environmental responses, which can represent complex adaptive phenotypes. C. elegans sits at the forefront of investigation into small-RNA gene regulatory mechanisms, or RNA interference (RNAi), and wild strains exhibit natural variation in RNAi competency following environmental triggers. Here, we examine how genetic differences among five wild strains affect the C. elegans transcriptome in general and after inducing RNAi responses to two germline target genes. Approximately 34% of genes were differentially expressed across strains; 411 genes were not expressed at all in at least one strain despite robust expression in others, including 49 genes not expressed in reference strain N2. Despite the presence of hyper-diverse hotspots throughout the C. elegans genome, reference mapping bias was of limited concern: over 92% of variably expressed genes were robust to mapping issues. Overall, the transcriptional response to RNAi was strongly strain-specific and highly specific to the target gene, and the laboratory strain N2 was not representative of the other strains. Moreover, the transcriptional response to RNAi was not correlated with RNAi phenotypic penetrance; the two germline RNAi incompetent strains exhibited substantial differential gene expression following RNAi treatment, indicating an RNAi response despite failure to reduce expression of the target gene. We conclude that gene expression, both generally and in response to RNAi, differs across C. elegans strains such that choice of strain may meaningfully influence scientific conclusions. To provide a public, easily accessible resource for querying gene expression variation in this dataset, we introduce an interactive website at https://wildworm.biosci.gatech.edu/rnai/ .
Recently published single-cell sequencing data from individual human sperm (n=41,189; 969–3377 cells from each of 25 donors) offer an opportunity to investigate questions of inheritance with improved statistical power, but require new methods tailored to these extremely low-coverage data (∼0.01× per cell). To this end, we developed a method, named rhapsodi, that leverages sparse gamete genotype data to phase the diploid genomes of the donor individuals, impute missing gamete genotypes, and discover meiotic recombination breakpoints, benchmarking its performance across a wide range of study designs. We then applied rhapsodi to the sperm sequencing data to investigate adherence to Mendel’s Law of Segregation, which states that the offspring of a diploid, heterozygous parent will inherit either allele with equal probability. While the vast majority of loci adhere to this rule, research in model and non-model organisms has uncovered numerous exceptions whereby ‘selfish’ alleles are disproportionately transmitted to the next generation. Evidence of such ‘transmission distortion’ (TD) in humans remains equivocal in part because scans of human pedigrees have been under-powered to detect small effects. After applying rhapsodi to the sperm data and scanning for evidence of TD, our results exhibited close concordance with binomial expectations under balanced transmission. Together, our work demonstrates that rhapsodi can facilitate novel uses of inferred genotype data and meiotic recombination events, while offering a powerful quantitative framework for testing for TD in other cohorts and study systems.
Mendel’s Law of Segregation states that the offspring of a diploid, heterozygous parent will inherit either allele with equal probability. While the vast majority of loci adhere to this rule, research in model and non-model organisms has uncovered numerous exceptions whereby “selfish” alleles are disproportionately transmitted to the next generation. Evidence of such “transmission distortion” (TD) in humans remains equivocal in part because scans of human pedigrees have been under-powered to detect small effects. Recently published single-cell sequencing data from individual human sperm ( n = 41,189; 969-3,377 cells from each of 25 donors) offer an opportunity to revisit this question with unprecedented statistical power, but require new methods tailored to extremely low-coverage data (∼0.01 × per cell). To this end, we developed a method, named rhapsodi, that leverages sparse gamete genotype data to phase the diploid genomes of the donor individuals, impute missing gamete genotypes, and discover meiotic recombination breakpoints, benchmarking its performance across a wide range of study designs. After applying rhapsodi to the sperm sequencing data, we then scanned the gametes for evidence of TD. Our results exhibited close concordance with binomial expectations under balanced transmission, in contrast to tenuous signals of TD that were previously reported in pedigree-based studies. Together, our work excludes the existence of even weak TD in this sample, while offering a powerful quantitative framework for testing this and related hypotheses in other cohorts and study systems.
Sperm-seq is a high-throughput, droplet-based single-sperm sequencing technology capable of generating thousands of cell-barcoded single-sperm sequencing libraries at one time. This protocol describes Sperm-seq library generation, featuring a full protocol for sperm preparation and suggestions for droplet-based sequencing methods from 10X Genomics to employ. Sperm preparation can be completed in half a day and the full protocol can be completed in 2-3 days, with several wait times and break points.
Meiosis, although essential for reproduction, is also variable and error-prone: rates of chromosome crossover vary among gametes, between the sexes, and among humans of the same sex, and chromosome missegregation leads to abnormal chromosome numbers (aneuploidy)1–8. To study diverse meiotic outcomes and how they covary across chromosomes, gametes and humans, we developed Sperm-seq, a way of simultaneously analysing the genomes of thousands of individual sperm. Here we analyse the genomes of 31,228 human gametes from 20 sperm donors, identifying 813,122 crossovers and 787 aneuploid chromosomes. Sperm donors had aneuploidy rates ranging from 0.01 to 0.05 aneuploidies per gamete; crossovers partially protected chromosomes from nondisjunction at the meiosis I cell division. Some chromosomes and donors underwent more-frequent nondisjunction during meiosis I, and others showed more meiosis II segregation failures. Sperm genomes also manifested many genomic anomalies that could not be explained by simple nondisjunction. Diverse recombination phenotypes—from crossover rates to crossover location and separation, a measure of crossover interference—covaried strongly across individuals and cells. Our results can be incorporated with earlier observations into a unified model in which a core mechanism, the variable physical compaction of meiotic chromosomes, generates interindividual and cell-to-cell variation in diverse meiotic phenotypes. Thousands of sperm genomes have been analysed with a new method called Sperm-seq, revealing interconnected meiotic variation at the single-cell and person-to-person levels, and suggesting chromosome compaction as a way to explain the relationships between diverse recombination phenotypes.
This repository holds crossover and aneuploidy data for 31,228 human sperm genomes sequenced with Sperm-seq. Data are described in the preprint/paper "Insights about variation in meiosis from 31,228 human sperm genomes." Several levels of data are available, e.g., each crossover (allcrossovers_hg38.txt.gz) and the numbers of crossovers detected per sperm cell (numbercrossoverspercell.txt.gz). Each data file is described in its own readme, so README_allgainsdivisionoforigin.txt explains the data present in allgainsdivisionoforigin.txt. All files are either text files or text files compressed via gzip. All data was generated as described in the preprint/paper, using scripts available in the companion repository (most recent version DOI: 10.5281/zenodo.3561080). This new version has updated sex chromosome ploidies after identifying a minor coding error affecting ~24 cells (sexchromosomeploidy.txt.gz) and includes more detail about possible crossover error modes (README_allcrossovers_hg38.txt.gz).
Many genomic segments vary in copy number among individuals of the same species, or between cancer and normal cells within the same person. Correctly measuring this copy number variation is critical for studying its genetic properties, its distribution in populations and its relationship to phenotypes. Droplet digital PCR (ddPCR) enables accurate measurement of copy number by partitioning a PCR reaction into thousands of nanoliter-scale droplets, so that a genomic sequence of interest-whose presence or absence in a droplet is determined by end-point fluorescence-can be digitally counted. Here, we describe how we analyze copy number variants using ddPCR and review the design of effective assays, the performance of ddPCR with those assays, the optimization of reactions, and the interpretation of data.
Type 2 diabetes (T2D) affects more than 415 million people worldwide, and its costs to the health care system continue to rise. To identify common or rare genetic variation with potential therapeutic implications for T2D, we analyzed and replicated genome-wide protein coding variation in a total of 8,227 individuals with T2D and 12,966 individuals without T2D of Latino descent. We identified a novel genetic variant in the IGF2 gene associated with ∼20% reduced risk for T2D. This variant, which has an allele frequency of 17% in the Mexican population but is rare in Europe, prevents splicing between IGF2 exons 1 and 2. We show in vitro and in human liver and adipose tissue that the variant is associated with a specific, allele-dosage–dependent reduction in the expression of IGF2 isoform 2. In individuals who do not carry the protective allele, expression of IGF2 isoform 2 in adipose is positively correlated with both incidence of T2D and increased plasma glycated hemoglobin in individuals without T2D, providing support that the protective effects are mediated by reductions in IGF2 isoform 2. Broad phenotypic examination of carriers of the protective variant revealed no association with other disease states or impaired reproductive health. These findings suggest that reducing IGF2 isoform 2 expression in relevant tissues has potential as a new therapeutic strategy for T2D, even beyond the Latin American population, with no major adverse effects on health or reproduction.
Schizophrenia is a heritable brain illness with unknown pathogenic mechanisms. Schizophrenia's strongest genetic association at a population level involves variation in the major histocompatibility complex (MHC) locus, but the genes and molecular mechanisms accounting for this have been challenging to identify. Here we show that this association arises in part from many structurally diverse alleles of the complement component 4 (C4) genes. We found that these alleles generated widely varying levels of C4A and C4B expression in the brain, with each common C4 allele associating with schizophrenia in proportion to its tendency to generate greater expression of C4A. Human C4 protein localized to neuronal synapses, dendrites, axons, and cell bodies. In mice, C4 mediated synapse elimination during postnatal development. These results implicate excessive complement activity in the development of schizophrenia and may help explain the reduced numbers of synapses in the brains of individuals with schizophrenia.