This protocol explains the default pipeline we recommend for analysing short-read data for Littorina marine snails. The pipeline was designed based on our experience processing WGS data, and careful consideration of the available tools. We are presenting this protocol as a suggestion of how Littorina sequencing data can be processed, and as a template for building more specific pipelines. Our hope is that through writing this protocol we can have more consistency among projects, leading to greater transparency for collaborators and providing a template which could be used to build pipelines for other organisms.
Key innovations are fundamental to biological diversification, but their genetic basis is poorly understood. A recent transition from egg-laying to live-bearing in marine snails ( Littorina spp.) provides the opportunity to study the genetic architecture of an innovation that has evolved repeatedly across animals. Individuals do not cluster by reproductive mode in a genome-wide phylogeny, but local genealogical analysis revealed numerous small genomic regions where all live-bearers carry the same core haplotype. Candidate regions show evidence for live-bearer–specific positive selection and are enriched for genes that are differentially expressed between egg-laying and live-bearing reproductive systems. Ages of selective sweeps suggest that live-bearer–specific alleles accumulated over more than 200,000 generations. Our results suggest that new functions evolve through the recruitment of many alleles rather than in a single evolutionary step.
Short-read whole-genome sequencing helps us to understand the genetic mechanisms underpinning evolution. To go from raw reads to variants requires a bioinformatic pipeline. Processing sequence data in bioinformatic pipelines is intricate, with several intermediate steps that have a plethora of parameters and options. A default variant-discovery pipeline saves time and effort and makes diverse studies more comparable. This protocol explains the default pipeline we recommend for analysing short-read data from marine snails in the genus Littorina. The pipeline was designed based on our experience processing WGS data, and careful consideration of the available tools. We are presenting this protocol as a suggestion of how Littorina sequencing data can be processed, which can be used as a template for building more specific pipelines. For more general information about filtering next generation sequencing data see Hemstrom et al. (2024). Our hope is that through writing this protocol we can have more consistency among projects, leading to greater transparency for collaborators and providing a template which could be used to build pipelines for other organisms.
Inversions are thought to play a key role in adaptation and speciation, suppressing recombination between diverging populations. Genes influencing adaptive traits cluster in inversions, and changes in inversion frequencies are associated with environmental differences. However, in many organisms, it is unclear if inversions are geographically and taxonomically widespread. The intertidal snail, Littorina saxatilis, is one such example. Strong associations between putative polymorphic inversions and phenotypic differences have been demonstrated between two ecotypes of L. saxatilis in Sweden and inferred elsewhere, but no direct evidence for inversion polymorphism currently exists across the species range. Using whole genome data from 107 snails, most inversion polymorphisms were found to be widespread across the species range. The frequencies of some inversion arrangements were significantly different among ecotypes, suggesting a parallel adaptive role. Many inversions were also polymorphic in the sister species, L. arcana, hinting at an ancient origin.
Key innovations are fundamental to biological diversification, but their genetic architecture is poorly understood. A recent transition from egg-laying to live-bearing in Littorina snails provides the opportunity to study the architecture of an innovation that has evolved repeatedly in animals. Samples do not cluster by reproductive mode in a genome-wide phylogeny, but local genealogical analysis revealed numerous genomic regions where all live-bearers carry the same core haplotype. Associated regions show evidence for live-bearer-specific positive selection, and are enriched for genes that are differentially expressed between egg-laying and live-bearing reproductive systems. Ages of selective sweeps suggest live-bearing alleles accumulated gradually, involving selection at different times in the past. Our results suggest that innovation can have a polygenic basis, and that novel functions can evolve gradually, rather than in a single step.
Comparing genome scans among species is a powerful approach for investigating the patterns left by evolutionary processes. In particular, this offers a way to detect candidate genes that drive convergent evolution. We compared genome scan results to investigate if patterns of genetic diversity and divergence are shared among divergent species within the stickleback order (Gasterosteiformes): the threespine stickleback (Gasterosteus aculeatus), ninespine stickleback (Pungitius pungitus), and tubesnout (Aulorhynchus flavidus). Populations were sampled from the southern and northern edges of each species' range, to identify patterns associated with latitudinal changes in genetic diversity. Weak correlations in genetic diversity (F ST and expected heterozygosity) and three different patterns in the genomic landscape were found among these species. Additionally, no candidate genes for convergent evolution were detected. This is a counterexample to the growing number of studies that have shown overlapping genetic patterns, demonstrating that genome scan comparisons can be noisy due to the effects of several interacting evolutionary forces.
Normalization is the first critical step in microbiome sequencing data analysis used to account for variable library sizes. Current RNA-Seq based normalization methods that have been adapted for microbiome data fail to consider the unique characteristics of microbiome data, which contain a vast number of zeros due to the physical absence or under-sampling of the microbes. Normalization methods that specifically address the zero inflation remain largely undeveloped. Here we propose GMPR-a simple but effective normalization method-for zero-inflated sequencing data such as microbiome data. Simulation studies and real datasets analyses demonstrate that the proposed method is more robust than competing methods, leading to more powerful detection of differentially abundant taxa and higher reproducibility of the relative abundances of taxa.
Recombination can impede ecological speciation with gene flow by mixing locally adapted genotypes with maladapted migrant genotypes from a divergent population. In such a scenario, suppression of recombination can be selectively favoured. However, in finite populations evolving under the influence of random genetic drift, recombination can also facilitate adaptation by reducing Hill-Robertson interference between loci under selection. In this case, increased recombination rates can be favoured. Although these two major effects on recombination have been studied individually, their joint effect on ecological speciation with gene flow remains unexplored. Using a mathematical model, we investigated the evolution of recombination rates in two finite populations that exchange migrants while adapting to contrasting environments. Our results indicate a two-step dynamic where increased recombination is first favoured (in response to the Hill-Robertson effect), and then disfavoured, as the cost of recombining locally with maladapted migrant genotypes increases over time (the maladaptive gene flow effect). In larger populations, a stronger initial benefit for recombination was observed, whereas high migration rates intensify the long-term cost of recombination. These dynamics may have important implications for our understanding of the conditions that facilitate incipient speciation with gene flow and the evolution of recombination in finite populations.