The current economics of scientific publishing reveal a profound imbalance: academia pays prices far exceeding the actual costs of publication. Rather than supporting research, much of this expenditure sustains the profits of a few dominant commercial publishers. Transitioning to responsible publishing is a collective challenge that requires raising awareness among scientists about the problem and the solutions available. We present DAFNEE, a database of academia-friendly journals in ecology, evolutionary biology and archaeology (https://dafnee.isem-evolution.fr/). DAFNEE includes information on over 600 journals (co)run by academic or non-profit institutions, aiming at helping to keep publishing funds within the academic community. The database details these journal's business models, article processing charges, citation rates and partnerships. We show that DAFNEE journals compare favourably to non-DAFNEE ones in terms of editorial and financial policy, while offering similar citation rates. Finally, we offer several recommendations aimed at encouraging authors, reviewers, and evaluators to adopt more responsible publishing practices.
GC-biased gene conversion (gBGC) is a widespread evolutionary force associated with meiotic recombination that favours the accumulation of deleterious AT to GC substitutions in proteins, moving them away from their fitness optimum. In many mammals recombination hotspots have a rapid turnover, leading to episodic gBGC, with the accumulation of deleterious mutations stopping when the recombination hotspot dies. Selection is therefore expected to act to repair the damage caused by gBGC episodes through compensatory evolution. However, this process has never been studied or quantified so far. Here, we analysed the nucleotide substitution pattern in coding sequences of a highly diversified group of Murinae rodents. Using phylogenetic analyses of about 70,000 coding exons, we identified numerous exon-specific, lineage-specific gBGC episodes, characterised by a clustering of synonymous AT to GC substitutions and by an increasing rate of non-synonymous AT to GC substitutions, many of which are potentially deleterious. Analysing the molecular evolution of the affected exons in downstream lineages, we found evidence for pervasive compensatory evolution after deleterious gBGC episodes. Compensation appears to occur rapidly after the end of the episode, and to be driven by the standing genetic variation rather than new mutations. Our results demonstrate the impact of gBGC on the evolution of amino-acid sequences, and underline the key role of epistasis in protein adaptation. This study contributes to a growing body of literature emphasizing that adaptive mutations, which arise in response to environmental changes, are just one subset of beneficial mutations, alongside mutations resulting from oscillations around the fitness optimum. ### Competing Interest Statement The authors have declared no competing interest.
In many eukaryotes, meiotic recombination occurs preferentially at discrete sites, called recombination hotspots. In various lineages, recombination hotspots are located in regions with promoter-like features and are evolutionarily stable. Conversely, in some mammals, hotspots are driven by PRDM9 that targets recombination away from promoters. Paradoxically, PRDM9 induces the self-destruction of its targets and this triggers an ultra-fast evolution of mammalian hotspots. PRDM9 is ancestral to all animals, suggesting a critical importance for the meiotic program, but has been lost in many lineages with surprisingly little effect on meiosis success. However, it is unclear whether the function of PRDM9 described in mammals is shared by other species. To investigate this, we analyzed the recombination landscape of several salmonids, the genome of which harbors one full-length PRDM9 and several truncated paralogs. We identified recombination initiation sites in Oncorhynchus mykiss by mapping meiotic DNA double-strand breaks (DSBs). We found that DSBs clustered at hotspots positioned away from promoters, enriched for the H3K4me3 and H3K36me3 and the location of which depended on the genotype of full-length Prdm9. We observed a high level of polymorphism in the zinc finger domain of full-length Prdm9, indicating diversification driven by positive selection. Moreover, population-scaled recombination maps in O. mykiss, Oncorhynchus kisutch and Salmo salar revealed a rapid turnover of recombination hotspots caused by PRDM9 target motif erosion. Our results imply that PRDM9 function is conserved across vertebrates and that the peculiar evolutionary runaway caused by PRDM9 has been active for several hundred million years.
A recommendation of: Florent Sylvestre, Nadia Aubin-Horth, Louis Bernatchez Sex-biased gene expression across tissues reveals unexpected differentiation in the gills of the threespine stickleback https://doi.org/10.1101/2024.06.09.597944
GC-biased gene conversion (gBGC) is a widespread evolutionary force associated with meiotic recombination that favors the accumulation of deleterious AT to GC substitutions in proteins, moving them away from their fitness optimum. In many mammals, recombination hotspots have a rapid turnover, leading to episodic gBGC, with the accumulation of deleterious mutations stopping when the recombination hotspot dies. Selection is therefore expected to act to repair the damage caused by gBGC episodes through compensatory evolution. However, this process has never been studied or quantified so far. Here, we analyzed the nucleotide substitution pattern in coding sequences of a highly diversified group of Murinae rodents. Using phylogenetic analyses of about 70,000 coding exons, we identified numerous exon-specific, lineage-specific gBGC episodes, characterized by a clustering of synonymous AT to GC substitutions and by an increasing rate of nonsynonymous AT to GC substitutions, many of which are potentially deleterious. Analyzing the molecular evolution of the affected exons in downstream lineages, we found evidence for pervasive compensatory evolution after deleterious gBGC episodes. Compensation appears to occur rapidly after the end of the episode and to be driven by the standing genetic variation rather than new mutations. Our results demonstrate the impact of gBGC on the evolution of amino-acid sequences and underline the key role of epistasis in protein adaptation. This study contributes to a growing body of literature emphasizing that adaptive mutations, which arise in response to environmental changes, are just 1 subset of beneficial mutations, alongside mutations resulting from oscillations around the fitness optimum.
The evolution of gene families is complex, involving gene-level evolutionary events such as gene duplication, horizontal gene transfer, and gene loss, and other processes such as incomplete lineage sorting (ILS). Because of this, topological differences often exist between gene trees and species trees. A number of models have been recently developed to explain these discrepancies, the most realistic of which attempts to consider both gene-level events and ILS. When unified in a single model, the interaction between ILS and gene-level events can cause polymorphism in gene copy number, which we refer to as copy number hemiplasy (CNH). In this paper, we extend the Wright-Fisher process to include duplications and losses over several species, and show that the probability of CNH for this process can be significant. We study how well two unified models-multilocus multispecies coalescent (MLMSC), which models CNH, and duplication, loss, and coalescence (DLCoal), which does not-approximate the Wright-Fisher process with duplication and loss. We then study the effect of CNH on gene family evolution by comparing MLMSC and DLCoal. We generate comparable gene trees under both models, showing significant differences in various summary statistics; most importantly, CNH reduces the number of gene copies greatly. If this is not taken into account, the traditional method of estimating duplication rates (by counting the number of gene copies) becomes inaccurate. The simulated gene trees are also used for species tree inference with the summary methods ASTRAL and ASTRAL-Pro, demonstrating that their accuracy, based on CNH-unaware simulations calibrated on real data, may have been overestimated.
The neutral and nearly neutral theories, introduced more than 50 yr ago, have raised and still raise passionate discussion regarding the forces governing molecular evolution and their relative importance. The debate, initially focused on the amount of within-species polymorphism and constancy of the substitution rate, has spread, matured, and now underlies a wide range of topics and questions. The neutralist/selectionist controversy has structured the field and influences the way molecular evolutionary scientists conceive their research.
Gene flow and incomplete lineage sorting are two distinct sources of phylogenetic conflict, i.e., gene trees that differ in topology from each other and from the species tree. Distinguishing between the two processes is a key objective of current evolutionary genomics. This is most often pursued via the so-called ABBA-BABA type of method, which relies on a prediction of symmetry of gene tree discordance made by the incomplete lineage sorting hypothesis. Gene flow, however, need not be asymmetric, and when it is not, ABBA-BABA approaches do not properly measure the prevalence of gene flow. I introduce Aphid, an approximate maximum-likelihood method aimed at quantifying the sources of phylogenetic conflict via topology and branch length analysis of three-species gene trees. Aphid draws information from the fact that gene trees affected by gene flow tend to have shorter branches, and gene trees affected by incomplete lineage sorting longer branches, than the average gene tree. Accounting for the among-loci variance in mutation rate and gene flow time, Aphid returns estimates of the speciation times and ancestral effective population size, and a posterior assessment of the contribution of gene flow and incomplete lineage sorting to the conflict. Simulations suggest that Aphid is reasonably robust to a wide range of conditions. Analysis of coding and non-coding data in primates illustrates the potential of the approach and reveals that a substantial fraction of the human/chimpanzee/gorilla phylogenetic conflict is due to ancient gene flow. Aphid also predicts older speciation times and a smaller estimated effective population size in this group, compared to existing analyses assuming no gene flow.
In asexual animals, female meiosis is modified to produce diploid oocytes. If meiosis still involves recombination, this is expected to lead to a rapid loss of heterozygosity, with adverse effects on fitness. Many asexuals, however, have a heterozygous genome, the underlying mechanisms being most often unknown. Cytological and population genomic analyses in the nematode Mesorhabditis belari revealed another case of recombining asexual being highly heterozygous genome-wide. We demonstrated that heterozygosity is maintained despite recombination because the recombinant chromatids of each chromosome pair cosegregate during the unique meiotic division. A theoretical model confirmed that this segregation bias is necessary to account for the observed pattern and likely to evolve under a wide range of conditions. Our study uncovers an unexpected type of non-Mendelian genetic inheritance involving cosegregation of recombinant chromatids.
Several studies have highlighted the presence of contaminated entries in public sequence repositories, calling for special attention to the associated metadata. Here, we propose and evaluate a fast and efficient kmer-based approach to assess the degree of mislabeling or contamination. We applied it to high-throughput whole-genome raw sequence data for 236 Ind-Seq and 22 Pool-Seq samples of the invasive species Drosophila suzukii. We first used CLARK software to build a dictionary of species-discriminating kmers from the curated assemblies of 29 target drosophilid species (including D. melanogaster, D. simulans, D. subpulchrella or D. biarmipes) and 12 common drosophila pathogens and commensals (including Wolbachia). Counting the number of k-mers composing each query sample sequence that matched a discriminating k-mer from the dictionary provided a simple criterion for assignment to target species and evaluation of the entire sample. Analyses of a wide range of samples, representative of both target and other drosophilid species, demonstrated very good performance of the proposed approach, both in terms of run time and accuracy of sequence assignment. Of the 236 D. suzukii individuals, five were reassigned to D. simulans and eleven to D. subpulchrella. Another four showed moderate to substantial microbial contamination. Similarly, among the 22 Pool-Seq samples analyzed, two from the native range were found to be contaminated with 1 and 7 D. subpulchrella individuals, respectively (out of 50), and one from Europe was found to be contaminated with 5 to 6 D. immigrans individuals (out of 100). Overall, the present analysis allowed the definition of a large curated dataset consisting of >60 population samples representative of the worldwide genetic diversity, which may be valuable for further population genetics studies on D. suzukii. More generally, while we advocate careful sample identification and verification prior to sequencing, the proposed framework is simple and computationally efficient enough to be included as a routine post-hoc quality check prior to any data analysis and prior to data submission to public repositories.
Knowledge of recombination rate variation along the genome provides important insights into genome and phenotypic evolution. Population genomic approaches offer an attractive way to infer the population-scaled recombination rate ρ=4 N e r using the linkage disequilibrium information contained in DNA sequence polymorphism data. Such methods have been used in a broad range of plant and animal species to build genome-wide recombination maps. However, the reliability of these inferences has only been assessed under a restrictive set of conditions. Here, we evaluate the ability of one of the most widely used coalescent-based programs, LDhelmet , to infer a genomic landscape of recombination with the biological characteristics of a human-like landscape including hotspots. Using simulations, we specifically assessed the impact of methodological (sample size, phasing errors, block penalty) and evolutionary parameters (effective population size ( N e ), demographic history, mutation to recombination rate ratio) on inferred map quality. We report reasonably good correlations between simulated and inferred landscapes, but point to limitations when it comes to detecting recombination hotspots. False positive and false negative hotspots considerably confound fine-scale patterns of inferred recombination under a wide range of conditions, particularly when N e is small and the mutation/recombination rate ratio is low, to the extent that maps inferred from populations sharing the same recombination landscape appear uncorrelated. We thus address a message of caution for the users of these approaches, at least for genomes with complex recombination landscapes such as in humans.
SignificanceThe dynamics of deleterious variation under contrasting demographic scenarios remain poorly understood in spite of their relevance in evolutionary and conservation terms. Here we apply a genomic approach to study differences in the burden of deleterious alleles between the endangered Iberian lynx (Lynx pardinus) and the widespread Eurasian lynx (Lynx lynx). Our analysis unveils a significantly lower deleterious burden in the former species that should be ascribed to genetic purging, that is, to the increased opportunities of selection against recessive homozygotes due to the inbreeding caused by its smaller population size, as illustrated by our analytical predictions. This research provides theoretical and empirical evidence on the evolutionary relevance of genetic purging under certain demographic conditions.
The shift from sexual reproduction to parthenogenesis has occurred repeatedly in animals, but how the loss of sex affects genome evolution remains poorly understood. We generated reference genomes for five independently evolved parthenogenetic species in the stick insect genus Timema and their closest sexual relatives. Using these references and population genomic data, we show that parthenogenesis results in an extreme reduction of heterozygosity and often leads to genetically uniform populations. We also find evidence for less effective positive selection in parthenogenetic species, suggesting that sex is ubiquitous in natural populations because it facilitates fast rates of adaptation. Parthenogenetic species did not show increased transposable element (TE) accumulation, likely because there is little TE activity in the genus. By using replicated sexual-parthenogenetic comparisons, our study reveals how the absence of sex affects genome evolution in natural populations, providing empirical support for the negative consequences of parthenogenesis as predicted by theory.
The academic journal publishing model is deeply unethical: today, a few major, for-profit conglomerates control more than 50% of all articles in the natural sciences and social sciences, driving subscription and open-access publishing fees above levels that can be sustainably maintained by publicly funded universities, libraries, and research institutions worldwide. About a third of the costs paid for publishing papers is profit for these dominant publishers' shareholders, and about half of them covers costs to keep the system running, including lobbying, marketing fees, and paywalls. The paywalls in turn restrict access of scientific outputs, preventing them from being freely shared with the public and other researchers. Thus, money that the public is told goes into science is actually being funneled away from it, or used to limit access to it. Alternatives to this model exist and have increased in popularity in recent years, including diamond open-access journals and community-driven recommendation models. These are free of charge for authors and minimize costs for institutions and agencies, while making peer-reviewed scientific results publicly accessible. However, for-profit publishing agents have made change difficult, by co-opting open-access schemes and creating journal-driven incentives that prevent an effective collective transition away from profiteering. Here, we give a brief overview of the current state of the academic publishing system, including its most important systemic problems. We then describe alternative systems. We explain the reasons why the move toward them can be perceived as costly to individual researchers, and we demystify common roadblocks to change. Finally, in view of the above, we provide a set of guidelines and recommendations that academics at all levels can implement, in order to enable a more rapid and effective transition toward ethical publishing.
Hybridization occupies a central role in many fundamental evolutionary processes, such as speciation or adaptation. Yet, despite its pivotal importance in evolution, little is known about the actual prevalence and distribution of hybridization across the tree of life. Here we develop and implement a new statistical method enabling the detection of F1 hybrids from single-individual genome sequencing data. Using simulations and sequencing data from known hybrid systems, we first demonstrate the specificity of the method, and identify its statistical limits. Next, we showcase the method by applying it to available sequencing data from more than 1500 species of Arthropods, including Hymenoptera, Hemiptera, Coleoptera, Diptera and Archnida. Among these taxa, we find Hymenoptera, and especially ants, to display the highest number of candidate F1 hybrids, suggesting higher rates of recent hybridization in these groups. The prevalence of F1 hybrids was heterogeneously distributed across ants, with taxa including many candidates tending to harbor specific ecological and life history traits. This work shows how large-scale genomic comparative studies of recent hybridization can be implemented, uncovering the determinants of hybridization frequency across whole taxa.
Genome sequence (fasta) files and annotation (gff) files for ten Timema species: T. bartmani, T. cristinae, T. poppensis, T. californicum, T. podura, T. tahoe, T. monikensis, T. douglasi, T. shepardi, and T. genevievae.Species are abbreviated as follows: Tbi = T. bartmani, Tce = T. cristinae, Tps = T. poppensis, Tcm = T. californicum, Tpa = T. podura, Tte = T. tahoe, Tms = T. monikensis, Tdi = T. douglasi, Tsi = T. shepardi, and Tge = T. genevievaeFor details of assembly and annotation see:Jaron, K. S*., Parker, D. J*., Anselmetti, Y., Tran Van, P. T., Bast, J., Dumas, Z., Figuet, E., François, C. M., Hayward, K., Rossier, V., Simion, P., Robinson-Rechavi, M., Galtier, N., Schwander, T. 2021. Convergent consequences of parthenogenesis on stick insect genomes. bioRxiv. doi: https://doi.org/10.1101/2020.11.20.391540 File list:Tbi_b3v08.fasta = T. bartmani genome sequence fileTbi_b3v08.max_arth_b2g_droso_b2g.gff = T. bartmani genome annotation fileTce_b3v08.fasta = T. cristinae genome sequence fileTce_b3v08.max_arth_b2g_droso_b2g.gff = T. cristinae genome annotation fileTcm_b3v08.fasta = T. bartmani genome sequence fileTcm_b3v08.max_arth_b2g_droso_b2g.gff = T. californicum genome annotation fileTdi_b3v08.fasta = T. douglasi genome sequence fileTdi_b3v08.max_arth_b2g_droso_b2g.gff = T. douglasi genome annotation fileTge_b3v08.fasta = T. genevievae genome sequence fileTge_b3v08.max_arth_b2g_droso_b2g.gff = T. genevievae genome annotation fileTms_b3v08.fasta = T. monikensis genome sequence fileTms_b3v08.max_arth_b2g_droso_b2g.gff = T. monikensis genome annotation fileTpa_b3v08.fasta = T. podura genome sequence fileTpa_b3v08.max_arth_b2g_droso_b2g.gff = T. podura genome annotation fileTps_b3v08.fasta = T. poppensis genome sequence fileTps_b3v08.max_arth_b2g_droso_b2g.gff = T. poppensis genome annotation fileTsi_b3v08.fasta = T. shepardi genome sequence fileTsi_b3v08.max_arth_b2g_droso_b2g.gff = T. shepardi genome annotation fileTte_b3v08.fasta = T. tahoe genome sequence fileTte_b3v08.max_arth_b2g_droso_b2g.gff = T. tahoe genome annotation file
Reconstructing ancestral characters on a phylogeny is an arduous task because the observed states at the tips of the tree correspond to a single realization of the underlying evolutionary process. Recently, it was proposed that ancestral traits can be indirectly estimated with the help of molecular data, based on the fact that life history traits influence substitution rates. Here we challenge these new approaches in the Cetartiodactyla, a clade of large mammals which, according to paleontology, derive from small ancestors. Analysing transcriptome data in 41 species, of which 22 were newly sequenced, we provide a dated phylogeny of the Cetartiodactyla and report a significant effect of body mass on the overall substitution rate, the synonymous vs. non-synonymous substitution rate and the dynamics of GC-content. Our molecular comparative analysis points toward relatively small Cetartiodactyla ancestors, in agreement with the fossil record, even though our data set almost exclusively consists of large species. This analysis demonstrates the potential of phylogenomic methods for ancestral trait reconstruction and gives credit to recent suggestions that the ancestor to placental mammals was a relatively large and long-lived animal.
Conservation policy in the giant Galpagos tortoise, an iconic endangered animal, has been assisted by genetic markers for 15 years: a dozen loci have been used to delineate thirteen (sub)species, between which hybridization is prevented. Here, comparative reanalysis of a previously published NGS data set reveals a conflict with traditional markers. Genetic diversity and population substructure in the giant Galpagos tortoise are found to be particularly low, questioning the genetic relevance of current conservation practices. Further examination of giant Galapagos tortoise population genomics is critically needed.
GC-biased gene conversion (gBGC) is a molecular evolutionary force that favours GC over AT alleles irrespective of their fitness effect. Quantifying the variation in time and across genomes of its intensity is key to properly interpret patterns of molecular evolution. In particular, the existing literature is unclear regarding the relationship between gBGC strength and species effective population size, Ne. Here we analysed the nucleotide substitution pattern in coding sequences of closely related species of mammals, thus accessing a high resolution map of the intensity of gBGC. Our maximum likelihood approach shows that gBGC is pervasive, highly variable among species and genes, and of strength positively correlated with Ne in mammals. We estimate that gBGC explains up to 60% of the total amount of synonymous ATGC substitutions. We show that the fine-scale analysis of gBGC-induced nucleotide substitutions has the potential to inform on various aspects of molecular evolution, such as the distribution of fitness effects of mutations and the dynamics of recombination hotspots.
Ostracods are one of the oldest crustacean groups with an excellent fossil record and high importance for phylogenetic analyses but genome resources for this class are still lacking. We have successfully assembled and annotated the first reference genomes for three species of nonmarine ostracods; two with obligate sexual reproduction (Cyprideis torosa and Notodromas monacha) and the putative ancient asexual Darwinula stevensoni. This kind of genomic research has so far been impeded by the small size of most ostracods and the absence of genetic resources such as linkage maps or BAC libraries that were available for other crustaceans. For genome assembly, we used an Illumina-based sequencing technology, resulting in assemblies of similar sizes for the three species (335-382 Mb) and with scaffold numbers and their N50 (19-56 kb) in the same orders of magnitude. Gene annotations were guided by transcriptome data from each species. The three assemblies are relatively complete with BUSCO scores of 92-96. The number of predicted genes (13,771-17,776) is in the same range as Branchiopoda genomes but lower than in most malacostracan genomes. These three reference genomes from nonmarine ostracods provide the urgently needed basis to further develop ostracods as models for evolutionary and ecological research.