Animal genomes exhibit a remarkable variation in size, but the evolutionary forces responsible for such variation are still debated. As the effective population size (Nee) reflects the intensity of genetic drift, it is expected to be a key determinant of the fixation rate of nearly-neutral mutations. Accordingly, the Mutational Hazard Hypothesis postulates lineages with low Nee to have bigger genome sizes due to the accumulation of slightly deleterious transposable elements (TEs), and those with high Nee to maintain streamlined genomes as a consequence of a more effective selection against TEs. However, the existence of both empirical confirmation and refutation using different methods and different scales precludes its general validation. Using high-quality public data, we estimated genome size, TE content, and rate of non-synonymous to synonymous substitutions (dN/dS) as Nee proxy for 807 species including vertebrates, molluscs, and insects. After collecting available life-history traits, we tested the associations among population size proxies, TE content, and genome size, while accounting for phylogenetic non-independence. Our results confirm TEs as major drivers of genome size variation, and endorse life-history traits and dN/dS as reliable proxies for Nee. However, we do not find any evidence for increased drift to result in an accumulation of TEs across animals. Within more closely related clades, only a few isolated and weak associations emerge in fishes and birds. Our results outline a scenario where TE dynamics vary according to lineage-specific patterns, lending no support for genetic drift as the predominant force driving long-term genome size evolution in animals.
A recommendation of: Jordan Teoli, Miriam Merenciano, Marie Fablet, Anamaria Necsulea, Daniel Siqueira-de-Oliveira, Alessandro Brandulas-Cammarata, Audrey Labalme, Hervé Lejeune, Jean-François Lemaitre, François Gueyffier, Damien Sanlaville, Claire Bardel, Cristina Vieira, Gabriel AB Marais, Ingrid Plotton Transposable element expression with variation in sex chromosome number supports a toxic Y effect on human longevity https://doi.org/10.1101/2023.08.03.550779
Retrotransposons can cause somatic genome variation in the human nervous system, which is hypothesized to have relevance to brain development and neuropsychiatric disease. However, the detection of individual somatic mobile element insertions presents a difficult signal-to-noise problem. Using a machine-learning method (RetroSom) and deep whole-genome sequencing, we analyzed L1 and Alu retrotransposition in sorted neurons and glia from human brains. We characterized two brain-specific L1 insertions in neurons and glia from a donor with schizophrenia. There was anatomical distribution of the L1 insertions in neurons and glia across both hemispheres, indicating retrotransposition occurred during early embryogenesis. Both insertions were within the introns of genes ( CNNM2 and FRMD4A ) inside genomic loci associated with neuropsychiatric disorders. Proof-of-principle experiments revealed these L1 insertions significantly reduced gene expression. These results demonstrate that RetroSom has broad applications for studies of brain development and may provide insight into the possible pathological effects of somatic retrotransposition.
Transposable Element MOnitoring with LOng-reads (TrEMOLO) is a new software that combines assembly- and mapping-based approaches to robustly detect genetic elements called transposable elements (TEs). Using high- or low-quality genome assemblies, TrEMOLO can detect most TE insertions and deletions and estimate their allele frequency in populations. Benchmarking with simulated data revealed that TrEMOLO outperforms other state-of-the-art computational tools. TE detection and frequency estimation by TrEMOLO were validated using simulated and experimental datasets. Therefore, TrEMOLO is a comprehensive and suitable tool to accurately study TE dynamics. TrEMOLO is available under GNU GPL3.0 at https://github.com/DrosophilaGenomeEvolution/TrEMOLO .
Efficiently detecting genomic structural variants (SVs) is a key step to grasp the “missing heritability” underlying complex traits involved in major evolutionary processes such as speciation, phenotypic plasticity, and adaptive responses. Yet, the SV-based genotype/trait association studies are still largely overlooked mainly due to the lack of reliable detection methods. Here, we present a random forest (RF) method for accurate deletion identification: RF4Del. By relying on the analysis of the mapping profiles, data already available in most sequencing projects, RF4Del can easily and quickly call deletions. Several classic and ensemble learning strategies were carefully evaluated using proper benchmark data. RF4Del was trained and tested on simulated data from the model species Drosophila melanogaster to detect deletions. The model consists of 13 features extracted from a mapping file. We show that RF4Del outperforms established SV callers (DELLY, Pindel) with higher overall performance (F1-score > 0.75; 6x-12x sequence coverage) and is less affected by low sequencing coverage and deletion size variations. RF4Del could learn from a compilation of sequence patterns linked to a given SV. Such models can then be combined to form a learning system able to detect all types of SVs in a given genome, beyond the one used in our study. https://github.com/alvesrcoo/eletric-scheep
Structural variations (SVs) constitute a significant source of genetic variability in virus genomes. Yet knowledge about SV variability and contribution to the evolutionary process in large double-stranded (ds)DNA viruses is limited. Cyprinid herpesvirus 3 (CyHV-3), also commonly known as koi herpesvirus (KHV), has the largest dsDNA genome within herpesviruses. This virus has become one of the biggest threats to common carp and koi farming, resulting in high morbidity and mortalities of fishes, serious environmental damage, and severe economic losses. A previous study analyzing CyHV-3 virulence evolution during serial passages onto carp cell cultures suggested that CyHV-3 evolves, at least in vitro, through an assembly of haplotypes that alternatively become dominant or under-represented. The present study investigates the SV diversity and dynamics in CyHV-3 genome during 99 serial passages in cell culture using, for the first time, ultra-deep whole-genome and amplicon-based sequencing. The results indicate that KHV polymorphism mostly involves SVs. These SVs display a wide distribution along the genome and exhibit high turnover dynamics with a clear bias towards inversion and deletion events. Analysis of the pathogenesis-associated ORF150 region in ten intermediate cell passages highlighted mainly deletion, inversion and insertion variations that deeply altered the structure of ORF150. Our findings indicate that SV turnovers and defective genomes represent key drivers in the viral population dynamics and in vitro evolution of KHV. Thus, the present study can contribute to the basic research needed to design safe live-attenuated vaccines, classically obtained by viral attenuation after serial passages in cell culture.
In 2020, the world faced the Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) pandemic that drastically altered people's lives. Since then, many countries have been forced to suspend public gatherings, leading to many conference cancellations, postponements, or reorganizations. Switching from a face-to-face to a remote conference became inevitable and the ultimate solution to sustain scientific exchanges at the national and the international levels. The same year, as a committee, we were in charge of organizing the major French annual conference that covers all computational biology areas: The "Journées Ouvertes en Biologie, Informatique et Mathématiques" (JOBIM). Despite the health crisis, we succeeded in changing the conference format from face to face to remote in a very short amount of time. Here, we propose 10 simple rules based on this experience to modify a conference format in an optimized and cost-effective way. In addition to the suggested rules, we decided to emphasize an unexpected benefit of this situation: a significant reduction in greenhouse gas (GHG) emissions related to travel for scientific conference attendance. We believe that even once the SARS-CoV-2 crisis is over, we collectively will have an opportunity to think about the way we approach such scientific events over the longer term.
In this paper, we investigate througth a premilinary study the influence of repeat elements during the assembly process. We analyze the link between the presence and the nature of one type of repeat element, called transposable element (TE) and misassembly events in genome assemblies. We propose to improve assemblies by taking into account the presence of repeat elements, including TEs, during the scaffolding step. We analyze the results and relate the misassemblies to TEs before and after correction.
Background Plasmids are mobile genetic elements that often carry accessory genes, and are vectors for horizontal transfer between bacterial genomes. Plasmid detection in large genomic datasets is crucial to analyze their spread and quantify their role in bacteria adaptation and particularly in antibiotic resistance propagation. Bioinformatics methods have been developed to detect plasmids. However, they suffer from low sensitivity (i.e., most plasmids remain undetected) or low precision (i.e., these methods identify chromosomes as plasmids), and are overall not adapted to identify plasmids in whole genomes that are not fully assembled (contigs and scaffolds). Results We developed PlasForest, a homology-based random forest classifier identifying bacterial plasmid sequences in partially assembled genomes. Without knowing the taxonomical origin of the samples, PlasForest identifies contigs as plasmids or chromosomes with a F1 score of 0.950. Notably, it can detect 77.4% of plasmid contigs below 1 kb with 2.8% of false positives and 99.9% of plasmid contigs over 50 kb with 2.2% of false positives. Conclusions PlasForest outperforms other currently available tools on genomic datasets by being both sensitive and precise. The performance of PlasForest on metagenomic assemblies are currently well below those of other k-mer-based methods, and we discuss how homology-based approaches could improve plasmid detection in such datasets.
Most of the current knowledge on the genetic basis of adaptive evolution is based on the analysis of single nucleotide polymorphisms (SNPs). Despite increasing evidence for their causal role, the contribution of structural variants to adaptive evolution remains largely unexplored. In this work, we analyzed the population frequencies of 1,615 Transposable Element (TE) insertions annotated in the reference genome of Drosophila melanogaster, in 91 samples from 60 worldwide natural populations. We identified a set of 300 polymorphic TEs that are present at high population frequencies, and located in genomic regions with high recombination rate, where the efficiency of natural selection is high. The age and the length of these 300 TEs are consistent with relatively young and long insertions reaching high frequencies due to the action of positive selection. Besides, we identified a set of 21 fixed TEs also likely to be adaptive. Indeed, we, and others, found evidence of selection for 84 of these reference TE insertions. The analysis of the genes located nearby these 84 candidate adaptive insertions suggested that the functional response to selection is related with the GO categories of response to stimulus, behavior, and development. We further showed that a subset of the candidate adaptive TEs affects expression of nearby genes, and five of them have already been linked to an ecologically relevant phenotypic effect. Our results provide a more complete understanding of the genetic variation and the fitness-related traits relevant for adaptive evolution. Similar studies should help uncover the importance of TE-induced adaptive mutations in other species as well.
Motivation: Transposable elements (TEs) constitute a significant proportion of the majority of genomes sequenced to date. TEs are responsible for a considerable fraction of the genetic variation within and among species. Accurate genotyping of TEs in genomes is therefore crucial for a complete identification of the genetic differences among individuals, populations and species. Results: In this work, we present a new version of T-lex, a computational pipeline that accurately genotypes and estimates the population frequencies of reference TE insertions using short-read high-throughput sequencing data. In this new version, we have re-designed the T-lex algorithm to integrate the BWA-MEM short-read aligner, which is one of the most accurate short-read mappers and can be launched on longer short-reads (e.g. reads >150bp). We have added new filtering steps to increase the accuracy of the genotyping, and new parameters that allow the user to control both the minimum and maximum number of reads, and the minimum number of strains to genotype a TE insertion. We also showed for the first time that T-lex3 provides accurate TE calls in a plant genome.
Transposable elements (TEs) are parasitic DNA sequences that threaten genome integrity by replicative transposition in host gonads. The Piwi-interacting RNAs (piRNAs) pathway is assumed to maintain Drosophila genome homeostasis by downregulating transcriptional and post-transcriptional TE expression in the ovary. However, the bursts of transposition that are expected to follow transposome derepression after piRNA pathway impairment have not yet been reported. Here, we show, at a genome-wide level, that piRNA loss in the ovarian somatic cells boosts several families of the endogenous retroviral subclass of TEs, at various steps of their replication cycle, from somatic transcription to germinal genome invasion. For some of these TEs, the derepression caused by the loss of piRNAs is backed up by another small RNA pathway (siRNAs) operating in somatic tissues at the post transcriptional level. Derepressed transposition during 70 successive generations of piRNA loss exponentially increases the genomic copy number by up to 10-fold.
Gene copy-number variations are widespread in natural populations, but investigating their phenotypic consequences requires contemporary duplications under selection. Such duplications have been found at the ace-1 locus (encoding the organophosphate and carbamate insecticides’ target) in the mosquito Anopheles gambiae (the major malaria vector); recent studies have revealed their intriguing complexity, consistent with the involvement of various numbers and types (susceptible or resistant to insecticide) of copies. We used an integrative approach, from genome to phenotype level, to investigate the influence of duplication architecture and gene-dosage on mosquito fitness. We found that both heterogeneous (i.e., one susceptible and one resistant ace-1 copy) and homogeneous (i.e., identical resistant copies) duplications segregated in field populations. The number of copies in homogeneous duplications was variable and positively correlated with acetylcholinesterase activity and resistance level. Determining the genomic structure of the duplicated region revealed that, in both types of duplication, ace-1 and 11 other genes formed tandem 203kb amplicons. We developed a diagnostic test for duplications, which showed that ace-1 was amplified in all 173 resistant mosquitoes analyzed (field-collected in several African countries), in heterogeneous or homogeneous duplications. Each type was associated with different fitness trade-offs: heterogeneous duplications conferred an intermediate phenotype (lower resistance and fitness costs), whereas homogeneous duplications tended to increase both resistance and fitness cost, in a complex manner. The type of duplication selected seemed thus to depend on the intensity and distribution of selection pressures. This versatility of trade-offs available through gene duplication highlights the importance of large mutation events in adaptation to environmental variation. This impressive adaptability could have a major impact on vector control in Africa.
DNA derived from transposable elements (TEs) constitutes large parts of the genomes of complex eukaryotes, with major impacts not only on genomic research but also on how organisms evolve and function. Although a variety of methods and tools have been developed to detect and annotate TEs, there are as yet no standard benchmarks-that is, no standard way to measure or compare their accuracy. This lack of accuracy assessment calls into question conclusions from a wide range of research that depends explicitly or implicitly on TE annotation. In the absence of standard benchmarks, toolmakers are impeded in improving their tools, annotators cannot properly assess which tools might best suit their needs, and downstream researchers cannot judge how accuracy limitations might impact their studies. We therefore propose that the TE research community create and adopt standard TE annotation benchmarks, and we call for other researchers to join the authors in making this long-overdue effort a success.