Mice have evolved a new dental plan with two additional cusps on the upper molar, while hamsters were retaining the ancestral plan. By comparing the dynamics of molar development with transcriptome time series, we found at least three early changes in mouse upper molar development. Together, they redirect spatio-temporal dynamics to ultimately form two additional cusps. The mouse lower molar has undergone much more limited phenotypic evolution. Nevertheless, its developmental trajectory evolved as much as that of the upper molar and co-evolved with it. Among the coevolving changes, some are clearly involved in the new upper molar phenotype. We found a similar level of coevolution in bat limbs. In conclusion, our study reveals how serial organ morphology has adapted through organ-specific developmental changes, as expected, but also through shared changes that have organ-specific effects on the final phenotype. This highlights the important role of developmental system drift in one organ to accommodate adaptation in another. Mice evolved a new shape in the upper, but not in the lower molar. In addition to developmental changes specific to the upper molar, this occurred through shared changes with the lower molar, the development of which has drifted to support upper molar adaptation.
Comparative genomic data can be used to look for substitutions in coding sequences that are associated with the variation of a particular phenotypic trait. A few statistical methods have been proposed to do so for phenotypes represented by discrete values. For continuous traits, no such statistical approach has been proposed, and researchers have resorted to sensible but uncharacterized criteria. Here, we investigate a phylogenetic model for coding sequences where amino acid preferences at a site are given by a continuous function of a quantitative trait. This function is inferred from the amino acids and the trait values in extant species and requires inferred point estimates of ancestral values of the trait at internal nodes. For detecting sites whose evolution is associated with this trait, we use a significance test against the hypothesis that amino acid preference does not depend on the trait. This procedure is compared to simpler strategies on simulated alignments. It displays an increased recall for low false positive rates, which is of special importance for performing whole-genome scans. This comes however at a much higher computational cost, and we suggest using a simple test to filter promising candidate sites. We then revisit a dataset of alignments for 62 species of mammals, using longevity as a phenotypic trait. We apply our method to three protein families that have previously been proposed to display sites associated with variation in lifespan in mammals. Using a graphical representation extracted from the detailed phylogenetic analysis of candidate sites, we suggest that the evidence for this in the sequence data alone is weak. The proposed method has been added to our Pelican software. It is available at and can now be used with both discrete and continuous phenotypes to search for sites associated with phenotypic variation, on data sets with thousands of alignments. ### Competing Interest Statement The authors have declared no competing interest.
It is often useful, in the field of molecular evolution, to identify the selective pressures acting on a particular site of a protein to better understand its function. This is typically done with likelihood-based approaches applied to codon sequences in a phylogenetic context. However, these approaches are computationally costly. Here we adapt a linear transformer neural network architecture, which has been shown to be able to reconstruct accurate phylogenies from sequence alignments, to identify selective pressures acting on individual amino acid sites. We design different versions of the architecture and train and test them on simulations. We compare the results of one of our best models to the state-of-the-art approach codeml and find that it outperforms it when it is applied to data that resemble its training data, but that it performs less well when applied to data that does not resemble the training data. In all cases, our approach operates at a fraction of codeml’s computational cost. These results suggest that such a neural network architecture, trained on realistic simulations, could compare very favorably to state-of-the-art approaches to characterize selection pressures acting on coding sequences. ### Competing Interest Statement The authors have declared no competing interest. Agence Nationale de la Recherche, https://ror.org/00rbzpz17, ANR-23-CE45-0027
Phylogenetic inference aims at reconstructing the tree describing the evolution of a set of sequences descending from a common ancestor. The high computational cost of state-of-the-art maximum likelihood and Bayesian inference methods limits their usability under realistic evolutionary models. Harnessing recent advances in likelihood-free inference and geometric deep learning, we introduce Phyloformer, a fast and accurate method for evolutionary distance estimation and phylogenetic reconstruction. Sampling many trees and sequences under an evolutionary model, we train the network to learn a function that enables predicting a tree from a multiple sequence alignment. On simulated data, we compare Phyloformer to FastME-a distance method-and two maximum likelihood methods: FastTree and IQTree. Under a commonly used model of protein sequence evolution and exploiting graphics processing unit (GPU) acceleration, Phyloformer outpaces all other approaches and exceeds their accuracy in the Kuhner-Felsenstein metric that accounts for both the topology and branch lengths. In terms of topological accuracy alone, Phyloformer outperforms FastME, but falls behind maximum likelihood approaches, especially as the number of sequences increases. When a model of sequence evolution that includes dependencies between sites is used, Phyloformer outperforms all other methods across all metrics on alignments with fewer than 80 sequences. On 3,801 empirical gene alignments from five different datasets, Phyloformer matches the topological accuracy of the two maximum likelihood implementations. Our results pave the way for the adoption of sophisticated realistic models for phylogenetic inference.
Interpreting transcriptomic data presents significant challenges, particularly in non-targeted approaches. While modern functional enrichment methods are well-suited for experimental designs involving two conditions, they are less applicable to data series. In this context, we developed Cluefish, a free and open-source, semi-automated R workflow designed for untargeted, comprehensive biological interpretation of transcriptomic data series. Cluefish applies over-representation analysis on pre-clustered protein-protein interaction networks, using clusters as anchors to identify smaller, more specific biological functions. Innovative features, including cluster merging and recovery of isolated genes through shared biological contexts, enable a more complete exploration of the data. We applied Cluefish to an in-house dataset with zebrafish exposed to a dose-gradient of dibutyl phthalate and to two published toxicology datasets featuring different organisms. Combined with DRomics, a tool for dose-response analysis-Cluefish identified gene clusters deregulated at low doses and linked to biological functions overlooked by the standard approach. Notably, it revealed that retinoid signaling disruption may be the most sensitive pathway affected by dibutyl phthalate during zebrafish development, potentially leading to morphological changes. The Cluefish workflow aims to provide valuable clues for biological hypothesis generation and experimental validation. It is freely available at https://github.com/ellfran-7/cluefish.
The human genome comprises tens of thousands of long non-coding RNAs (lncRNAs), whose functionality is highly debated. In the field of hepatocellular carcinoma (HCC) research, as for other cancer types, lncRNAs are increasingly reported to act as oncogenes or tumor suppressors and are put forward as useful diagnostic or prognostic biomarkers. Here, we investigate the reliability of these claims by performing a meta-analysis of the associations between HCC and lncRNAs reported in the scientific literature, and by assessing lncRNA expression patterns in two HCC patient cohorts. While nowadays up to 6% of all HCC-related publications cite lncRNAs, we show that most reported associations between HCC and lncRNAs have not been reproduced. In general, lncRNAs are less often differentially expressed between HCC tissues and controls compared to protein-coding genes. However, HCC-associated lncRNAs are frequently up-regulated in tumor samples, consistent with the fact that they are often selected based on transcriptome-wide comparative analyses. We perform a detailed examination of the 25 lncRNAs that are most frequently cited in association with HCC. For 10 out of these 25 lncRNAs, including well known lncRNAs such as MALAT1, NEAT1, H19 and XIST, we identify important conflicts between the biological roles and expression patterns previously reported for them in HCC and the expression patterns that we observe here. Finally, we observe that HCC-associated publications that cite lncRNAs are retracted three times more often than publications that cite protein-coding genes. Our results thus highlight the poor reproducibility of lncRNA-related claims in association with HCC, which is problematic in a context where new biomarkers and molecular targets for therapy are greatly needed. ### Competing Interest Statement The authors have declared no competing interest.
Quantifying oxidative stress has garnered extensive interest in evolutionary ecology and physiology since proposed as a mediator of life histories. However, while the theoretical framework of oxidative stress ecology is wellsupported by laboratory-based studies, results obtained in wild populations on oxidative damage and antioxidant biomarkers have shown inconsistent trends. We propose that red blood cell lysis could be a source of bias affecting measurements of oxidative stress biomarkers, distorting the conclusions drawn from them. Using an experimental approach consisting of enriching plasma from roe deer with lysed red blood cells, we show that the values of commonly used oxidative stress biomarkers linearly increase with the degree of haemolysis - assayed by haemoglobin concentration. This result concerns oxidized proteins (carbonyls) and lipids (TBARS), as well as enzymatic (superoxide dismutase) and non-enzymatic (trolox assay, OXY assay) antioxidant markers. Based on 707 roe deer blood samples collected in the field, we next show that the occurrence of haemolysis in plasma samples is negatively related to age. Finally, we illustrate that considering the variance explained by age-related haemolysis improves explanatory models for inter-individual variability in plasma oxidative stress biomarkers, without substantially altering the estimates of the parameters studied here. Our results raise the question of the veracity of the conclusions if the degree of haemolysis in plasma is not considered in animal models such as roe deer, for which the occurrence and severity of haemolysis vary according to individual characteristics. We recommend measuring and controlling for the degree of haemolysis be considered in future studies that investigate the causes and consequences of oxidative stress in ecophysiological studies.
Over the past decade, neural networks have been successful at making predictions from biological sequences, especially in the context of regulatory genomics. As in other fields of deep learning, tools have been devised to extract features such as sequence motifs that can explain the predictions made by a trained network. Here we intend to go beyond explainable machine learning and introduce SEISM, a selective inference procedure to test the association between these extracted features and the predicted phenotype. In particular, we discuss how training a one-layer convolutional network is formally equivalent to selecting motifs maximizing some association score. We adapt existing sampling-based selective inference procedures by quantizing this selection over an infinite set to a large but finite grid. Finally, we show that sampling under a specific choice of parameters is sufficient to characterize the composite null hypothesis typically used for selective inference—a result that goes well beyond our particular framework. We illustrate the behavior of our method in terms of calibration, power and speed and discuss its power/speed trade-off with a simpler data-split strategy. SEISM paves the way to an easier analysis of neural networks used in regulatory genomics, and to more powerful methods for genome wide association studies (GWAS).
We know that morphogenesis evolves along with final morphologies, but have little integrated understanding of morphogenesis evolution, except for old principles such as recapitulation and heterochronies. To revisit such principles, we monitored the developmental dynamics of mouse and hamster molars by combining transcriptome timeseries and morphological quantifications. The mouse upper molar evolved a new dental plan with two more cusps. They form last in mouse upper molar, recapitulating their appearance in the fossil record, but divergence in transcriptome dynamics is already visible in early stages. Three corresponding early developmental changes, including heterochronies and changes in cell proportions, combine and result in a new developmental trajectory culminating with the late addition of two cusps. The biggest surprise came from the lower molar, which was initially included as an additional control, but whose developmental trajectories evolved as much as upper molar’s. Their transcriptome dynamics markedly co-evolved, including spatio-dynamic aspects of cusp formation which are obviously involved in the new upper molar phenotype. Hence upper molar innovation relies in part on non-specific changes which impact morphogenesis in a concerted manner in both molars, but have little impact on lower molar phenotype. This counter-intuitive observation was confirmed in bat limbs. By bridging concerted transcriptomic evolution with concerted evolution of developmental mechanisms, our study introduces a principle for the evolution of organ-specific morphological innovation, with early and pan-organ developmental changes. This changes our expectations on the underlying genetic evolution, and highlights the important role of developmental drift in one organ to accommodate adaptation in another.
Transposable elements (TEs) are parasite DNA sequences that are able to move and multiply along the chromosomes of all genomes. They can be controlled by the host through the targeting of silencing epigenetic marks, which may affect the chromatin structure of neighboring sequences, including genes. In this study, we used transcriptomic and epigenomic high-throughput data produced from ovarian samples of several Drosophila melanogaster and Drosophila simulans wild-type strains, in order to finely quantify the influence of TE insertions on gene RNA levels and histone marks (H3K9me3 and H3K4me3). Our results reveal a stronger epigenetic effect of TEs on ortholog genes in D. simulans compared with D. melanogaster. At the same time, we uncover a larger contribution of TEs to gene H3K9me3 variance within genomes in D. melanogaster, which is evidenced by a stronger correlation of TE numbers around genes with the levels of this chromatin mark in D. melanogaster. Overall, this work contributes to the understanding of species-specific influence of TEs within genomes. It provides a new light on the considerable natural variability provided by TEs, which may be associated with contrasted adaptive and evolutionary potentials.
Identifying the footprints of selection in coding sequences can inform about the importance and function of individual sites. Analyses of the ratio of nonsynonymous to synonymous substitutions (dN/dS) have been widely used to pinpoint changes in the intensity of selection, but cannot distinguish them from changes in the direction of selection, that is, changes in the fitness of specific amino acids at a given position. A few methods that rely on amino-acid profiles to detect changes in directional selection have been designed, but their performances have not been well characterized. In this paper, we investigate the performance of six of these methods. We evaluate them on simulations along empirical phylogenies in which transition events have been annotated and compare their ability to detect sites that have undergone changes in the direction or intensity of selection to that of a widely used dN/dS approach, codeml's branch-site model A. We show that all methods have reduced performance in the presence of biased gene conversion but not CpG hypermutability. The best profile method, Pelican, a new implementation of Tamuri AU, Hay AJ, Goldstein RA. (2009. Identifying changes in selective constraints: host shifts in influenza. PLoS Comput Biol. 5(11):e1000564), performs as well as codeml in a range of conditions except for detecting relaxations of selection, and performs better when tree length increases, or in the presence of persistent positive selection. It is fast, enabling genome-scale searches for site-wise changes in the direction of selection associated with phenotypic changes.
The SARS-CoV-2 epidemic in France has focused a lot of attention as it has had one of the largest death tolls in Europe. It provides an opportunity to examine the effect of the lockdown and of other events on the dynamics of the epidemic. In particular, it has been suggested that municipal elections held just before lockdown was ordered may have helped spread the virus. In this manuscript we use Bayesian models of the number of deaths through time to study the epidemic in 13 regions of France. We found that the models accurately predict the number of deaths 2 to 3 weeks in advance, and recover estimates that are in agreement with recent models that rely on a different structure and different input data. In particular, the lockdown reduced the viral reproduction number by ≈ 80%. However, using a mixture model, we found that the lockdown had had different effectiveness depending on the region, and that it had been slightly more effective in decreasing the reproduction number in denser regions. The mixture model predicts that 2.08 (95% CI: 1.85-2.47) million people had been infected by May 11, and that there were 2567 (95% CI: 1781-5182) new infections on May 10. We found no evidence that the reproduction numbers differ between week-ends and week days, and no evidence that the reproduction numbers increased on the election day. Finally, we evaluated counterfactual scenarios showing that ordering the lockdown 1 to 7 days sooner would have resulted in 19% to 76% fewer deaths, but that ordering it 1 to 7 days later would have resulted in 21% to 266% more deaths. Overall, the predictions of the model indicate that holding the elections on March 15 did not have a detectable impact on the total number of deaths, unless it motivated a delay in imposing the lockdown.
While the genomic positions and the patterns of activity of enhancer elements can now be efficiently determined, predicting their target genes remains a difficult task. Although chromatin conformation capture data has revealed numerous long-range regulatory interactions between genes and enhancers, enhancer target genes are still traditionally inferred based on genomic proximity. This approach is at the basis of GREAT, a widely used tool for analyzing the functional significance of sets of cis -regulatory elements. Here, we propose a new tool, named GOntact, which infers regulatory relationships using promoter-capture Hi-C (PCHi-C) data and uses these predictions to derive Gene Ontology enrichments for sets of cis -regulatory elements. We apply GOntact on enhancer and PCHi-C data from human and mouse and show that it generates functional annotations that are coherent with the patterns of activity of enhancers. We find that there is substantial overlap between functional predictions obtained with GREAT and GOntact but that each method can yield unique testable hypotheses, reflecting the underlying differences in target gene assignment. With the increasing availability of high-resolution chromatin contact data, we believe that GOntact can provide better-informed functional predictions for cis -regulatory elements.
Interspecific hybridization is often seen as a genomic stress that may lead to new gene expression patterns and deregulation of transposable elements (TEs). The understanding of expression changes in hybrids compared with parental species is essential to disentangle their putative role in speciation processes. However, to date we ignore the detailed mechanisms involved in genomic deregulation in hybrids. We studied the ovarian transcriptome and epigenome of the Drosophila buzzatii and Drosophila koepferae species together with their F-1 hybrid females. We found a trend toward underexpression of genes and TE families in hybrids. The epigenome in hybrids was highly similar to the parental epigenomes and showed intermediate histone enrichments between parental species in most cases. Differential gene expression in hybrids was often associated only with changes in H3K4me3 enrichments, whereas differential TE family expression in hybrids may be associated with changes in H3K4me3, H3K9me3, or H3K27me3 enrichments. We identified specific genes and TE families, which their differential expression in comparison with the parental species was explained by their differential chromatin mark combination enrichment. Finally, cis-trans compensatory regulation could also contribute in some way to the hybrid deregulation. This work provides the first study of histone content in Drosophila interspecific hybrids and their effect on gene and TE expression deregulation.
The search for new biomarkers and drug targets for hepatocellular carcinoma (HCC) has spurred an interest in long non-coding RNAs (lncRNAs), often proposed as oncogenes or tumor suppressors. Furthermore, lncRNA expression patterns can bring insights into the global de-regulation of cellular machineries in tumors. Here, we examine lncRNAs in a large HCC cohort, comprising RNA-seq data from paired tumor and adjacent tissue biopsies from 114 patients. We find that numerous lncRNAs are differentially expressed between tumors and adjacent tissues and between tumor progression stages. Although we find strong differential expression for most lncRNAs previously associated with HCC, the expression patterns of several prominent HCC-associated lncRNAs disagree with their previously proposed roles. We examine the genomic characteristics of HCC-expressed lncRNAs and reveal an enrichment for repetitive elements among the lncRNAs with the strongest expression increases in advanced-stage tumors. This enrichment is particularly striking for lncRNAs that overlap with satellite repeats, a major component of centromeres. Consistently, we find increased non-coding RNA transcription from centromeres in tumors, in the majority of patients, suggesting that aberrant centromere activation takes place in HCC.
Transposable elements (TEs) are the main components of genomes. However, due to their repetitive nature, they are very difficult to study using data obtained with short-read sequencing technologies. Here, we describe an efficient pipeline to accurately recover TE insertion (TEI) sites and sequences from long reads obtained by Oxford Nanopore Technology (ONT) sequencing. With this pipeline, we could precisely describe the landscapes of the most recent TEIs in wild-type strains of Drosophila melanogaster and Drosophila simulans. Their comparison suggests that this subset of TE sequences is more similar than previously thought in these two species. The chromosome assemblies obtained using this pipeline also allowed recovering piRNA cluster sequences, which was impossible using short-read sequencing. Finally, we used our pipeline to analyze ONT sequencing data from a D. melanogaster unstable line in which LTR transposition was derepressed for 73 successive generations. We could rely on single reads to identify new insertions with intact target site duplications. Moreover, the detailed analysis of TEIs in the wild-type strains and the unstable line did not support the trap model claiming that piRNA clusters are hotspots of TE insertions.
ABSTRACTSerial appendages are similar organs found at different places in the body, such as fore/hindlimbs or different teeth. They are bound to develop with the same pleiotropic genes, apart from identity genes. These identity genes have logically been implicated in cases where a single appendage evolved a drastically new shape while the other retained an ancestral shape, by enabling developmental changesspecificallyin one organ. Here, we showed that independent evolution involved developmental changes happeningin bothorgans, in two well characterized model systems.Mouse upper molars evolved a new dental plan with two more cusps on the lingual side, while the lower molar kept a much more ancestral morphology, as did the molars of hamster, our control species. We obtained quantitative timelines of cusp formation and corresponding transcriptomic timeseries in the 4 molars. We found that a molecular and morphogenetic identity of lower and upper molars predated the mouse and hamster divergence and likely facilitated the independent evolution of molar’s lingual side in the mouse lineage. We found 3 morphogenetic changes which could combine to cause the supplementary cusps in the upper molar and a candidate gene,Bmper. Unexpectedly given its milder morphological divergence, we observed extensive changes in mouse lower molar development. Its transcriptomic profiles diverged as much as, and co-evolved extensively with, those of the upper molar. Consistent with the transcriptomic quantifications, two out of the three morphogenetic changes also impacted lower molar development.Moving to limbs, we show the drastic evolution of the bat wing also involved gene expression co-evolution and a combination of specific and pleiotropic changes. Independent morphological innovation in one organ therefore involves concerted developmental evolution of the other organ. This is facilitated by evolutionary flexibility of its development, a phenomenon known as Developmental System Drift.AUTHOR SUMMARYSerial organs, such as the different wings of an insect or the different limbs or teeth of a vertebrate, can develop into drastically different shapes due to the position-specific expression of so-called “identity” genes. Often during evolution, one organ evolves a new shape while another retains a conserved shape. It was thought that identity genes were responsible for these cases of independent evolution, by enabling developmental changes specifically in one organ. Here, we showed that developmental changes evolvedin bothorgans to enable the independent evolution of the upper molar in mice and the wing in bats. In the organ with the new shape, several developmental changes combine. In the organ with the conserved shape, part of these developmental changes are seen as well. This modifies the development but is not sufficient to drastically change the phenotype, a phenomenon known as “Developmental System Drift”, DSD. Thus, the independent evolution of one organ relies on concerted molecular changes, which will contribute to adaptation in one organ and be no more than DSD in another organ. This concerted evolution could apply more generally to very different body parts and explain previous observations on gene expression evolution.
Serial appendages are similar organs found at different places in the body, such as fore/hindlimbs or different teeth. They are bound to develop with the same pleiotropic genes, apart from identity genes. These identity genes have logically been implicated in cases where a single appendage evolved a drastically new shape while the other retained an ancestral shape, by enabling developmental changes specifically in one organ. Here, we showed that independent evolution involved developmental changes happening in both organs, in two well characterized model systems.Mouse upper molars evolved a new dental plan with two more cusps on the lingual side, while the lower molar kept a much more ancestral morphology, as did the molars of hamster, our control species. We obtained quantitative timelines of cusp formation and corresponding transcriptomic timeseries in the 4 molars. We found that a molecular and morphogenetic identity of lower and upper molars predated the mouse and hamster divergence and likely facilitated the independent evolution of molar’s lingual side in the mouse lineage. We found 3 morphogenetic changes which could combine to cause the supplementary cusps in the upper molar and a candidate gene, Bmper . Unexpectedly given its milder morphological divergence, we observed extensive changes in mouse lower molar development. Its transcriptomic profiles diverged as much as, and co-evolved extensively with, those of the upper molar. Consistent with the transcriptomic quantifications, two out of the three morphogenetic changes also impacted lower molar development.Moving to limbs, we show the drastic evolution of the bat wing also involved gene expression co-evolution and a combination of specific and pleiotropic changes. Independent morphological innovation in one organ therefore involves concerted developmental evolution of the other organ. This is facilitated by evolutionary flexibility of its development, a phenomenon known as Developmental System Drift.AUTHOR SUMMARY Serial organs, such as the different wings of an insect or the different limbs or teeth of a vertebrate, can develop into drastically different shapes due to the position-specific expression of so-called “identity” genes. Often during evolution, one organ evolves a new shape while another retains a conserved shape. It was thought that identity genes were responsible for these cases of independent evolution, by enabling developmental changes specifically in one organ. Here, we showed that developmental changes evolved in both organs to enable the independent evolution of the upper molar in mice and the wing in bats. In the organ with the new shape, several developmental changes combine. In the organ with the conserved shape, part of these developmental changes are seen as well. This modifies the development but is not sufficient to drastically change the phenotype, a phenomenon known as “Developmental System Drift”, DSD. Thus, the independent evolution of one organ relies on concerted molecular changes, which will contribute to adaptation in one organ and be no more than DSD in another organ. This concerted evolution could apply more generally to very different body parts and explain previous observations on gene expression evolution.### Competing Interest StatementThe authors have declared no competing interest.
In evolutionary genomics, researchers have taken an interest in identifying substitutions that subtend convergent phenotypic adaptations. This is a difficult question that requires distinguishing foreground convergent substitutions that are involved in the convergent phenotype from background convergent substitutions. Those may be linked to other adaptations, may be neutral or may be the consequence of mutational biases. Furthermore, there is no generally accepted definition of convergent substitutions. Various methods that use different definitions have been proposed in the literature, resulting in different sets of candidate foreground convergent substitutions. In this article, we first describe the processes that can generate foreground convergent substitutions in coding sequences, separating adaptive from non-adaptive processes. Second, we review methods that have been proposed to detect foreground convergent substitutions in coding sequences and expose the assumptions that underlie them. Finally, we examine their power on simulations of convergent changes-including in the presence of a change in the efficacy of selection-and on empirical alignments. This article is part of the theme issue 'Convergent evolution in the genomics era: new insights and directions'.
Stefano Nolfi合作论文数Institute of Cognitive Sciences and Technologies, National Research Council2