
Dietary adaptations are fundamental drivers of evolutionary innovations. Still, the morphological and molecular mechanisms underlying these transitions remain poorly understood. True bugs (Heteroptera), one of the most diverse clades of hemimetabolous insects, have evolved a vast array of trophic niches. They exhibit remarkable structural and functional diversity in their salivary glands, providing an ideal system for investigating how dietary variation shapes evolutionary trajectories. Using morphometric, proteo-transcriptomic, and comparative phylogenetic approaches, we examined the evolutionary patterns of salivary glands across 220 hemipteran species, encompassing all seven infraorders and six dietary categories of Heteroptera. Our results reveal strong correlations between dietary shifts, salivary gland size and allometric scaling. Ancestral state estimation analysis demonstrates multiple evolutionary transitions in salivary gland morphology, highlighting structural adaptations associated with different diets. Furthermore, we found that distantly related species with similar diets share dominant salivary protein groups, implying convergent molecular solutions to shared feeding challenges. Notably, we identify venom protein family 2 as a key adaptation. This family originated through lineage-specific co-option of an ancestral protein into heteropteran salivary systems, followed by extensive expansion that facilitated dietary diversification. These findings provide novel insights into the morphological, physiological, and molecular adaptations of salivary glands, emphasizing the dynamic interplay between feeding ecology and macroevolution.
The eukaryotic genome has been described as a collection of different phylogenetic histories. For most phylogenomic analyses the primary goal is to identify the species tree, the singular history that underlies and shapes the "gene trees" of individual loci. Discordance among gene trees and with the species tree is expected due to deep coalescence/lineage sorting, while also resulting from various technical causes (e.g., long branch attraction, pseudo-orthology), or, of greater interest, by introgression and horizontal transfer. Where do competing phylogenetic signals reside in gene tree topology space-that part of tree space occupied by the gene trees reconstructed for a particular dataset of taxa and genetic loci? We explored this question in the small (~30 species) leguminous plant genus, Glycine, which has extensive genomic resources due to the inclusion of cultivated soybean (G. max). Glycine genomes are highly duplicated due to relatively recent (~10 million years) ancestral polyploidy and have extensive nuclear-cytoplasmic discordance. We explored Glycine gene tree topology space using a set of 2389 nuclear genes and 61 representative accessions selected from a 570-taxon x 100 gene concatenation supergene tree, reconstructing gene trees for all nuclear loci and from complete plastid genomes and partial mitochondrial genomes. Species trees (ASTRAL) and maximum likelihood (ML) concatenation trees were congruent with one another but were discordant with organellar genome trees, which were incongruent with one another. Individual loci all had unique topologies for the 61-taxon dataset and for a reduced dataset of 27 taxa. No locus tracked either the species tree or the plastome topology in the resulting "flat" gene tree topology space of either dataset, nor did clustering identify any regional differentiation of gene tree topology space populated by loci with similar topologies. Only when the dataset was reduced to six Glycine species, chosen because they have complete genome sequences, and an outgroup was a topological landscape produced in which most loci tracked the species tree topology, with secondary peaks that included, most prominently, the discordant plastome topology. There was no evidence of pseudo-orthology in this landscape, and synteny-based assessment of thousands of loci across these six genomes identified few candidate pseudo-orthologs. Thus, while it is true that the Glycine genome is indeed a collection of different historical signals, those signals are complex and exist at the level of clades within trees rather than as entire gene trees. Although phylogenomic methods can reconstruct the species tree from signals scattered among many loci, even loci with very low resolution, other biologically relevant signals are much more difficult to localize without an explicit starting hypothesis.
Bayesian tip-dating under fossilized birth-death process has become an increasingly popular approach to reconstruct time-scaled phylogenies. This method implements the ages of fossils included in the analysis to calibrate clock models used to estimate divergence times. The use of morphological data and morphological clock models are fundamental in the reconstruction of the positions of fossil taxa and in the timing of their evolution in Bayesian tip-dating, but how morphological clocks perform with calibration types other than the ages from fossil tips are currently unknown. We tested the impact of varying fossil taxon sampling approaches and clock models on divergence time estimations under the morphological clock using a newly constructed matrix of ruminants. Our results indicate that increased fossil taxon sampling had only a minor effect on morphological divergence times; choice of implementing an independent or autocorrelated rate clock model revealed generally similar age estimates. Comparison among tip-dated and fossil calibrated node-dated analyses reveal both approaches estimated node ages generally consistent with previous age estimates from the molecular clock. However, increased sampling did improve precision of age estimates and prevented root age artefacts like deep-root attraction, the tendency for the root age in clock analyses to be much older than expected from the fossil record. Also, the inclusion of more fossil taxa nesting within nodes dominated by extant taxa recovered more realistic ages for these nodes. With the analysis of the largest morphological dataset of fossil and extant ruminants, the first to assess the phylogenetics of the primitive "traguline" and derived crown pecoran ruminants together, we confirm the monophyly of the total-groups of each living family, the inclusion of the blastomerycines in Moschidae and the separation of Palaeomerycidae and Dromomerycidae in addition to clarifying the positions of many fossil ruminant taxa. We discuss the most notable issues within ruminant systematics including the origin of the group, the position of the most controversial fossil taxa (i.e. Archaeomeryx, Palaeomerycidae, Dromomerycidae) and the origin of their cranial appendages. Our analyses highlight the importance of using a genomic backbone in morphological/total evidence analysis of Ruminantia because of the conflicting signals between different datatypes and from within datatypes; we further extend this to analyses of other groups with similar conflicts among data.
Abstract The Y chromosome has been omitted from almost all phylogenomic analyses to date due to its complex structure, which makes accurate assembly and alignment across species challenging. Yet the Y chromosome is, in theory, an optimal phylogenetic marker. Y-chromosomal genes have a smaller effective population size than autosomal genes, reducing the likelihood of incomplete lineage sorting. Likewise, heterogametic hybrid offspring of different species are generally sterile, creating a barrier to Y-chromosomal gene flow. To overcome difficulties in using the Y chromosome, we developed a novel approach to identify orthologous Y-linked sequences in placental mammals, enabling us to generate a 63-species Y-chromosome alignment. We aligned ~80 kilobases of predominantly noncoding sequence from conserved genes that are broadly expressed regulators of protein synthesis and spermatogenesis—the X-degenerate genes. Our phylogenetic reconstructions demonstrate that noncoding X-degenerate gene sequences recapitulate the same phylogeny derived from genome-wide noncoding, neutrally evolved sequences across the biparentally inherited autosomes and the X chromosome. We find strong support for the superordinal clades Euarchonta, Scrotifera, Fereuungulata, and Zooamata—groupings that have been historically considered controversial. Our results demonstrate that not only can the Y chromosome be aligned across species divergences spanning more than 100 million years, but it also performs robustly in a comparative phylogenetic context. We also present evidence that interchromosomal gene conversion between the non-recombining ZFX and ZFY genes has independently occurred in multiple lineages. ZFY has a proposed function as a meiotic executioner, and interchromosomal gene conversion may serve as a compensatory mechanism to prevent genetic decay of this essential gene.
Abstract Change in the haploid chromosome number commonly generates reproductive isolation between diverging species. If changes to the haploid chromosome number (karyotype) most often generate new species, variation in the rate of chromosome number evolution is expected to predict variation in diversification rates. While this correlation has been supported in some plants, we know less about how the mode and tempo of karyotype change evolve. Methods to address the evolution of diversification and chromosome number transition rates are computationally expensive and analyses are typically restricted to small clades or avoided all together. We identify and describe variation in the mode and tempo of karyotype evolution (via dysploidy and polyploidy) in the Polypodiales—a species and karyotype-rich order of ferns—by extending the Chromosome Number and Hidden State-dependent Speciation and Extinction model (ChromoHiSSE) to include whole genome duplication. Using the extended ChromoHiSSE model we estimate rates of karyotype evolution across 962 leptosporangiate ferns of the Polypodiales. We recover two hidden modes of chromosome number evolution between which the rates of karyotype evolution differ by more than an order of magnitude. Our rate estimates and the stochastic mapping of these modes across fern evolution suggests lineages with high karyotype lability are less likely to persist in the long-term. These results reinforce the theory that modern fern diversity is shaped substantially by polyploid speciation but challenge the expectation that diversification rates are enhanced by karyotype-driven reproductive isolation.
Recent advances in systematics and evolutionary biology have revealed extraordinary complexity in phylogenetic diversification, challenged our traditional views of the Tree of Life, and changed our understanding of the nature of phylogenetic entities themselves. These empirical advances call for an updated conceptual basis for phylogenetics. Phylogeny is best conceptualized as a complex multidimensional system - encompassing spatial, temporal, and hierarchical dimensions. The hierarchical dimension confers important but poorly studied influences - such as emergence, constraint, non-linearity, and non-extrapolationism - on phylogenetic possibilities. Spatially and temporally, phylogenetic diversification patterns should be viewed as resulting from multimodal phenomena - i.e., as a consequence of both vertical and horizontal modes of evolution. Further, phylogeny should not be viewed as a single absolute history - it comprises interacting, multi-layered lineage histories (i.e., genomes, cells, organisms, species). Causal interactions and feedback dynamics amongst these dimensions represent a classic biocomplexity system. Therefore, we suggest that this is an opportune time for development of a phylogenetic systems theory. Because our view of phylogeny comprises a complex interacting system of multiple lineage histories, we also argue that a process ontology provides the logical foundation for contemporary phylogenetics. A process ontology framework also reconciles recent observations that the boundaries of biological entities (including the boundaries of genomes, cells, organisms, and species lineages) are fuzzy and open to varying degrees of exchange at various points in their histories. While these fuzzy boundaries accommodate the biocomplexity and multimodal evolutionary processes described above, they also pose challenges for our notions of phylogenetic entities themselves and their biological individuality - leading to exciting new areas of study in systematics, theoretical biology and philosophy of biology.
Multivariate phylogenetic comparative methods for modelling high-dimensional traits such as 3D shapes or gene expression profiles have been recently developed. However, these approaches are impractical and almost impossible to use when the number of traits exceeds a few thousands, as they become computationally prohibitive. We overcome these limitations by proposing a new maximum likelihood approach based on the Empirical Bayes framework. This approach takes into account the information of the complete covariances (among species and traits) to infer parameters and compare trait evolution models for high-dimensional datasets. Through simulations, we demonstrate that the proposed approach can accurately estimate parameters of various trait evolution models, even when the number of traits is ten times larger than the number of lineages; it requires less memory and is at least 10 times faster than currently available approaches. This efficient framework enabled us to extend the high-dimensional multivariate phylogenetic comparative toolkit by including an Ornstein-Uhlenbeck process with multiple optima to study adaptation to various selective regimes. Applying our approach to the evolution of jaw morphology in relation to dietary adaptation in mammals, we demonstrate morphological convergence in carnivorous and herbivorous lineages. The proposed Empirical Bayes approach, implemented in the R package mvMORPH, enables phylogenetic comparative methods to efficiently handle high-dimensional datasets and complex models of trait evolution.
The pace of evolutionary diversification varies among groups, between environments, and over geologic timescales. Radiations may be linked to phenotypic innovations and ecological opportunity but for highly diverse groups the factors contributing to accelerated diversification are often less clear. We examine the dynamics and drivers of diversification in gobies, an exceptionally species-rich group with a cosmopolitan distribution in aquatic habitats that exhibits extensive ecological variation coupled with low phenotypic diversity. We establish their evolutionary timescale using a phylogenomic hypothesis of relationships calibrated with multiple fossils and identify a pronounced increase in diversification rate in the clade containing Gobiidae and Oxudercidae. This shift is correlated with both phenotypic and ecological traits; elevated diversification is found among marine lineages and those on reefs as well as in those possessing a fused ventral pelvic disc, a key innovation that facilitates stability on benthic substrates. Our analyses additionally reveal diversification fluctuations at the Eocene-Oligocene transition and in the mid-Miocene, periods of extinction and recovery in global marine habitats. Achieving even near-complete taxon sampling is unlikely in hyperdiverse groups such as gobies, but we show that evolutionary dynamics may still be reliably inferred with less than complete taxon sampling by validating our diversification shift patterns using completely sampled, stochastically resolved trees and confirming the robustness of our rate estimates across congruence classes.
Biogeography is intrinsically linked to evolution as a process and to systematics as a practice. Phylogenetic biogeography, in particular, studies the distribution of life in space over time through the lens of common ancestry. Over the past century, new biological and geological discoveries, theoretical frameworks, and methodological techniques revolutionized how we understand why species have come to live where they do. This perspective piece orients readers to major advances, changes, and conflicts from the phylogenetic biogeography literature, much of which was centered around articles published in this very journal. As part of our survey, we also highlight areas that were historically active, remain biologically significant, and deserve renewed attention.
One of the most transformative advances in molecular systematics over the past several decades has been our understanding of how population processes like incomplete lineage sorting (ILS) give rise to phylogenetic discordance. Fossil clades, which are known almost exclusively from morphology, have benefited little from these developments. An emerging body of work has shown how processes related to those that cause discordance at the molecular level can also shape patterns at the phenotypic level. Conflicting signal can arise when phenotypic polymorphisms maintained in an ancestral taxon sort among descendant taxa or persist beyond speciation. Here, I explore how ancestral polymorphisms contribute to the landscape of phylogenetic discordance in Pleistocene Homo and late-Cretaceous Micraster. I found rampant discordance stemming from the sorting and persistence of ancestral polymorphisms in both clades. This suggests that mechanisms related to those that drive gene-tree discordance may also shape morphological patterns. Simulations show that explicitly modelling the processes that give rise to this discordance can dramatically improve reconstructions. Further studying how biological processes like ILS generate discordance in phenotypic characters and incorporating them into phylogenetic models can improve tree reconstruction and generate novel biological insights in fossil clades.
Plants with amphitropical distributions have closely related populations in both Northern and Southern Hemispheres, but are absent from the intervening tropics. They provide a unique opportunity to study the constraints shaping the distribution of temperate lineages through time. Using grasses from the ecologically diverse supertribe Melicodae, an emerging study system with species distributed throughout the temperate regions, we test the hypothesis that geography and/or environmental niche constrain which lineages successfully cross the tropics to establish in the opposite hemisphere. Biogeographic and evolutionary modelling was conducted on well resolved plastid and nuclear phylogenies constructed from whole-genome sequencing of 178 accessions of 103 Melicodae species. Results show that species from cold regions are much less likely to successfully cross the tropics, with successful lineages all sharing warmer niches that evolved prior to their establishment in the opposite hemisphere. Evidence suggests that this result is explained both by the greater distances that high-latitude, cold-origin lineages must disperse to cross the tropics, and inherent limitations associated with colder thermal niches. In particular, our results suggest that traits allowing species to cope with cold winters, rather than an inability to cope with warm summers, limit their ability to establish in the opposite hemisphere, hinting at important trade-offs between cold-tolerance and biogeographic potential. These results provide insight into the drivers of the distribution and diversity of plants, and the challenges facing cold-origin lineages in a rapidly warming world. If cold-origin species occupy a smaller proportion of their potential range, and are unlikely to establish in new areas with suitable climates, their ability to track preferred habitat as climates warm may be worse than currently expected.
The grass subtribe Loliinae has great ecological and economic importance, as it includes community-dominant species of mountain grasslands and the most extensively cultivated pasture, fodder and turf grasses (fescues, ryegrasses). Resolving the phylogeny of recently evolved Loliinae lineages has proven challenging due to frequent introgressions and polyploidizations that occurred throughout their history. Here we present the first target-capture phylogeny of Loliinae using 270 orthologous single-copy nuclear coding loci for a large sample of 132 representative taxa, covering all its 29 evolutionary lineages. Additionally, we assembled plastome sequences to complement the inferred hybrid speciation history of the Loliinae. Concatenated maximum likelihood and multispecies coalescent trees of ortho-homeolog single-copy genes showed well-supported relationships for major lineages, which were generally consistent across analyses and genomes, and with previous taxonomic and phylogenetic findings. However, they also revealed high levels of both nuclear and cyto-nuclear discordances estimated to be caused by hybridization and incomplete lineage sorting. We complemented this with a phylogenetic network analysis with representative samples from the main clades to infer reticulation events in the evolution of these grasses. Furthermore, we performed gene tree - species tree reconciliation methods using gene duplication and loss models and multi-labeled trees for polyploidy analysis to estimate the proportion of duplicated genes and the nature of polyploidy of the major Loliinae lineages. These analyses agreed that the Fine-Leaved (FL) Loliinae clade could have evolved from hybridization between more ancestral Broad-Leaved (BL) Loliinae lineages and that both groups underwent ancestral and recent hybridizations. The time-calibrated phylogeny of for the main Loliinae clades supports an early Miocene origin for Loliinae and Mid-Late Miocene splits for its main BL and FL lineages, while current species-rich groups radiated in the Late Miocene. Hybridization tests of nuclear data and topological incongruence assays between the nuclear single-copy genes and plastome-based trees using various approaches and different sampling subsets confirmed the rampant hybridization experienced by Loliinae at deep and shallow nodes. However, hybridization rates differed from lineage to lineage within the major clades and were not correlated with time or ploidy level, but rather depended on their different propensities to hybridize with species within and/or outside their own clade. Our analyses detected high hybridization rates in four broad-leaved (Subulatae-Hawaiian, Tropical-South African, Mexico-Central-South American (MCSA I-II) and Leucopoa p. p.) and five fine-leaved Loliinae lineages (American II, Aulaxyper, Afroalpine, American-Neozeylandic, Australia-Tasmania) containing rogue species that probably originated from trans-clade crosses and are more likely to hybridize greatly. In contrast, they recovered low hybridization rates in four broad-leaved (Schedonorus-Lolium, Subbulbosae, Drymanthele-Pseudoscariosae-Lojaconoa, Leucopoa p. p.) and six fine-leaved lineages (Festuca, Psilurus-Vulpia, American I, Exaratae-Loretia, American-Vulpia-Pampas, Eskia), with species derived from single ancestors that hybridize only with close congeners. Inferences from gene duplications and allopolyploidizations, along with inheritance probabilities from the phylogenetic network, point to the BL Loliinae lineages as ancient hybrids or paleo-allopolyploids, while the FL lineages, especially those of the core FL clade, correspond to more recent meso- or neo-allopolyploids. [allopolyploidization, lineage-specific hybridization rates, nuclear and cyto-nuclear discordances, Festuca and related genera, phylogenomics, nuclear single-copy genes, plastome].
Phenotypes serve as the interface between organisms and their environments and are thus pivotal for comprehensive biological understanding. However, comparative analyses of species' phenotypes must account for the non-independence of characters imposed by the branching pattern of macroevolution. Methods to account for this phylogenetic non-independence have historically been conceived of as their own subfield: phylogenetic comparative methods (PCMs). In this conceptualization, a researcher takes a pre-existing phylogeny and uses it to correct for the non-independence of the character data that they wish to analyze. Here, we argue that parallel developments in scientific philosophy, data availability, computational capacity, and model development have enabled a new paradigm of character evolution, where patterns of character evolution are seen as intrinsically related to the diversification process, and thus should be inferred jointly with the tree rather than reconstructed on an existing phylogeny in a two-step process. In the context of this paradigm shift, we review historical milestones in studies of character macroevolution and discuss major recent conceptual and methodological advances, with an emphasis on the opportunities and insights provided by the joint-inference perspective. We include primers on current topics in character evolution where joint inference is particularly effective, including: (1) state-dependent speciation and extinction models and the importance of cladogenetic change; (2) jointly modeling discrete and continuous characters; (3) accounting for hidden process variation in character evolution; (4) joint inference in divergence-time estimation; (5) joint inference of phylogeny and ancestral states; and (6) joint inference of alignment and phylogeny. The article concludes with a reflection on the future trajectory of these methods, emphasizing the interconnectedness of character evolution with broader processes in biology.
The symbiosis between clownfishes (or anemonefishes) and their host sea anemones ranks among the most recognizable animal interactions on the planet. Found on coral reef habitats across the Indian and Pacific Oceans, 28 recognized species of clownfishes adaptively radiated from a common ancestor to live obligately with only 10 nominal species of host sea anemones. Are the host sea anemones truly less diverse than clownfishes? Did the symbiosis with clownfishes trigger a reciprocal co-evolutionary response to the mutualism? To address these questions, we combined fine- and broad-scale biogeographic sampling with multiple independent genomic datasets for the bubble-tip sea anemone, Entacmaea quadricolor-the most common clownfish host anemone throughout the Indo-West Pacific. Fine-scale sampling and restriction site associated DNA sequencing (RADseq) throughout the Japanese Archipelago revealed three highly divergent cryptic species: two of which co-occur throughout the Ryukyu Islands and can be differentiated by the clownfish species they host. Remarkably, broader biogeographic sampling and bait-capture sequencing reveals that this pattern is not simply the result of local ecological processes unique to Japan, but part of a deeper evolutionary signal where some species of E. quadricolor serve as host to the generalist clownfish species Amphiprion clarkii and others serve as host to the specialist clownfish A. frenatus. In total, we delimit six cryptic species in E. quadricolor that have diversified within the last five million years. The rapid speciation of E. quadricolor combined with functional ecological and phenotypic differentiation supports the hypothesis that this diversification is an evolutionary response to mutualism with clownfishes. Clownfishes are not merely settling in locally available hosts but recruiting to specialized host lineages with which they have co-evolved. These findings have important implications for understanding how the clownfish-sea anemone symbiosis has evolved and will shape future research agendas on this iconic model system.
Accurate phylogenetic inference requires models that account for heterogeneity in molecular evolution. Mitochondrial protein-coding genes, which encode membrane-bound proteins composed of multiple transmembrane α-helices, exhibit considerable compositional and functional variation across structural regions, variation that is often overlooked in standard partitioning strategies. Here, we introduce TRAMPO (TRAnsMembrane Protein Order), a novel pipeline that incorporates predicted secondary structural features (i.e. matrix-facing, transmembrane, and intermembrane-facing domains) into phylogenetic partitioning schemes. We applied TRAMPO to seven mitochondrial datasets spanning crustaceans, hexapods, and vertebrates, and evaluated eight partitioning strategies based on combinations of codon position, strand, and secondary structure. Transmembrane helices, especially at second codon positions, showed pronounced thymine enrichment and hydrophobic amino-acid composition, reflecting domain-specific evolutionary constraints. To assess whether these structural patterns influence phylogenetic reconstruction, we performed maximum likelihood analyses under standard and Lie Markov models, General Heterogeneous evolution On a Single Topology, and profile mixture models. We also evaluated different models of among-site rate variation (including the proportion of invariant sites, gamma distributions, and FreeRates, which approximates rate heterogeneity using flexible discrete rate categories) to examine their interaction with partitioning strategies and overall model performance. Incorporating structural information into partitioning schemes consistently improved model fit and reduced apparent heterogeneity, as reflected in lower AIC values and more compositionally homogeneous partitions. These improvements translated into more consistent and topologically congruent phylogenetic trees across most datasets, while also reducing computational time. Notably, second codon positions within transmembrane helices were consistently retained as distinct partitions during model optimization, even in Mammals and Vertebrates, where secondary structure contributed little to overall model performance, underscoring their strong and conserved evolutionary signal. Surveys of tree space using quartet distances further supported these findings, with structurally informed models yielding more tightly clustered and internally consistent tree topologies. The benefits of structural partitioning were most pronounced in lineages of intermediate evolutionary depth and declined in ancient vertebrate and mammalian clades, where substitutional saturation accumulates with evolutionary time and strand asymmetry tends to emerge more frequently. In some cases, models with the lowest AIC did not yield the most congruent topologies, underscoring the limitations of information criteria when comparing models of different complexity. Overall, our findings demonstrate that secondary structural features, particularly the repetitive architecture of transmembrane helices, harbour meaningful phylogenetic signal. Incorporating this information into partitioning schemes improves tree reconstruction and mitigates underlying heterogeneity. TRAMPO provides a scalable, open-source tool to implement this approach in mitochondrial phylogenetics.
The abundant discordance between evolutionary relationships across the genome has rekindled interest in methods for comparing and averaging trees on a shared leaf set. However, compared to tree topology, where much progress has been made, handling branch lengths has been more challenging. Species tree branch lengths can be measured in various units, often different from gene trees. Moreover, rates of evolution change across the genome, the species tree, and specific branches of gene trees. These factors compound the stochasticity of coalescence times and estimation noise, making branch lengths highly heterogeneous across the genome. For many downstream applications in phylogenomic analyses, branch lengths are as important as the topology, and yet, existing tools to compare and combine weighted trees are limited. In this paper, we address the question of matching one tree to another, accounting for their branch lengths. We define a series of computational problems called Topology-Constrained Metric Matching (TCMM) that seek to transform the branch lengths of a query tree based on a reference tree. We show that TCMM problems can be solved efficiently using a linear algebraic formulation coupled with dynamic programming preprocessing. While many applications can be imagined for this framework, we explore two applications in this paper: embedding leaves of gene trees in Euclidean space to find outliers potentially indicative of estimation errors, and summarizing gene tree branch lengths onto the species tree. In these applications, our method, when paired with existing methods, increases their accuracy at limited computational expense.
The classification of organisms into species is fundamental to the study of life. Contrary to popular belief, simple and quantitative standards for species delineation are often lacking, and debates about species boundaries create obstacles for conservation biology, agriculture, legislation, and education. We chose butterflies as a model system to address this key biological question. We sequenced and analyzed transcriptomes of 186 butterfly specimens representing 25 pairs of species, representing close but clearly distinct species, conspecific populations, and taxa with debated relationships. We found that species are robustly separated from conspecific populations by the combination of two measures computed on Z-linked genes: a fixation index, which detects genetic gaps between species, and the extent of gene flow, which quantifies reproductive isolation. Applying these criteria suggest that all nine butterfly pairs with debated relationships are distinct species, not populations or subspecies. Furthermore, we found that elevated divergence and positive selection in proteins involved in DNA interaction, circadian clock, pheromone sensing, development, and immune response recurrently correlate with speciation. A significant fraction of these divergent proteins are encoded by the Z chromosome, which appears to be more resistant to introgression than autosomes. Taken together, our findings point to potential common speciation mechanisms in butterflies, provide additional support for the important role of the Z chromosome in speciation, and suggest quantitative criteria for species delimitation, which is vital for the exploration of biodiversity.
Inference of interspecific gene flow using genomic data is important to reliable reconstruction of species phylogenies and to our understanding of the speciation process. Gene flow is harder to detect if it involves sister lineages than nonsisters; for example, most heuristic methods based on data summaries are unable to infer gene flow between sisters. Likelihood-based methods can identify introgression between sisters but the test exhibits several nonstandard features, including boundary problems, indeterminate parameters, and multiple routes from the alternative to the null hypotheses. In the Bayesian test, those irregularities pose challenges to the use of the Savage-Dickey (S-D) density ratio to calculate the Bayes factor. Here we develop a theory for applying the S-D approach under nonstandard conditions. We show that the Bayesian test of introgression between sister lineages has low false-positive rates and high power. We discuss issues surrounding the estimation of the rate of gene flow between sister lineages, especially at very low or very high rates, and suggest that evidence for gene flow between sisters be assessed via a Bayesian test. We find that the species split time has a major impact on the information content in the data, with more information at deeper divergence. We use a genomic dataset from Sceloporus lizards to illustrate the test of gene flow between sister lineages.
Traditional phylogenetically aware correlation methods perform well under gradual evolutionary processes. However, abrupt evolutionary shifts-or macroevolutionary jumps, characteristic of punctuated evolution-can produce extreme phylogenetically independent contrasts (PIC), leading to inflated false positives or increased false negatives in trait correlation analyses. We introduce O(D)GC (Outlier- and Distribution-Guided Correlation), a flexible workflow that identifies outliers in PICs using a distribution-free boxplot criterion and applies Spearman correlation whenever influential outliers are detected. If no outliers are detected, Pearson correlation is used-automatically for large datasets (n ≥ 30), or guided by normality testing in smaller samples. We systematically compared PIC-O(D)GC with five widely applied phylogenetic correlation methods-PIC-Pearson, PIC-MM, PGLS (phylogenetic generalized least squares), MR-PMM (multi-response phylogenetic mixed model), and Corphylo-on 322,000 simulated datasets spanning five evolutionary scenarios (two shift settings: single-trait shifts and dual-trait co-directional jumps; and three no-shift gradual evolution settings), including both fixed-depth and randomly located shifts, tested across 11 shift or noise gradients, three tree sizes (16, 128, 256 tips), and both balanced and random topologies. Overall, PIC-O(D)GC achieved error rates comparable to-or noticeably higher than-those of PIC-MM, while yielding substantially lower error rates than most alternative methods. Under no-shift conditions, it retained power similar to other methods. Analyses of three empirical datasets likewise showed that PIC-O(D)GC and PIC-MM corrected shift-induced distortions that misled conventional methods. Moreover, PIC-O(D)GC offers a conceptually simple framework and incurs markedly lower computational cost. By design, its correlation-only output provides less mechanistic detail than regression-based approaches like PGLS. However, when paired with PIC diagnostics, this outlier-guided strategy highlights evolutionary jumps, distinguishes coupled from decoupled shifts, and-via clade partitioning or tip pruning-recovers background correlations, offering biologically informative insights into how punctuated events interact with gradual trends in trait evolution.
Massively parallel algorithms leveraging graphics processing units (GPUs) have significantly accelerated inference in statistical phylogenetics, with applications in understanding pathogen evolution, population dynamics, natural selection, and evolutionary timescales using ancient genomes. Continued advancements in GPU hardware necessitate innovative algorithms to fully exploit their potential. Here, we introduce three novel algorithms that accelerate matrix multiplication operations using tensor cores on NVIDIA GPUs to calculate the observed sequence data likelihood and the gradient of the log-likelihood with respect to branch-length-specific parameters under continuous-time Markov chain models of evolution. The algorithms presented in this paper deliver 2 to 3-fold gains in performance for amino acid and codon models compared to existing GPU-based massively parallel algorithms. Notably, these performance gains are accompanied by a ~2-fold reduction in energy usage, demonstrating the potential of these algorithms to lower the carbon footprint of evolutionary computing. We make our new algorithms available to the broader phylogenetics community through the high-performance, open source library BEAGLE v4.0.0.