
Given a tree topology and an assignment of character states to its leaves, the Small Parsimony Problem (SPP) consists in assigning character states to the internal nodes in a way maximizing a certain parsimony or probabilistic criterion. In the genome rearrangement field, tree leaves are permutations of gene sets, and the problem is to infer permutations at internal nodes minimizing a rearrangement distance. Almost all genome rearrangement models lead to intractable problems for the SPP. Considering only the numerical profiles on a phylogeny, the Count package (Csűrös, 2010) can be used to predict the size of gene families on internal nodes under a parsimony or probabilistic model. Here, we present a tractable version of the SPP: given a tree topology leaf-labeled by unordered gene sets, infer gene sets at internal nodes in a way minimizing the number of gain and loss episodes on the edges of the tree, while having a single gain point for each gene (i.e. under Dollo’s law). We show that the entire solution space is covered by testing four possible cases on each internal node’s content, leading to a linear-time dynamic programming algorithm for obtaining an optimal solution. We apply our InOutParsimony software to clusters of orthologous mitochondrial protein-coding genes (MitoCOGs) in both the mitochondrial and nuclear genomes of 11 land plant species. The results are discussed considering the Endosymbiotic Gene Transfer events shaping the mitochondria and nucleus contents and compared with Count’s returned numerical profiles.
To reconstruct the gene orders of the constituent genomes in an ancestral pangenome, we propose an analysis of raccroche maximum weight matching output of contigs based on adjacency pairs from phylogenetically disparate genomes. The key idea is to use the multiple solutions to the matching optimization as a sample of the set of constituent genomes. We identify those gene-order contigs present in all the solutions as the “core” of the pangenome, and those absent in some of the solutions as the pangenome “shell”. Different cliques of mutually compatible shell contigs identified different constituent genomes. A significant proportion of shell genes in each pangenome was inherited from the set of shell genes in its ancestor pangenome. As a hint to the chromosomal structure, we performed hierarchical clustering on the combined set of contigs based on the number of solutions shared by each pair of contigs, in the search for chromosomal fragments, large clusters present in many ancestors, and some limited but clear results were obtained.
Species tree estimation from multi-copy gene family trees, including both paralogs and orthologs, is a challenging task due to the gene tree discordance caused by biological processes such as incomplete lineage sorting (ILS) and gene duplication and loss (GDL). Quartet-based species tree estimation methods, such as ASTRAL, Quartet MaxCut (QMC), and Quartet Fiduccia–Mattheyses (QFM) frameworks have gained substantial popularity for their accuracy and statistical guarantee. However, most of these methods rely on single-copy gene trees and model only ILS, which limits their applicability to large genomic datasets. ASTRAL-Pro incorporates both orthology and paralogy for species tree inference under GDL by employing a refined quartet similarity measure based on the concept of species-driven quartets (SQs). In this study, we show that these SQ-based techniques can be effectively leveraged within the QFM framework. This required substantial algorithmic re-engineering, including the development of efficient techniques for computing the initial bipartition in QFM and novel combinatorial methods for computing refined quartet scores directly from gene family trees. We extensively evaluated our method, wQFM-GDL, on simulated and real biological datasets and compared it with leading methods ASTRAL-Pro3 and SpeciesRax. wQFM-GDL outperforms all other methods in 113 out of 124 model conditions considered in this study, with performance differences becoming more pronounced as dataset size increases. In particular, for larger datasets with 200 and 500 taxa, wQFM-GDL significantly outperforms all leading methods in all 72 out of 72 model conditions and achieves, on average, nearly a 25 https://github.com/abdur-rafi/wQFM-GDL .
Synthetic data are becoming increasingly important for computational studies of cophylogeny, including machine learning, benchmarking, and method testing. Several generators have been proposed to produce such data, but each relies on different assumptions about host-symbiont coevolution. These assumptions are often implicit and rarely examined, even though results can depend strongly on the synthetic model being used. In this article, we present a systematic structural analysis of representative cophylogeny generators under controlled scenarios. The goal is to make their assumptions explicit and to understand how these choices shape the synthetic data they produce as well as the conclusions that may be drawn from them.
The study of evolution between species (phylogenetics) and the study of evolution within a species (population genetics) are highly related, as the same biological mechanisms are fundamental to both fields. Although both have been studied for a long time, their joint study in a unified setting has been prevented by the different time scales they consider and the different data types they employ. A similar discrepancy holds for their whole-genome specializations, comparative genomics and pangenomics. Two active areas in these fields are genome rearrangement studies and graphical pangenomics, respectively. Since the emergence of graphical pangenomics, these have existed as separate fields, despite observations that central data structures representing genomic variants in both fields are highly similar. While there exists a wealth of theoretical results for various rearrangement models in comparative genomics, the application to pangenomic data is hampered by the limitations of rearrangement problem formulations. On the practical side, pangenomes typically contain too many individual genomes for classical problems, such as the often NP-hard parsimony problems, to be solved, or for all-vs-all comparisons using rearrangement distances to be performed. On the theoretical side, some assumptions in the formulation of rearrangement problems, such as the assumption of an underlying tree, are inadequate for many pangenomes. In this work, we propose the Complete Ancestral Reconstruction for Pangenomes (CARP) problem, which overcomes these limitations while retaining intuitive relationships to both classical rearrangement problems and pangenome graphs.
Gene tree parsimony (GTP) is a common approach for efficient reconciliation of multiple discordant gene tree phylogenies for the inference of a single species tree. However, despite the popularity of GTP methods due to their low computational costs, prior work has shown that some commonly employed parsimony costs are statistically inconsistent under the multispecies coalescent process. Furthermore, a fine-grained analysis of the inconsistency has indicated potentially complimentary behavior of duplication and deep coalescence costs for symmetric and asymmetric species trees. In this work, we prove inconsistency of GTP estimators for all linear combinations of duplication, loss and deep coalescence scores. We also explore empirical implications of this result evaluating inference results of several GTP cost schemes under varying levels of incomplete lineage sorting.
The breakpoint distance is a fundamental measure for comparing gene orders and plays a central role in genome rearrangement phylogenetics. For more than two genomes, breakpoint medians are widely used to infer ancestral gene orders, yet their typical behavior is difficult to characterize due to the complex, non-geodesic geometry of genome space. A conjecture of Haghighi and Sankoff states that for random genomes, breakpoint medians tend to lie near one of the input genomes rather than constituting genuinely intermediate solutions. In this paper, we study breakpoint medians of independently and uniformly sampled signed multichromosomal genomes. We prove the Haghighi–Sankoff conjecture in this setting for the first time. We show that with high probability, any breakpoint median must remain close to a single input genome, rather than drawing adjacencies from multiple genomes. Our analysis relies on probabilistic bounds for usable adjacencies in random signed genomes and structural constraints of the signed multichromosomal model. We further provide a game-theoretic interpretation that explains why compromise strategies lead to increased total distance and the emergence of anti-medians.
Abundant discordance among gene trees is widely documented, but the causes of this heterogeneity are varied. Discordance among estimated gene trees can stem from real sources such as coalescent processes, hybridization, and horizontal gene transfer (HGT). It can also stem from errors in data, such as hidden paralogy, mistaken homology, bad alignment, and contamination. While some of these processes create stochastic and subtle changes in gene tree topologies (e.g., human closer to gorilla than to chimp), others can produce unexpected patterns (e.g., guinea pig sister to gorilla). Given a large number of gene trees and a median species tree, one could attempt to automatically find these outliers among gene trees. In this paper, we develop a method that uses quartet-based subtree-prune-and-regraft (SPR) moves, paired with gradient-boosted decision trees, to predict whether parts of a gene tree disagree with species trees in unusual ways. We show that our method, GBOD, is quite accurate in finding HGT events, but less so in other scenarios. Nevertheless, this combination of machine learning and phylogenetic features provides a promising framework for outlier detection.
Phylogenetic networks capture complex evolutionary relationships, including hybridization and gene flow, that are not easily represented by tree-based models. Orchards are a class of phylogenetic networks that have proven valuable for reconstructing evolution from ancestral profiles, and in which reticulation arcs model lateral gene transfers. Several methods exist to infer orchards, and distance measures between networks can be used to evaluate them against simulations or gold-standard datasets. One such measure is the maximum agreement cherry-reduced subnetwork (MACRS) of two orchards. Recently, Landry et al. proved that finding such a MACRS of two binary level-1 orchards is fixed-parameter tractable (FPT) with respect to the number of reticulations in both networks. In this paper, we show that this problem can be solved on arbitrary binary orchards, in FPT time with respect to the maximum level of the two networks. In particular, a MACRS of two binary level-1 orchards can be found in polynomial time.
Phylogenetic distance estimation and distance-based phylogeny reconstruction are well-studied cornerstone topics in phylogenetics. Classical approaches for both utilize mathematical or probabilistic graphical models of biomolecular sequence evolution. But model violations can occur and model-based analysis can be impacted as a result. Recent advances in statistical machine learning using deep neural networks provide an alternative in the form of representation learning. Newer applications of deep learning to phylogenetic distance estimation have followed. A number of challenges in this area remain, since state-of-the-art methods are often restricted to pairwise or subset-based analyses and retain other simplifying assumptions. In fact, classical model-based methods and representation learning are orthogonal and, as we show, their combination can be greater than the sum of its parts. We bridge these different approaches by synthesizing mathematical and logical constraints with statistical machine learning – an approach from physics-informed machine learning (PIML). Our algorithmic solution takes the form of a Transformer-based framework for learning phylogeny-informed representations directly from MSAs, which we apply to the task of phylogenetic distance estimation. The result is FIREFLY, a computational framework for “PHYlogeny-informed REpresentation learning to estimate PHyLogenetic dIstances”). We benchmarked FIREFLY’s performance against other state-of-the-art methods using simulated and empirical datasets. We found that FIREFLY improves both pairwise distance estimation accuracy and downstream phylogenetic inference compared with state-of-the-art methods. The gains are particularly pronounced under high indel rates and on estimated MSAs, where alignment errors and gap-induced uncertainty are most severe. Our results highlight the value of integrating phylogeny-based inductive bias into deep representation learning and suggest that MSA-level modeling offers a robust foundation for evolutionary inference under challenging conditions.
We address the problem of plasmid binning, that aims to group contigs—from a draft short-read assembly for a bacterial sample—into bins each expected to correspond to a plasmid present in the sequenced bacterial genome. We formulate the plasmid binning problem as a network multi-flow problem in the assembly graph and describe a Mixed-Integer Linear Program to solve it. We compare our new method, PlasBin-HMF, with state-of-the-art methods, MOB-recon, gplasCC, and PlasBin-flow, on a dataset of more than 500 bacterial samples, and show that PlasBin-HMF outperforms the other methods, by preserving the explainability.
Phylogenetic trees represent the evolutionary histories of taxa and support tasks such as clustering and Tree of Life reconstruction. Many established comparison methods, including the Robinson-Foulds (RF) distance, assume identical taxon sets. A methodological gap remains for trees with distinct but overlapping taxa. Existing approaches either prune non-common leaves, which can discard information, or complete both trees such that they share the same taxa. Completion is more comprehensive, but current methods typically ignore branch lengths, which are essential for identifying evolutionary patterns. This paper introduces k-Nearest Common Leaves (k-NCL), an algorithm for completing rooted phylogenetic trees defined on different but overlapping taxa. The method uses branch lengths and topological characteristics and does not rely on a specific distance measure. The k-NCL algorithm is designed to preserve evolutionary relationships in the trees under comparison. The running time is O(n^2) , where n is the size of the union of the two leaf sets. Additional properties include preservation of original distances and topology, symmetry, and uniqueness of the completion. Implemented in Python, k-NCL is evaluated on biological datasets of amphibians, birds, mammals, and sharks. Experimental results show that RF combined with k-NCL improves phylogenetic tree clustering performance compared to the RF(+) tree completion approach.
De novo genome assembly, the reconstruction of complete DNA sequences from reads without the use of a reference genome, remains one of the most challenging and fundamental problems in computational biology. A common method for de novo genome assembly involves creating an assembly graph of reads and defining a graph traversal that represents the genomic sequence. Recently, the first deep learning-based method, GNNome, was proposed to tackle this problem. Starting from an assembly graph, GNNome performs de novo assembly in two steps: binary edge classification and a greedy walk. However, we observe that the decisions of the greedy agent only matter in 0.86
In this paper, we study the genomic median of 3 problem with respect to the rank distance, which looks for a genome that minimizes the sum of the rank distances to three given genomes. We advance the knowledge on the mathematical properties of this problem, and settle its computational complexity, showing it is NP-hard. We also prove that the gap between exact and relaxed solutions of the problem can be arbitrarily large.
Short tandem repeats (STRs) are low-entropy regions in the genome, consisting of a short (1–6 bp) unit that is consecutively repeated multiple times. They are known for high mutational instability, due to so-called stutter-mutations, in which the number of units in the run increases or decreases. In particular, STRs with repeat unit length of 1–2 bp are prone to mutate even within several cell divisions. The extremely rapid accumulation of variation makes them interesting phylogenetic markers for retrospective single-cell lineage reconstruction. Here we model their mutational dynamics at the level of individual repeat unit motif and then aggregate length variations over many STR loci with the aim of obtaining a very fast “molecular clock”. We calibrate our model based on several datasets with known lineage structure prepared from cultured cells. We find that the mutational dynamics of STRs are reasonably consistent for a given cell-line, but vary among different ones. This suggests that the dynamics are not entirely explained by mutations in caretaker genes, rather, various other factors play a role—possibly tissue origin and differentiation state. Further data and research is necessary to asses their relative effects.
Phylogenetic networks are widespread representations of evolutionary histories for taxa that undergo hybridization or Lateral-Gene Transfer (LGT) events. There are now many tools to reconstruct such networks, but no clearly established metric to compare them. Such metrics are needed, for example, to evaluate predictions against a simulated ground truth. Despite years of effort in developing metrics, known dissimilarity measures either do not distinguish all pairs of different networks, or are extremely difficult to compute. Since it appears challenging, if not impossible, to create the ideal metric for all classes of networks, it may be relevant to design them for specialized applications. In this article, we introduce a metric on LGT networks, which consist of trees with additional arcs that represent lateral gene transfer events. Our metric is based on edit operations, namely the addition/removal of transfer arcs, and the contraction/expansion of arcs of the base tree, allowing it to connect the space of all LGT networks. We show that it is linear-time computable if the order of transfers along a branch is unconstrained but NP-hard otherwise, in which case we provide a fixed-parameter tractable (FPT) algorithm in the level. We implemented our algorithms and demonstrate their applicability on three numerical experiments. Full online version: https://www.biorxiv.org/content/10.1101/2025.11.20.689557
De novo genome assembly remains a central challenge in computational biology, particularly for diploid genomes where maternal and paternal haplotypes must be accurately resolved. Existing assemblers achieve impressive results through carefully designed heuristics, yet modern deep learning methods remain largely unexplored in the diploid setting. We present DipGNNome, the first deep learning–based framework for diploid de novo genome assembly. Our approach formulates genome assembly as an edge classification and graph traversal problem, given haplotype-aware assembly graphs. We train a graph neural network (GNN) to guide contig construction as the layout phase in an Overlap-Layout-Consensus genome assembly pipeline. To enable this, we establish a novel pipeline for generating diploid graphs with ground-truth edge labels, providing the first systematic way to produce training data for machine learning models in this domain. This framework creates a foundation for applying and extending graph-based deep learning to diploid assembly. DipGNNome creates assemblies comparable to SotA and demonstrates the feasibility of deep learning for diploid assembly and introduces a paradigm that bridges algorithmic genomics with graph representation learning. Our code, dataset and trained model are openly available at https://github.com/lbcb-sci/DipGNNome .
How DNA sequence encodes gene regulation remains a central challenge in regulatory genomics. Transcription factors (TFs) are key mediators of this process, binding to specific sequence motifs to control gene expression. Yet, predicting where they bind from sequence alone remains a challenging problem. A cross-species angle offers two complementary benefits: it tests whether trained models have learned conserved, biochemically grounded rules that generalize across species, and it enables binding prediction in species where experimental data is scarce. Key challenges in this context are that TF binding sites undergo rapid evolutionary turnover, and that there are systematic distributional differences between species’ genomes. We present MORALE, a domain adaptation framework for cross-species TF binding prediction. By aligning the first and second moments of sequence embeddings across species during training, MORALE learns species-invariant representations without adversarial training, additional parameters, or architectural changes. Applied to liver ChIP-seq data from two species (human, mouse) and five species (adding rhesus macaque, rat, and dog), MORALE consistently matches or outperforms gradient reversal (the adversarial baseline) across all TFs, and avoids the performance degradation below the no-adaptation baseline that gradient reversal can exhibit. In the five-species setting, MORALE surpasses a human-only model, demonstrating that moment alignment can unlock cross-species generalization that neither multi-species training nor adversarial adaptation achieves alone. MORALE also recovers TF binding motifs more faithfully than the adversarial approach, suggesting its representations capture biologically meaningful sequence features. As a closed-form, parameter-free regularizer, MORALE integrates into any embedding-based sequence model.
Likelihood-based inference under the multispecies coalescent provides accurate estimates of species trees. However, maximum likelihood and Bayesian inference are both computationally very demanding. Pseudo-likelihood has been previously proposed as a computationally efficient alternative to full phylogenetic likelihood calculations in the context of maximum likelihood estimation. However, theoretical and practical aspects of pseudo-likelihood in the context of Bayesian inference have not been explored. In this work, we provide strong theoretical guarantees and an empirical evaluation of pseudo-likelihood-based Bayesian inference of species trees under the multispecies coalescent. Our contributions are threefold. First, we prove a bound on the convergence rate for species tree topology inference and a Bernstein-von Mises result for branch lengths under model misspecification. Second, we provide an empirical comparison of full- and pseudo-likelihood-based Bayesian inference on synthetic data. Finally, we demonstrate the practical scalability of pseudo-likelihood-based inference by analyzing two biological datasets.
Network Phylogenetic Diversity (Network-PD) is a measure for the diversity of a set of species based on a rooted phylogenetic network (with branch lengths and inheritance probabilities on the reticulation edges) describing the evolution of those species. We consider the Max-Network-PD problem: Given such a network, find k species with maximum Network-PD score. We show that this problem is fixed-parameter tractable (FPT) for binary networks, by describing an optimal algorithm running in O(2r log(k)(n + r)) time, with n the total number of species in the network and r its reticulation number. Furthermore, we show that Max-Network-PD is NP-hard for level-1 networks, proving that, unless P=NP, the FPT approach cannot be extended by using the level as parameter instead of the reticulation number.