We introduce two novel distances for comparing rooted phylogenetic networks based on the ⊖-operator, which removes a vertex while preserving the ancestor relations among the remaining vertices. The distance d_⊖ measures the minimum number of such removals needed to obtain isomorphic networks, whereas d_⊖^- ignores shortcut arcs and therefore compares the induced ancestry structures. We show that d_⊖ is a metric up to leaf-fixing isomorphism and that d_⊖^- is a metric up to shortcut-free isomorphism. Moreover, both distances extend the Robinson–Foulds distance on phylogenetic trees and are bounded below by the hardwired cluster distances. For several broad network classes, including tree-child, normal, level-1, and regular networks, d_⊖^- can be computed in polynomial time. In contrast, computing d_⊖ is NP-hard, W[2]-hard when parameterized by the distance value, and admits no polynomial-time constant-factor approximation unless P=NP. Although computing d_⊖^- is NP-hard in general, for distinct-cluster networks it reduces to Vertex Cover, yielding a fixed-parameter algorithm and a polynomial-time 2-approximation.
REvolutionH-tl is a fast, scalable, and integrated software platform for inferring orthology relationships, gene trees, species trees, and reconciled evolutionary scenarios directly from sequence data. Built upon the formal framework of best match graphs (BMGs), REvolutionH-tl predicts orthogroups and orthologous gene pairs with high accuracy, requiring neither precomputed trees nor multiple external tools. The software reconstructs event-labeled gene and species trees, seamlessly integrating reconciliation to produce fast, accurate, and biologically insightful evolutionary scenarios. Through extensive benchmarking on synthetic datasets with known ground truth, REvolutionH-tl outperforms or matches the accuracy of established tools such as OrthoFinder, Proteinortho, RAxML, GeneRax, and RANGER-DTL, while achieving significantly lower runtimes. A key innovation of REvolutionH-tl is its built-in support for detailed, publication-ready visualizations, which allow users to explore genome evolution dynamics, orthogroup composition, and reconciliation results with clarity and ease. These visual features position REvolutionH-tl as the first platform of its kind to combine analytical precision with intuitive interpretability. The software is open-source, cross-platform, and freely available at https://pypi.org/project/revolutionhtl/ , providing a robust solution for large-scale evolutionary analyses in comparative genomics.
Evolutionary histories are often represented by rooted phylogenetic networks, whose leaves correspond to extant taxa and whose internal vertices represent ancestral lineages. Since such histories must usually be inferred from incomplete data, in particular from genomic sequences of present-day taxa, one often obtains only local information about relative evolutionary proximity. For instance, sequence data may suggest that two taxa x and y are more closely related to each other than either is to a third taxon z. This information is classically encoded by a rooted triple xy|z. In this paper, we study rooted triples in phylogenetic networks under an ancestor-based interpretation: xy|z is displayed if the unique least common ancestor (LCA) of x and y lies strictly below the unique LCA of x and z, respectively of y and z, and the latter two LCAs coincide. We also introduce anchored triples xy|z, which retain only the asymmetric comparison that the LCA of x and y lies below the LCA of x and z. This relaxation is natural in networks, where different pairwise ancestral relationships need not behave as they do in trees. We consider several variants of consistency problems for ordinary and anchored triples, both with and without forbidden triples. Somewhat surprisingly, these ancestor-based consistency questions for triples in phylogenetic networks do not appear to have been addressed before despite their direct biological interpretation and the fact that such constraints can be inferred naturally from genomic sequence data. By translating these questions into realization problems for required and forbidden LCA-constraints, we show that all resulting problems can be solved in polynomial time. Moreover, whenever a solution exists, a suitable realizing DAG and phylogenetic network can be constructed within the same time bound.
We explore the connections between clusters and least common ancestors (LCAs) in directed acyclic graphs (DAGs), focusing on the interplay between so-called 7-lcarelevant DAGs and DAGs with the 7-lca-property. Here, 7 denotes a set of integers. In 7-lca-relevant DAGs, each vertex is the unique LCA for some subset A of leaves of size |A| E 7, whereas in a DAG with the 7-lca-property there exists a unique LCA for every subset A of leaves satisfying |A| E 7. We elaborate on the difference between these two properties and establish their close relationship to pre-7-ary and 7-ary set systems. This, in turn, generalizes results established for (pre-) binary and kary set systems. Moreover, we build upon recently established results that use a simple operator e, enabling the transformation of arbitrary DAGs into 7-lca-relevant DAGs. This process reduces unnecessary complexity while preserving key structural properties of the original DAG. The set CG consists of all clusters in a DAG G, where clusters correspond to the descendant leaves of vertices. While in some cases CH = CG when transforming G into an 7-lca-relevant DAG H, it often happens that certain clusters in CG do not appear as clusters in H. To understand this phenomenon in detail, we characterize the subset of clusters in CG that remain in H for DAGs G with the 7-lca-property. Furthermore, we show that the set W of vertices required to transform G into H = G e W is uniquely determined for such DAGs. This, in turn, allows us to show that the "shortcutfree" version of the transformed DAG H is always a tree or a galled-tree whenever CG represents the clustering system of a tree or galled-tree and G has the 7-lca-property. In the latter case CH = CG always holds. (c) 2025 The Author(s). Published by Elsevier B.V. This is an open access article under the CC BY license (http://creativecommons.org/licenses/by/4.0/).
Phylogenetic networks provide a framework for representing evolutionary histories involving reticulate events such as hybridization or horizontal gene transfer. A central problem is to infer such networks from local structural information. In this paper, we study network inference from least common ancestor (LCA) constraints, which specify relative ancestral relationships between pairs of taxa. While previous work has characterized when a set of required LCA constraints can be realized by a phylogenetic network, practical applications may also involve constraints that must be explicitly avoided, for example due to biological prior knowledge. We therefore consider the realization problem for pairs (R,F), where R is a set of required LCA-constraints and F is a set of forbidden ones. Since there are several natural ways to formalize what it means for a network to avoid a forbidden LCA-constraint, we study three such variants. For each of them, we characterize exactly when there exists a phylogenetic network that realizes all constraints in R while avoiding all constraints in F in the respective sense. Based on these characterizations, we derive polynomial-time algorithms that decide the existence of such networks and construct one whenever it exists.
Phylogenetic networks and, more generally, directed acyclic graphs (DAGs) represent hierarchical structure beyond trees, for instance in the presence of reticulate evolutionary events such as hybridization or horizontal gene transfer. A central question is which parts of such graphs are essential with respect to leaf-observable information, and which parts can be removed without changing this information. Resolving this question can lead to principled simplification methods for phylogenetic networks, such as the recent normalization approach of Francis et al. In this paper, we study this question from three related perspectives: clusters displayed by a DAG G, least common ancestors (LCAs) of subsets of its leaf set, and visibility, a path-based property of vertices. We first introduce an LCA-based simplification procedure called i-regularization. For a DAG G and i≥ 1, the DAG _i(G) retains precisely those vertices that occur as unique LCAs of leaf subsets of size at most i, removes the remaining non-leaf vertices by a graph-editing operation ⊖, and then deletes shortcuts. We show that _i(G) preserves all such LCAs, is i-lca-relevant, and admits a cluster-level description: it is regular, i.e., isomorphic to the Hasse diagram of the corresponding lca-clusters. We then compare LCA-based regularization with normalization. Using the same ⊖-operator, we describe the cover construction underlying normalization, identify visible vertices that are nevertheless removed, and characterize when regularization and normalization coincide. Together, these results provide a unified framework for cluster-based, LCA-based, and visibility-based simplifications of DAGs and phylogenetic networks.
A least common ancestor (LCA) of two leaves in a directed acyclic graph (DAG) is a vertex that is an ancestor of both of the leaves, and for which no proper descendant of this vertex is also an ancestor of the two leaves. LCAs play a central role in representing hierarchical relationships in rooted trees and, more generally, in DAGs. In 1981, Aho et al. introduced the problem of determining whether a collection of pairwise LCA constraints on a set X, of the form (i, j) < (k, l) with i, j, k, l is an element of X, can be realized by a rooted tree with leaf set X such that, whenever (i, j) < (k, l) holds, the LCA of i and j in the tree is a descendant of the LCA of k and l. In particular, they presented a polynomial-time algorithm, called Build, to solve this problem. In many cases, however, such constraints cannot be realized by any tree. In these situations, it is natural to ask whether they can still be realized by a more general directed acyclic graph (DAG). In this paper, we extend Aho et al.'s problem from trees to DAGs, providing a theoretical and algorithmic framework for reasoning about LCA constraints in this more general setting. More specifically, given a collection R of LCA constraints, we introduce the notion of the +-closure R+ which captures additional LCA relations implied by R. Using this closure, we then associate canonical DAG G(R) to R and show that a collection of constraints R can be realized by a DAG in terms of LCAs if and only if R is realized in this way by G(R). In addition, we adapt this construction to obtain phylogenetic networks, a special class of DAGs widely used in phylogenetics. In particular, we define a canonical phylogenetic network N-R for realizing R, and prove that N-R is regular, that is, it coincides with the Hasse diagram of the set system underlying N-R. Finally, we show that, for any DAG-realizable collection R, the classical closure of R, that is the collection of all LCA constraints that must hold in every DAG realizing R, coincides with its +-closure. All constructions introduced in this paper can be computed in polynomial time, and we provide explicit algorithms for each. All algorithms developed in this paper are implemented in the freely available Python package RealLCA.
Horizontal gene transfer is an important contributor to evolution. Following Walter M. Fitch, two genes are xenologs if at least one HGT separates them. More formally, the directed Fitch graph has a set of genes as its vertices, and directed edges (x, y) for all pairs of genes x and y for which y has been horizontally transferred at least once since it diverged from the last common ancestor of x and y. Subgraphs of Fitch graphs can be inferred by comparative sequence analysis. In many cases, however, only partial knowledge about the “full” Fitch graph can be obtained. Here, we characterize Fitch-satisfiable graphs that can be extended to a biologically feasible “full” Fitch graph and derive a simple polynomial-time recognition algorithm. We then proceed to show that several versions of finding the Fitch graph with total maximum (confidence) edge-weights are NP-hard. In addition, we provide a greedy-heuristic for “optimally” recovering Fitch graphs from partial ones. Somewhat surprisingly, even if ∼ 80% of information of the underlying input Fitch-graph G is lost (i.e., the partial Fitch graph contains only ∼ 20% of the edges of G), it is possible to recover ∼ 90% of the original edges of G on average.
Encoding phylogenetic networks by suitable substructures is a central problem in phylogenetic combinatorics. We study encodings based on least common ancestor (LCA) constraints. For a directed acyclic graph (DAG) $G$ with leaf set $X$, we consider the relation on pairs of leaves in which $(ab,xy)$ records that the LCAs of $a,b$ and $x,y$ are well-defined and that the former is a descendant of the latter. We first identify precisely which part of $G$ is determined by this relation. To this end, we compare the canonical DAG constructed from the LCA relation with the 2-regularization of $G$, obtained by removing all vertices that are not LCAs of one or two leaves and then deleting shortcut edges. We prove that these two DAGs are isomorphic. Hence the obstruction to encoding a graph by its LCA relation is exactly the information lost under 2-regularization. This yields a general reconstruction principle, which we apply to several natural classes of phylogenetic networks. In particular, we show that shortcut-free 2-LCA-relevant DAGs, phylogenetic trees, regular level-1 networks, regular networks with binary clustering systems, regular networks whose clustering systems are closed weak hierarchies, strong-phylogenetic normal networks, separated phylogenetic normal networks, and binary normal networks are encoded by their LCA relations. We also introduce a sparse triple-like restriction consisting only of comparisons of the form $(ab,ac)$, where $a,b,c\in X$ are pairwise distinct. For graphs with the 2-LCA property, we show that this sparse relation, together with the leaf set, determines the full LCA relation after a natural closure operation. Consequently, several of the above classes can be reconstructed, up to isomorphism, from the sparse relation in polynomial time.
Orthologous genes, which arise through speciation, play a key role in comparative genomics and functional inference. In particular, graph-based methods allow for the inference of orthology estimates without prior knowledge of the underlying gene or species trees. This results in orthology graphs, where each vertex represents a gene, and an edge exists between two vertices if the corresponding genes are estimated to be orthologs. Orthology graphs inferred under a tree-like evolutionary model must be cographs. However, real-world data often deviate from this property, either due to noise in the data, errors in inference methods or, simply, because evolution follows a network-like rather than a tree-like process. The latter, in particular, raises the question of whether and how orthology graphs can be derived from or, equivalently, are explained by phylogenetic networks. In this work, we study the constraints imposed on orthology graphs when the underlying evolutionary history follows a phylogenetic network instead of a tree. We show that any orthology graph can be represented by a sufficiently complex level-k network. However, such networks lack biologically meaningful constraints. In contrast, level-1 networks provide a simpler explanation, and we establish characterizations for level-1 explainable orthology graphs, i.e., those derived from level-1 evolutionary histories. To this end, we employ modular decomposition, a classical technique for studying graph structures. Specifically, an arbitrary graph is level-1 explainable if and only if each primitive subgraph is a near-cograph (a graph in which the removal of a single vertex results in a cograph). Additionally, we present a linear-time algorithm to recognize level-1 explainable orthology graphs and to construct a level-1 network that explains them, if such a network exists. Finally, we demonstrate the close relationship of level-1 explainable orthology graphs to the substitution operation, weakly chordal and perfect graphs, as well as graphs with twin-width at most 2.
The Djokovi\'{c}-Winkler relation $\Theta$ is a binary relation defined on the edge set of a given graph that is based on the distances of certain vertices and which plays a prominent role in graph theory. In this paper, we explore the relatively uncharted ``reflexive complement'' $\overline\Theta$ of $\Theta$, where $(e,f)\in \overline\Theta$ if and only if $e=f$ or $(e,f)\notin \Theta$ for edges $e$ and $f$. We establish the relationship between $\overline\Theta$ and the set $\Delta_{ef}$, comprising the distances between the vertices of $e$ and $f$ and shed some light on the intricacies of its transitive closure $\overline\Theta^*$. Notably, we demonstrate that $\overline\Theta^*$ exhibits multiple equivalence classes only within a restricted subclass of complete multipartite graphs. In addition, we characterize non-trivial relations $R$ that coincide with $\overline\Theta$ as those where the graph representation is disconnected, with each connected component being the (join of) Cartesian product of complete graphs. The latter results imply, somewhat surprisingly, that knowledge about the distances between vertices is not required to determine $\overline\Theta^*$. Moreover, $\overline\Theta^*$ has either exactly one or three equivalence classes.
Rooted phylogenetic networks, or more generally, directed acyclic graphs (DAGs), are widely used to model species or gene relationships that traditional rooted trees cannot fully capture, especially in the presence of reticulate processes or horizontal gene transfers. Such networks or DAGs are typically inferred from observable data (e.g., genomic sequences of extant species), providing only an estimate of the true evolutionary history. However, these inferred DAGs are often complex and difficult to interpret. In particular, many contain vertices that do not serve as least common ancestors (LCAs) for any subset of the underlying genes or species, thus may lack direct support from the observable data. In contrast, LCA vertices are witnessed by historical traces justifying their existence and thus represent ancestral states substantiated by the data. To reduce unnecessary complexity and eliminate unsupported vertices, we aim to simplify a DAG to retain only LCA vertices while preserving essential evolutionary information. In this paper, we characterize LCA -relevant and lca -relevant DAGs, defined as those in which every vertex serves as an LCA (or unique LCA) for some subset of taxa. We introduce methods to identify LCAs in DAGs and efficiently transform any DAG into an LCA -relevant or lca -relevant one while preserving key structural properties of the original DAG or network. This transformation is achieved using a simple operator " ⊖ " that mimics vertex suppression.
Directed acyclic graphs (DAGs) are fundamental structures used across many scientific fields. A key concept in DAGs is the least common ancestor (LCA), which plays a crucial role in understanding hierarchical relationships. Surprisingly little attention has been given to DAGs that admit a unique LCA for every subset of their vertices. Here, we characterize such global lca-DAGs and provide multiple structural and combinatorial characterizations. We show that global lca-DAGs have a close connection to join semi-lattices and establish a connection to forbidden topological minors. In addition, we introduce a constructive approach to generating global lca-DAGs and demonstrate that they can be recognized in polynomial time. We investigate their relationship to clustering systems and other set systems derived from the underlying DAGs.
In phylogenetics, reconstructing rooted trees from distances between taxa is a common task. Böcker and Dress generalized this concept by introducing symbolic dating maps δ :X × X →Υ , where distances are replaced by symbols, and showed that there is a one-to-one correspondence between symbolic ultrametrics and labeled rooted phylogenetic trees. Many combinatorial structures fall under the umbrella of symbolic dating maps, such as 2-dissimilarities, symmetric labeled 2-structures, or edge-colored complete graphs, and are here referred to as strudigrams. Strudigrams have a unique decomposition into non-overlapping modules, which can be represented by a modular decomposition tree (MDT). In the absence of prime modules, strudigrams are equivalent to symbolic ultrametrics, and the MDT fully captures the relationships δ (x,y) between pairs of vertices x,y ∈ X through the label of their least common ancestor in the MDT. However, in the presence of prime modules, appearing as prime vertices in the MDT, this information is generally hidden. To provide this missing structural information, we aim to locally replace the prime vertices in the MDT to obtain networks that capture full information about the strudigrams. While starting with the general framework of prime-vertex replacement networks, we then focus on a specific type of such networks obtained by replacing prime vertices with so-called galls, resulting in labeled galled-trees. We introduce the concept of galled-tree explainable (GaTEx) strudigrams, provide their characterization, and demonstrate that recognizing these structures and reconstructing the labeled networks that explain them can be achieved in polynomial time.
We present an exact algorithm for computing the connected Maximum Common Subgraph (MCS) across multiple graphs, where edges or vertices may additionally be labeled to account for possible atom types or bond types, a classical labeling used in molecular graphs. Our approach leverages modular product graphs and a modified Bron-Kerbosch algorithm to enumerate maximal cliques, ensuring all intermediate solutions are retained. A pruning heuristic efficiently reduces the modular product size, improving computational feasibility. Additionally, we introduce a graph ordering strategy based on graph-kernel similarity measures to optimize the search process. Our method is particularly relevant for bioinformatics and cheminformatics, where identifying conserved structural motifs in molecular graphs is crucial. Empirical results on molecular datasets demonstrate that our approach is exact, scalable and fast.
Median graphs are connected graphs in which for all three vertices there is a unique vertex that belongs to shortest paths between each pair of these three vertices. To be more formal, a graph $G$ is a median graph if, for all $\mu, u,v\in V(G)$, it holds that $|I(\mu,u)\cap I(\mu,v)\cap I(u,v)|=1$ where $I(x,y)$ denotes the set of all vertices that lie on shortest paths connecting $x$ and $y$. In this paper we are interested in a natural generalization of median graphs, called $k$-median graphs. A graph $G$ is a $k$-median graph, if there are $k$ vertices $\mu_1,\dots,\mu_k\in V(G)$ such that, for all $u,v\in V(G)$, it holds that $|I(\mu_i,u)\cap I(\mu_i,v)\cap I(u,v)|=1$, $1\leq i\leq k$. By definition, every median graph with $n$ vertices is an $n$-median graph. We provide several characterizations of $k$-median graphs that, in turn, are used to provide many novel characterizations of median graphs.
Most genes are part of larger families of evolutionary-related genes. The history of gene families typically involves duplications and losses of genes as well as horizontal transfers into other organisms. The reconstruction of detailed gene family histories, i.e., the precise dating of evolutionary events relative to phylogenetic tree of the underlying species has remained a challenging topic despite their importance as a basis for detailed investigations into adaptation and functional evolution of individual members of the gene family. The identification of orthologs, moreover, is a particularly important subproblem of the more general setting considered here. In the last few years, an extensive body of mathematical results has appeared that tightly links orthology, a formal notion of best matches among genes, and horizontal gene transfer. The purpose of this chapter is to broadly outline some of the key mathematical insights and to discuss their implication for practical applications. In particular, we focus on tree-free methods, i.e., methods to infer orthology or horizontal gene transfer as well as gene trees, species trees, and reconciliations between them without using a priori knowledge of the underlying trees or statistical models for the inference of phylogenetic trees. Instead, the initial step aims to extract binary relations among genes.
Orthology detection from sequence similarity remains a difficult and computationally expensive problem for gene families with large numbers of gene duplications and losses. REvolutionH-tl implements a new graph-based approach to identify orthogroups, orthology, and paralogy relationships first, and it uses this information in a second step to infer event-labeled gene trees and their reconciliation with an inferred species tree. It avoids using gene trees and species trees upon input and settles for a maximal subtree reconciliation in cases where noise or horizontal gene transfer precludes a global reconciliation. The accuracy of the tool is comparable to competing tools at substantially reduced computational cost. REvolutionH-tl is freely available at https://pypi.org/project/revolutionhtl/ .
Polygons are cycles embedded into the plane; their vertices are associated with x- and y-coordinates and the edges are straight lines. Here, we consider a set of polygons with pairwise non-overlapping interior that may touch along their boundaries. Ideas of the sweep line algorithm by Bajaj and Dey for non-touching polygons are adapted to accommodate polygons that share boundary points. The algorithms established here achieves a running time of 𝒪(n+Nlog N), where n is the total number of vertices and N<n is the total number of "maximal outstretched segments" of all polygons. It is asymptotically optimal if the number of maximal outstretched segments per polygon is bounded. In particular, this is the case for convex polygons.
Phylogenetic networks play an important role in evolutionary biology as, other than phylogenetic trees, they can be used to accommodate reticulate evolutionary events such as horizontal gene transfer and hybridization. Recent research has provided a lot of progress concerning the reconstruction of such networks from data as well as insight into their graph theoretical properties. However, methods and tools to quantify structural properties of networks or differences between them are still very limited. For example, for phylogenetic trees, it is common to use balance indices to draw conclusions concerning the underlying evolutionary model, and more than twenty such indices have been proposed and are used for different purposes. One of the most frequently used balance index for trees is the so-called total cophenetic index, which has several mathematically and biologically desirable properties. For networks, on the other hand, balance indices are to-date still scarce.In this contribution, we introduce the weighted total cophenetic index as a generalization of the total cophenetic index for trees to make it applicable to general phylogenetic networks. As we shall see, this index can be determined efficiently and behaves in a mathematical sound way, i.e., it satisfies so-called locality and recursiveness conditions. In addition, we analyze its extremal properties and, in particular, we investigate its maxima and minima as well as the structure of networks that achieve these values within the space of so-called level-1 networks. We finally briefly compare this novel index to the two other network balance indices available so-far.
Vincent Moulton合作论文数University of East Anglia;School of Computing Sciences7