There are few, if any, algorithms in statistical phylogenetics which are used more heavily than Felsenstein's 1973 pruning method for computing the likelihood of a tree. We present LvD, (Likelihood via Decomposition), an alternative to Felsenstein's algorithm based on a different decomposition of the underlying phylogeny. It works for all standard nucleotide models. The new algorithm allows updates of the likelihood calculation in worst case O(log n) time with n taxa, as opposed to worst case O(n) time for existing methods. In practice this leads to appreciable improvements in likelihood calculations, the extent of speed-up depending on how balanced or unbalanced the trees are. We explore implications for parallel computing, and show that the approach allows likelihoods to be computed in O(log n) parallel time per site, compared to (worst case) O(n) time. We implemented and applied the algorithm to large numbers of simulated and empirical data sets and showed that these theoretical advances lead to a significant practical speed-up, although the extent of the improvement depends on how balanced the phylogenies already are.
Submodular functions and their close relatives play a key role in combinatorial optimization, decision theory and potential theory. Part of their importance and usefulness stems from the connections with convex functions and polytopes. Here we explore connections between these functions and metric theory, with the bridge provided by diversities, a recently developed generalization of metric spaces to (finite) sets rather than just pairs. Both submodular functions and strongly submodular functions correspond to natural classes of diversities. Submodular diversities, as we define them here, are essentially non-decreasing, intersecting submodular functions which vanish on singletons. We prove new geometric embedding results for these diversities. In particular we show that submodular, strongly submodular, and XOS functions can be represented by the generalized circumradius, a set function in convex analysis equal to the amount a given convex body needs to be stretched to cover a set of points.
In phylogenetics and other areas of classification, the Buneman graph is commonly used to represent a collection of bipartitions or splits of a (finite) set X in order to display evolutionary relationships. The set X usually corresponds to a set of taxa (or species), and the splits are usually derived from molecular sequence data associated to the taxa. One issue with this approach is that missing molecular data can lead to bipartitions of subsets of X or partial splits, instead of splits of the full set X. In this paper, we show that the definition of the Buneman graph can be naturally extended to collections of partial splits of a set X. Just as with splits, we show that the graph so obtained is an X-labeled median graph but, in contrast to the usual Buneman graph, the elements in X are represented by convex subsets of the vertex set of the graph instead of single vertices. We also show that the Buneman graph for a collection of partial splits is closely related to subtree distances. In particular, for a collection S of weighted partial splits that satisfies a certain pairwise compatibility condition, we show that the corresponding edge-weighted Buneman graph is the unique minimal tree that represents the subtree distanced corresponding to S. Moreover, we show that in this special situation the Buneman graph can also be considered as a type of configuration space for the set of all tree-metrics that minimally extend the subtree distanced. (c) 2025 The Author(s). Published by Elsevier B.V. This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/).
Diversities are an extension of the concept of a metric space which assign a non-negative value to every finite set of points, rather than just pairs. A general theory of diversities has been developed which exhibits many deep analogies to metric space theory but also veers off in new directions. Just as many of the most important aspects of metric space theory involve metrics defined on ℝ^k, many applications of diversity theory require a specialized theory for diversities defined on ℝ^k, as we develop here. We focus on two fundamental classes of diversities defined on ℝ^k: those that are Minkowski linear and those that are Minkowski sublinear. Many well-known functions in convex analysis belong to these classes, including diameter, circumradius and mean width. We derive surprising characterizations of these classes, and establish elegant connections between them. Motivated by classical results in metric geometry, and connections with combinatorial optimization, we then examine embeddability of finite diversities into ℝ^k. We prove that a finite diversity can be embedded into a linear diversity exactly when it is of negative type and that it can be embedded into a sublinear diversity exactly when it corresponds to a generalized circumradius.
We characterize when a set of distances d(x, y) between elements in a set X have a subtree representation, a real tree T and a collection {Sx}x is an element of X of subtrees of T such that d(x, y) equals the length of the shortest path in T from a point in Sx to a point in Sy for all x, y is an element of X. The characterization was first established for finite X by Hirai (2006) using a tight span construction defined for distance spaces, metric spaces without the triangle inequality. To extend Hirai's result beyond finite X we establish fundamental results of tight span theory for general distance spaces, including the surprising observation that the tight span of a distance space is hyperconvex. We apply the results to obtain the first characterization of when a diversity-a generalization of a metric space which assigns values to all finite subsets of X, not just to pairs-has a tight span which is tree-like. (c) 2025 Elsevier B.V. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
The first comparative pre-treatment study of Miscanthus (Mxg) and sugarcane bagasse (SCB) using steam explosion (SE) and pressurised disc refining (PDR) pretreatment to optimise xylose and xylo-oligosaccharide release is described. The current investigation aimed to 1) Develop optimised batch-wise steam explosion parameters for Mxg and SCB, 2) Scale from static batch steam explosion to dynamic continuous pressurised disc refining, 3) Identify, understand, and circumvent scale-up production hurdles.Optimised SE parameters released 82% (Mxg) and 100% (SCB) of the available xylan. Scaling to PDR, Miscanthus yielded 85% xylan, highlighting how robust scouting assessments for boundary process parameters can result in successful technical transfer. In contrast, SCB technical transfer was not straightforward, with significant differences observed between the two processes, 100% (SE) and 58% (PDR).This report underlines the importance of feedstock-specific pretreatment strategies to underpin process development, scale-up, and optimisation of carbohydrate release from biomass.
Solid-state fermentation (SSF) is a sustainable method to convert food waste and plant biomass into novel foods for human consumption. Surplus bread crusts (BC) have the structural capacity to serve as an SSF scaffold, and their nutritional value could be increased in combination with perennial ryegrass (PRG), a biorefining feedstock with high-quality protein but an unpleasant sensory profile. SSF with Rhizopus oligosporus was investigated with these substrates to determine if the overall nutritional value could be increased. The BC-PRG SSFs were conducted for up to 72 h, over which time the starch content had decreased by up to 89.6%, the total amino acid (AA) content increased by up to 141.9%, and the essential amino acid (EAA) content increased by up to 54.5%. The BC-PRG SSF demonstrated that this process could potentially valorise BC and PRG, both widely available but underexplored substrates, for the production of alternative proteins.
NeighborNet constructs phylogenetic networks to visualize distance data. It is a popular method used in a wide range of applications. While several studies have investigated its mathematical features, here we focus on computational aspects. The algorithm operates in three steps. We present a new simplified formulation of the first step, which aims at computing a circular ordering. We provide the first technical description of the second step, the estimation of split weights. We review the third step by constructing and drawing the network. Finally, we discuss how the networks might best be interpreted, review related approaches, and present some open questions.
The standard models of sequence evolution on a tree determine probabilities for every character or site pattern. A flattening is an arrangement of these probabilities into a matrix, with rows corresponding to all possible site patterns for one set $A$ of taxa and columns corresponding to all site patterns for another set $B$ of taxa. Flattenings have been used to prove difficult results relating to phylogenetic invariants and consistency and also form the basis of several methods of phylogenetic inference. We prove that the rank of the flattening equals $r^{\ell_T(A|B)}$, where $r$ is the number of states and $\ell_T(A|B)$ is the parsimony length of the binary character separating $A$ and $B$. This result corrects an earlier published formula and opens up new applications for old parsimony theorems. Since completing this work, we have learnt that an equivalent result has been proved much earlier by Casanellas and Fern\'andez-S\'anchez, using a different proof strategy.
The generalized circumradius of a set of points A⊆ℝ^d with respect to a convex body K equals the minimum value of λ≥ 0 such that a translate of λ K contains A . Each choice of K gives a different function on the set of bounded subsets of ℝ^d ; we characterize which functions can arise in this way. Our characterization draws on the theory of diversities , a recently introduced generalization of metrics from functions on pairs to functions on finite subsets. We additionally investigate functions which arise by restricting the generalized circumradius to a finite subset of ℝ^d . We obtain elegant characterizations in the case that K is a simplex or parallelotope.
AbstractXylitol has been recognized by the US Department of Energy (DOE) as one of the top 12 value-added chemicals obtained from biomass, with a world market of 200,000 tonnes per year. The global xylitol market is expected to reach a value of US$ 1 Billion by 2026 growing at a compound annual growth rate (CAGR) of 5.8% during 2021–2026. Historically, the commercial xylitol production process has been dependent on the chemical hydrogenation of xylose. Several xylitol production plants, mainly in China that use the chemical process have had to reduce their production capacity to address regulations governing sustainability and environmental standards. In this chapter, key challenges and possible solutions for fermentative xylitol production at commercial scale are discussed in terms of: (1) Feedstock supply for commercial production plants; (2) Industrial biomass pretreatment; and (3) Lessons learned from industrial operations. These are drawn together to identify technology gaps and scaling-up challenges in light of the capital expenditure required to build a state-of-the art xylitol industrial biotechnology (IB) production facility and the potential to reduce climate change impact and contribute towards achieving net-zero targets.
In many phylogenetic applications, such as cancer and virus evolution, time trees, evolutionary histories where speciation events are timed, are inferred. Of particular interest are clock-like trees, where all leaves are sampled at the same time and have equal distance to the root. One popular approach to model clock-like trees is coalescent theory, which is used in various tree inference software packages. Methodologically, phylogenetic inference methods require a tree space over which the inference is performed, and the geometry of this space plays an important role in statistical and computational aspects of tree inference algorithms. It has recently been shown that coalescent tree spaces possess a unique geometry, different from that of classical phylogenetic tree spaces. Here we introduce and study a space of discrete coalescent trees. They assume that time is discrete, which is natural in many computational applications. This tree space is a generalisation of the previously studied ranked nearest neighbour interchange space, and is built upon tree-rearrangement operations. We generalise existing results about ranked trees, including an algorithm for computing distances in polynomial time, and in particular provide new results for both the space of discrete coalescent trees and the space of ranked trees. We establish several geometrical properties of these spaces and show how these properties impact various algorithms used in phylogenetic analyses. Our tree space is a discretisation of a previously introduced time tree space, called t-space, and hence our results can be used to approximate solutions to various open problems in t-space.
AbstractThe general theory developed by Ben Yaacov for metric structures provides Fraïssé limits which are approximately ultrahomogeneous. We show here that this result can be strengthened in the case of relational metric structures. We give an extra condition that guarantees exact ultrahomogenous limits. The condition is quite general. We apply it to stochastic processes, the class of diversities, and its subclass of $L_1$ diversities.
We describe a new and computationally efficient Bayesian methodology for inferring species trees and demographics from unlinked binary markers. Likelihood calculations are carried out using diffusion models of allele frequency dynamics combined with novel numerical algorithms. The diffusion approach allows for analysis of data sets containing hundreds or thousands of individuals. The method, which we call Snapper, has been implemented as part of the BEAST2 package. We conducted simulation experiments to assess numerical error, computational requirements, and accuracy recovering known model parameters. A reanalysis of soybean SNP data demonstrates that the models implemented in Snapp and Snapper can be difficult to distinguish in practice, a characteristic which we tested with further simulations. We demonstrate the scale of analysis possible using a SNP data set sampled from 399 fresh water turtles in 41 populations. [Bayesian inference; diffusion models; multi-species coalescent; SNP data; species trees; spectral methods.].
Hidden Markov models (HMMs) are general purpose models for time-series data widely used across the sciences because of their flexibility and elegance. Fitting HMMs can often be computationally demanding and time consuming, particularly when the number of hidden states is large or the Markov chain itself is long. Here we introduce a new Graphical Processing Unit (GPU)-based algorithm designed to fit long-chain HMMs, applying our approach to a model for low-frequency tremor events. Even on a modest GPU, our implementation resulted in an increase in speed of several orders of magnitude compared to the standard single processor algorithm. This permitted a full Bayesian inference of uncertainty related to model parameters and forecasts based on posterior predictive distributions. Similar improvements would be expected for HMM models given large number of observations and moderate state spaces (< 80 states with current hardware). We discuss the model, general GPU architecture and algorithms and report performance of the method on a tremor dataset from the Shikoku region, Japan. The new approach led to improvements in both computational performance and forecast accuracy, compared to existing frequentist methodology.
A bstract Microbial studies typically involve the sequencing and assembly of draft genomes for individual microbes or whole microbiomes. Given a draft genome, one first task is to determine its phylogenetic context, that is, to place it relative to the set of related reference genomes. We provide a new interactive graphical tool that addresses this task using Mash sketches to compare against all bacterial and archaeal representative genomes in the GTDB taxonomy, all within the framework of SplitsTree5. The phylogenetic context of the query sequences is then displayed as a phylogenetic outline, a new type of phylogenetic network that is more general that a phylogenetic tree, but significantly less complex than other types of phylogenetic networks. We propose to use such networks, rather than trees, to represent phylogenetic context, because they can express uncertainty in the placement of taxa, whereas a tree must always commit to a specific branching pattern. We illustrate the new method using a number of draft genomes of different assembly quality.
We define a set inner product to be a function on pairs of convex bodies which is symmetric, Minkowski linear in each dimension, positive definite, and satisfies the natural analogue of the Cauchy-Schwartz inequality (which is not implied by the other conditions). We show that any set inner product can be embedded into an inner product space on the associated support functions, thereby extending fundamental results of Hormander and Radstrom. The set inner product provides a geometry on the space of convex bodies. We explore some of the properties of that geometry, and discuss an application of these ideas to the reconstruction of ancestral ecological niches in evolutionary biology.
Computational inference of dated evolutionary histories relies upon various hypotheses about RNA, DNA, and protein sequence mutation rates. Using mutation rates to infer these dated histories is referred to as molecular clock assumption. Coalescent theory is a popular class of evolutionary models that implements the molecular clock hypothesis to facilitate computational inference of dated phylogenies. Cancer and virus evolution are two areas where these methods are particularly important. Methodologically, phylogenetic inference methods require a tree space over which the inference is performed, and geometry of this space plays an important role in statistical and computational aspects of tree inference algorithms. It has recently been shown that molecular clock, and hence coalescent, trees possess a unique geometry, different from that of classical phylogenetic tree spaces which do not model mutation rates. Here we introduce and study a space of discrete coalescent trees, that is, we assume that time is discrete, which is inevitable in many computational formalisations. We establish several geometrical properties of the space and show how these properties impact various algorithms used in phylogenetic analyses. Our tree space is a discretisation of a known time tree space, called t-space, and hence our results can be used to approximate solutions to various open problems in t-space. Our tree space is also a generalisation of another known trees space, called the ranked nearest neighbour interchange space, hence our advances in this paper imply new and generalise existing results about ranked trees.
Vincent Moulton合作论文数University of East Anglia;School of Computing Sciences8
Olivier Gascuel合作论文数Methodes et Algorithmes pour la Bioinformatique
LIRMM4
Rita Casadio合作论文数Bologna Biocomputing Unit3