Phylogenetic networks provide a means of describing the evolutionary history of sets of species believed to have undergone hybridization or gene flow during their evolution. The mutation process for a set of such species can be modeled as a Markov process on a phylogenetic network. Previous work has shown that a site-pattern probability distributions from a Jukes-Cantor phylogenetic network model must satisfy certain algebraic invariants. As a corollary, aspects of the phylogenetic network are theoretically identifiable from site-pattern frequencies. In practice, because of the probabilistic nature of sequence evolution, the phylogenetic network invariants will rarely be satisfied, even for data generated under the model. Thus, using network invariants for inferring phylogenetic networks requires some means of interpreting the residuals, or deviations from zero, when observed site-pattern frequencies are substituted into the invariants. In this work, we propose a method of utilizing invariant residuals and support vector machines to infer 4-leaf level-one phylogenetic networks, from which larger networks can be reconstructed. Given data for a set of species, the support vector machine is first trained on model data to learn the patterns of residuals corresponding to different network structures to classify the network that produced the data. We demonstrate the performance of our method on simulated data from the specified model and primate data.
Inspired by the work of Amdeberhan, Can, and Moll on broken necklaces, we define a broken bracelet as a linear arrangement of marked and unmarked vertices and introduce a generalization called n-stars, which is a collection of n broken bracelets whose final (unmarked) vertices are identified. Through these combinatorial objects, we provide a new framework for the study of Kostant's partition function, which counts the number of ways to express a vector as a nonnegative integer linear combination of the positive roots of a Lie algebra. Our main result establishes that (up to reflection) the number of broken bracelets with a fixed number of unmarked vertices with nonconsecutive marked vertices gives an upper bound for the value of Kostant's partition function for multiples of the highest root of a Lie algebra of type A. We connect this work to multiplex juggling sequences, as studied by Benedetti, Hanusa, Harris, Morales, and Simpson, by providing a correspondence to an equivalence relation on n-stars.
An important consideration for a model-based method of phylogenetic network inference is the identifiability of the network parameter of the model. A recurring theme in previous works exploring this issue is that it is often difficult to identify the orientation of edges in a triangle of the network. In fact, it has been shown that for some models it is impossible to determine the orientation of triangle edges utilizing the standard algebraic technique of phylogenetic invariants. In this work, we consider one such model with a Jukes-Cantor site-substitution process and no coalescence. We give a complete semialgebraic description of three, 3-leaf Jukes-Cantor phylogenetic network models with embedded triangles. By describing these base cases, we resolve several questions about the identifiability of networks with embedded triangles. We show that for any pair of models, the intersection and set differences of the models are full-dimensional regions of the space of site-pattern probability distributions. Thus, despite being algebraically indistinguishable, these network models are not identical, nor are they identifiable (or generically identifiable). Our results also yield a straightforward biological interpretation–that the signal from a hybridization event may be immediately detectable but decays over time until it is impossible to identify the orientation of edges in the triangle of a network.
Changes in environmental or system parameters often drive major biological transitions, including ecosystem collapse, disease outbreaks, and tumor development. Analyzing the stability of steady states in dynamical systems provides critical insight into these transitions. This paper introduces an algebraic framework for analyzing the stability landscapes of ecological models defined by systems of first-order autonomous ordinary differential equations with polynomial or rational rate functions. Using tools from real algebraic geometry, we characterize parameter regions associated with steady-state feasibility and stability via three key boundaries: singular, stability (Routh-Hurwitz), and coordinate boundaries. With these boundaries in mind, we employ routing functions to compute the connected components of parameter space in which the number and type of stable steady states remain constant, revealing the stability landscape of these ecological models. As case studies, we revisit the classical Levins-Culver competition-colonization model and a recent model of coral-bacteria symbioses. In the latter, our method uncovers complex stability regimes, including regions supporting limit cycles, that are inaccessible via traditional techniques. These results demonstrate the potential of our approach to inform ecological theory and intervention strategies in systems with nonlinear interactions and multiple stable states.
Phylogenetic networks describe the evolution of a set of taxa for which reticulate events have occurred at some point in their evolutionary history. Of particular interest is when the evolutionary history between a set of just three taxa has a reticulate event. In molecular phylogenetics, substitution models can model the process of evolution at the genetic level, and the case of three taxa with a reticulate event can be modeled using a substitution model on a semi-directed graph called a 3-sunlet. We investigate a class of substitution models called group-based phylogenetic models on 3-sunlet networks. In particular, we investigate the discrete geometry of the parameter space and how this relates to the dimension of the phylogenetic variety associated to the model. This enables us to give a dimension formula for this variety for general group-based models when the order of the group is odd.
Algebraic techniques in phylogenetics have historically been successful at proving identifiability results and have also led to novel reconstruction algorithms. In this paper, we study the ideal of phylogenetic invariants of the Cavender-Farris-Neyman (CFN) model on a phylogenetic network with the goal of providing a description of the invariants which is useful for network inference. It was previously shown that to characterize the invariants of any level-1 network, it suffices to understand all sunlet networks, which are those consisting of a single cycle with a leaf adjacent to each cycle vertex. We show that the parameterization of an affine open patch of the CFN sunlet model, which intersects the probability simplex, factors through the space of skew-symmetric matrices via Pfaffians. We then show that this affine patch is isomorphic to a determinantal variety and give an explicit Gröbner basis for the associated ideal, which involves only n2 coordinates rather than 2^n. Lastly, we show that sunlet networks with at least 6 leaves are identifiable using only these polynomials and run extensive simulations, which show that these polynomials can be used to accurately infer the correct network from DNA sequence data.
Log-linear exponential random graph models are a specific class of statistical network models that have a log-linear representation. This class includes many stochastic blockmodel variants. In this paper, we focus on β-stochastic blockmodels, which combine the β-model with a stochastic blockmodel. Here, using recent results by Almendra-Hernández, De Loera, and Petrović, which describe a Markov basis for β-stochastic block model, we give a closed form formula for the maximum likelihood degree of a β-stochastic blockmodel. The maximum likelihood degree is the number of complex solutions to the likelihood equations. In the case of the β-stochastic blockmodel, the maximum likelihood degree factors into a product of Eulerian numbers.
Watanabe's singular learning theory provides a framework for asymptotic analysis of Bayesian model selection for statistical models with singularities, where traditional statistical regularity assumptions fail. Learning coefficients, also known as real log canonical thresholds, play a central role in singular learning, as they govern the asymptotic behavior of Bayesian marginal likelihood integrals in settings where the Laplace approximations used for regular statistical models are not applicable. Learning coefficients are algebraic invariants that quantify the geometric complexity of a model and reveal how the singular structure impacts the model's generalization properties. In this paper, we apply algebraic methods to study the learning coefficients of factor analysis models, which are widely used latent variable models for continuously distributed data. Our main result provides exact formulas for learning coefficients of factor analysis models. Moreover, we study the singularity types of specific factor analysis models in detail.
We consider the fundamental question of which evolutionary histories can potentially be reconstructed from sufficiently long DNA sequences, by studying the identifiability of phylogenetic networks from data generated under Markov models of DNA evolution. This topic has previously been studied for phylogenetic trees and for phylogenetic networks that are level-1, which means that reticulate evolutionary events were restricted to be independent in the sense that the corresponding cycles in the network are non-overlapping. In this paper, we study the identifiability of phylogenetic networks from DNA sequence data under Markov models of DNA evolution for more general classes of networks that may contain pairs of tangled reticulations. Our main result is generic identifiability, under the Jukes-Cantor model, of binary semi-directed level-2 phylogenetic networks that satisfy two additional conditions called triangle-free and strongly tree-child. We also consider level-1 networks and show stronger identifiability results for this class than what was known previously. In particular, we show that the number of reticulations in a level-1 network is identifiable under the Jukes-Cantor model. Moreover, we prove general identifiability results that do not restrict the network level at all and hold for the Jukes-Cantor as well as for the Kimura-2-Parameter model. We show that any two binary semi-directed phylogenetic networks are distinguishable if they do not display exactly the same 4-leaf subtrees, called quartets. This has direct consequences regarding the blobs of a network, which are its reticulated components. We show that the tree-of-blobs of a network, the global branching structure of the network, is always identifiable, as well as the circular ordering of the subnetworks around each blob, for networks in which edges do not cross and taxa are on the outside. ### Competing Interest Statement The authors have declared no competing interest. National Science Foundation, https://ror.org/021nxhr62, DMS-1929284, DMS-1945584 Dutch Research Council, , OCENW.M.21.306, OCENW.KLEIN.125 Wisconsin Alumni Research Foundation, , Georgia Benkart Legacy Fund, ,
The steady-state degree of a chemical reaction network is the number of complex steady-states for generic rate constants and initial conditions. One way to bound the steady-state degree is through the mixed volume of the steady-state system or an equivalent system. In this work, we show that for partionable binomial networks, whose resulting steady-state systems are given by a set of binomials and a set of linear (not necessarily binomial) conservation equations, computing the mixed volume is equivalent to finding the volume of a single mixed cell that is the translate of a parallelotope. We then turn our attention to identifying cycles with binomial steady-state ideals. To this end, we give a coloring condition on directed cycles that guarantees the network has a binomial steady-state ideal. We highlight both of these theorems using a class of networks referred to as species-overlapping networks and give a formula for the mixed volume of these networks.
Recently, Sturma, Drton, and Leung proposed a general-purpose stochastic method for hypothesis testing in models defined by polynomial equality and inequality constraints. Notably, the method remains theoretically valid even near irregular points, such as singularities and boundaries, where traditional testing approaches often break down. In this paper, we evaluate its practical performance on a collection of biologically motivated models from phylogenetics. While the method performs remarkably well across different settings, we catalogue a number of issues that should be considered for effective application.
Motivated by the question of how biological systems maintain homeostasis in changing environments, Shinar and Feinberg introduced in 2010 the concept of absolute concentration robustness (ACR). A biochemical system exhibits ACR in some species if the steady-state value of that species does not depend on initial conditions. Thus, a system with ACR can maintain a constant level of one species even as the environment changes. Despite a great deal of interest in ACR in recent years, the following basic question remains open: How can we determine quickly whether a given biochemical system has ACR? Although various approaches to this problem have been proposed, we show that they are incomplete. Accordingly, we present new methods for deciding ACR, which harness computational algebra. We illustrate our results on several biochemical signaling networks.
Phylogenetic networks represent evolutionary histories of sets of taxa where horizontal evolution or hybridization has occurred. Placing a Markov model of evolution on a phylogenetic network gives a model that is particularly amenable to algebraic study by representing it as an algebraic variety. In this paper, we give a formula for the dimension of the variety corresponding to a triangle-free level-1 phylogenetic network under a group-based evolutionary model. On our way to this, we give a dimension formula for codimension zero toric fiber products. We conclude by illustrating applications to identifiability.
In the last quarter of a century, algebraic statistics has established itself as an expanding field which uses multilinear algebra, commutative algebra, computational algebra, geometry, and combinatorics to tackle problems in mathematical statistics. These developments have found applications in a growing number of areas, including biology, neuroscience, economics, and social sciences. Naturally, new connections continue to be made with other areas of mathematics and statistics. This paper outlines three such connections: to statistical models used in educational testing, to a classification problem for a family of nonparametric regression models, and to phase transition phenomena under uniform sampling of contingency tables. We illustrate the motivating problems, each of which is for algebraic statistics a new direction, and demonstrate an enhancement of related methodologies.
We discuss the role computational algebraic geometry and symbolic computation has played in regards to statistical problems related to inferring phylogenetic networks with a focus on identifiability. This article is an accompanying extended abstract of my talk at ISSAC 2024. The article summarizes work completed with several co-authors, including: Leo van Iersel, Remie Janssen, Mark Jones, Robert Krone, Colby Long, Samuel Martin, and Yukihiro Murakami.
Reconstructing the network of life from molecular data is a complicated task. How can computational algebraic geometry play a role?
Network-based modeling of complex systems and data using the language of graphs has become an essential topic across a range of different disciplines. Arguably, this graph-based perspective derives its success from the relative simplicity of graphs: A graph consists of nothing more than a set of vertices and a set of edges, describing relationships between pairs of such vertices. This simple combinatorial structure makes graphs interpretable and flexible modeling tools. The simplicity of graphs as system models, however, has been scrutinized in the literature recently. Specifically, it has been argued from a variety of different angles that there is a need for higher-order networks, which go beyond the paradigm of modeling pairwise relationships, as encapsulated by graphs. In this survey article we take stock of these recent developments. Our goals are to clarify (i) what higher-order networks are, (ii) why these are interesting objects of study, and (iii) how they can be used in applications.
A foundational question in the theory of linear compartmental models is how to assess whether a model is structurally identifiable – that is, whether parameter values can be inferred from noiseless data – directly from the combinatorics of the model. Our main result completely answers this question for models (with one input and one output) in which the underlying graph is a bidirectional tree; moreover, identifiability of such models can be verified visually. Models of this structure include two families of models often appearing in biological applications: catenary and mammillary models. Our analysis of such models is enabled by two supporting results, which are significant in their own right. One result gives the first general formula for the coefficients of input-output equations (certain equations that can be used to determine identifiability) that allows for input and output to be in distinct compartments. In another supporting result, we prove that identifiability is preserved when a model is enlarged and altered in specific ways involving adding a new compartment with a bidirected edge to an existing compartment.
Jan Verschelde合作论文数University of Illinois at Chicago; Statistics and Computer Science ;Department of Mathematics2