Finitely generated modules over the polynomial ring in n indeterminates are isomorphic to quotients of finite rank free modules. We introduce a theory of relative Gröbner bases for those quotients of free modules and, equivalently, for pairs of submodules; we prove corresponding Buchberger- and Schreyer-type theorems. As applications of this theory, we consider three problems in persistence theory, which can be solved by relative Gröbner bases. First, we show that the relative Schreyer's theorem can be used to compute free presentations of complexes of finitely generated torsion-free modules. In contrast to previous approaches, this allows computation of free presentations for multicritical persistent homology directly at the chain module level without additional topological constructions. Second, any finitely generated Artinian module embeds in an Artinian injective hull, giving rise to a flat-injective presentation. We represent the embedding of the module in this injective hull by a quotient of a free module and apply the relative Schreyer's theorem to construct an algorithm for the computation of a free presentation from a flat-injective presentation. Third, we investigate how free presentations, and more generally free resolutions, obtained by the two preceding applications can be minimized by standard reduction techniques.
The Isometry Theorem of Chazal et al. and Lesnick is a fundamental result in persistence theory, which states that the interleaving distance between two one-parameter persistence modules is equal to the bottleneck distance between their barcodes. Significant effort has been devoted to extending this result to modules defined over more general posets. As these modules do not generally admit nice decompositions, one must restrict attention to the class of interval-decomposable modules in order to define an appropriate notion of bottleneck distance. Even with this assumption, it is known that bottleneck distance may not be equivalent to interleaving distance, but that it is Lipschitz stable under certain, fairly restrictive, assumptions. In this paper, we consider the more basic question of stability of the Hausdorff distance with respect to interleaving distance for interval-decomposable modules. Our main theorem is a Lipschitz stability result, which holds in a fairly general setting of interval-decomposable modules over arbitrary posets, where intervals are assumed to be taken from any family satisfying certain closure conditions. Along the way, we develop some new tools and results for interval-decomposable modules over arbitrary posets, in the form of geometrically-flavored characterizations of the existence of morphisms and interleavings between interval modules.
Flat-injective presentations were introduced by Miller (2020) to combinatorially describe Zn-graded modules. We consider them for local graded rings R, with grading over any abelian group, and give a criterion for minimality of them. In the special case of the polynomial ring with Zn-grading, this criterion reduces to a family of k-linear equations, and we are able to give an algorithmic procedure for reduction. Furthermore, we provide the description of a flat-injective presentation, which can be constructed from the structure maps of a given finitely generated R-module. This solves the construction problem for flat-injective presentations under strong finiteness assumptions.
In real world, mutations of genetic sequences are often accompanied by their recombinations. Such joint phenomena are modeled by phylogenetic networks. Nakkleh formulated the phylogenetic network reconstruction problem (PNRP) as follows: Given a family of phylogenetic trees over a common set of taxa, is there a unique minimal phylogenetic network whose set of spanning trees contains the family? There are different answers to PNRP, since there are different ways to define what a minimal network is (based on different optimization criteria). Inspired by ideas from topological data analysis (TDA), we devise lattice-diagram models for the visualization of phylogenetic networks and of filtrations, called the cliquegram and the facegram, respectively, both generalizing the dendrogram model of phylogenetic trees. Both models allow us to solve the PNRP in a rigorous way and free of choosing optimization criteria. The solution to the phylogenetic network and filtration reconstruction process is obtained by taking the join operation of the dendrograms on the lattice of cliquegrams, and of facegrams, respectively. Furthermore, we show that computing the join-facegram from a given set of dendrograms is polynomial in the size and number of the input trees. We propose two novel invariants of facegrams, (i) the face-Reeb graph and (ii) the mergegram of a facegram. We show the mergegram is 1-Lipschitz stable, while the face-Reeb graph is not. In particular, we show that the mergegram is invariant of weak equivalences of filtrations (a stronger form of homotopy equivalence). This new TDA-signature, the mergegram, can be used as a computable proxy for phylogenetic networks and also, more broadly, for filtrations of datasets, which might be of independent interest to TDA. To illustrate the utility of those new TDA-tools to phylogenetics, we provide experiments with artificial and benchmark biological data.
In this paper, we introduce the persistence transformation, a novel methodology in Topological Data Analysis (TDA) for applications in time series data which can be obtained in various areas such as science, politics, economy, healthcare, engineering, and beyond. This approach captures the enduring presence or `persistence' of signal peaks in time series data arising from Morse functions while preserving their positional information. Through rigorous analysis, we demonstrate that the proposed persistence transformation exhibits stability and outperforms the persistent diagram of Morse functions (with respect to filtration, e.g., the upper levelset filtration). Moreover, we present a modified version of the persistence transformation, termed the reduced persistence transformation, which retains stability while enjoying dimensionality reduction in the data. Consequently, the reduced persistence transformation yields faster computational results for subsequent tasks, such as classification, albeit at the cost of reduced overall accuracy compared to the persistence transformation. However, the reduced persistence transformation finds relevance in specific domains, e.g., MALDI-Imaging, where positional information is of greater significance than the overall signal height. Finally, we provide a conceptual outline for extending the persistence diagram to accommodate higher-dimensional input while assessing its stability under these modifications.
BACKGROUND:Matrix-assisted laser desorption/ionization mass spectrometry imaging (MALDI MSI) displays significant potential for applications in cancer research, especially in tumor typing and subtyping. Lung cancer is the primary cause of tumor-related deaths, where the most lethal entities are adenocarcinoma (ADC) and squamous cell carcinoma (SqCC). Distinguishing between these two common subtypes is crucial for therapy decisions and successful patient management.RESULTS:We propose a new algebraic topological framework, which obtains intrinsic information from MALDI data and transforms it to reflect topological persistence. Our framework offers two main advantages. Firstly, topological persistence aids in distinguishing the signal from noise. Secondly, it compresses the MALDI data, saving storage space and optimizes computational time for subsequent classification tasks. We present an algorithm that efficiently implements our topological framework, relying on a single tuning parameter. Afterwards, logistic regression and random forest classifiers are employed on the extracted persistence features, thereby accomplishing an automated tumor (sub-)typing process. To demonstrate the competitiveness of our proposed framework, we conduct experiments on a real-world MALDI dataset using cross-validation. Furthermore, we showcase the effectiveness of the single denoising parameter by evaluating its performance on synthetic MALDI images with varying levels of noise.CONCLUSION:Our empirical experiments demonstrate that the proposed algebraic topological framework successfully captures and leverages the intrinsic spectral information from MALDI data, leading to competitive results in classifying lung cancer subtypes. Moreover, the framework's ability to be fine-tuned for denoising highlights its versatility and potential for enhancing data analysis in MALDI applications.
Given a multiparameter filtration of simplicial complexes, we consider the problem of explicitly constructing generators for the multipersistent homology groups with arbitrary PID coefficients. We propose the use of spanning trees as a tool to identify such generators by introducing a condition for persistent spanning trees, which is accompanied by an existence result for cofiltrations consisting of spanning trees. We also introduce a generalization of spanning trees, called spanning complexes, for dimensions higher than one, and we establish their existence as a first step towards this direction.
One-dimensional persistent homology is arguably the most important and heavily used computational tool in topological data analysis. Additional information can be extracted from datasets by studying multi-dimensional persistence modules and by utilizing cohomological ideas, e.g. the cohomological cup product. In this work, given a single parameter filtration, we investigate a certain 2-dimensional persistence module structure associated with persistent cohomology, where one parameter is the cup-length ℓ≥ 0 and the other is the filtration parameter. This new persistence structure, called the persistent cup module , is induced by the cohomological cup product and adapted to the persistence setting. Furthermore, we show that this persistence structure is stable. By fixing the cup-length parameter ℓ , we obtain a 1-dimensional persistence module, called the persistent ℓ -cup module, and again show it is stable in the interleaving distance sense, and study their associated generalized persistence diagrams. In addition, we consider a generalized notion of a persistent invariant , which extends both the rank invariant (also referred to as persistent Betti number ), Puuska’s rank invariant induced by epi-mono-preserving invariants of abelian categories, and the recently-defined persistent cup-length invariant , and we establish their stability. This generalized notion of persistent invariant also enables us to lift the Lyusternik-Schnirelmann (LS) category of topological spaces to a novel stable persistent invariant of filtrations, called the persistent LS-category invariant .
Metrics of interest in topological data analysis (TDA) are often explicitly or implicitly in the form of an interleaving distance $d_{\mathrm{I}}$ between poset maps (i.e. order-preserving maps), e.g. the Gromov-Hausdorff distance between metric spaces can be reformulated in this way. We propose a representation of a poset map $\mathbf{F}:\mathcal{P}\to\mathcal{Q}$ as a join (i.e. supremum) $\bigvee_{b\in B} \mathbf{F}_b$ of simpler poset maps $\mathbf{F}_b$ (for a join dense subset $B\subset \mathcal{Q}$) which in turn yields a decomposition of $d_{\mathrm{I}}$ into a product metric. The decomposition of $d_{\mathrm{I}}$ is simple, but its ramifications are manifold: (1) We can construct a geodesic path between any poset maps $\mathbf{F}$ and $\mathbf{G}$ with $d_{\mathrm{I}}(\mathbf{F},\mathbf{G})<\infty$ by assembling geodesics between all $\mathbf{F}_b$s and $\mathbf{G}_b$s via the join operation. This construction generalizes at least three constructions of geodesic paths that have appeared in the literature. (2) We can extend the Gromov-Hausdorff distance to a distance between simplicial filtrations over an arbitrary poset with a flow, preserving its universality and geodesicity. (3) We can clarify equivalence between several known metrics on multiparameter hierarchical clusterings. (4) We can illuminate the relationship between the erosion distance by Patel and the graded rank function by Betthauser, Bubenik, and Edwards, which in turn takes us to an interpretation on the representation $\bigvee_b \mathbf{F}_b$ as a generalization of persistence landscapes and graded rank functions.
Let X be a closed subspace of a metric space M. It is well known that, under mild hypotheses, one can estimate the Betti numbers of X from a finite set $$P\subset M$$ of points approximating X. In this paper, we show that one can also use P to estimate much more detailed topological properties of X. We achieve this by proving the stability of $$A_\infty $$ -persistent homology. In its most general case, this stability means that given a continuous function $$f:Y\rightarrow {\mathbb {R}}$$ on a topological space Y, small perturbations in the function f imply at most small perturbations in the family of $$A_\infty $$ -barcodes. This work can be viewed as a proof of the stability of cup-product and generalized-Massey-products persistence. The technical key of this paper consists of figuring out a setting which makes $$A_\infty $$ -persistence functorial.
Cohomological ideas have recently been injected into persistent homology and have for example been used for accelerating the calculation of persistence diagrams by the software Ripser. The cup product operation which is available at cohomology level gives rise to a graded ring structure that extends the usual vector space structure and is therefore able to extract and encode additional rich information. The maximum number of cocycles having non-zero cup product yields an invariant, the cup-length, which is useful for discriminating spaces. In this paper, we lift the cup-length into the persistent cup-length function for the purpose of capturing ring-theoretic information about the evolution of the cohomology (ring) structure across a filtration. We show that the persistent cup-length function can be computed from a family of representative cocycles and devise a polynomial time algorithm for its computation. We furthermore show that this invariant is stable under suitable interleaving-type distances.
Cohomological ideas have recently been injected into persistent homology and have been utilized for both enriching and accelerating the calculation of persistence diagrams. For instance, the software Ripser fundamentally exploits the computational advantages offered by cohomological ideas. The cup product operation which is available at cohomology level gives rise to a graded ring structure which extends the natural vector space structure and is therefore able to extract and encode additional rich information. The maximum number of cocycles having non-zero cup product yields an invariant, the Cup-Length, which is efficient at discriminating spaces. In this paper, we lift the cup-length into the Persistent Cup-Length invariant for the purpose of extracting non-trivial information about the evolution of the cohomology ring structure across a filtration. We show that the Persistent Cup-Length can be computed from a family of representative cocycles and devise a polynomial time algorithm for the computation of the Persistent Cup-Length invariant. We furthermore show that this invariant is stable under suitable interleaving-type distances. Along the way, we identify an invariant which we call the Cup-Length Diagram, which is stronger than persistent cuplength but can still be computed efficiently. In addition, by considering the l-fold product of persistent cohomology rings, we identify certain persistence modules, which are also stable and can be used to evaluate the persistent cup-length.
We define a persistent cohomology invariant called persistent cup-length which is able to extract non trivial information about the evolution of the cohomology ring structure across a filtration. We also devise algorithms for the computation of this invariant and we furthermore show that the persistent cup-length is 2-Lipschitz continuous with respect to the homotopy interleaving and Gromov-Hausdorff distances.
Inspired by the interval decomposition of persistence modules and the extended Newick format of phylogenetic networks, we show that, inside the larger category of partially ordered Reeb graphs , every Reeb graph with n leaves and first Betti number s , can be identified with a coproduct of at most 2^s partially ordered trees with (n + s) leaves. Reeb graphs are therefore classified up to isomorphism by their tree-decomposition. An implication of this result, is that the isomorphism problem for Reeb graphs is fixed parameter tractable when the parameter is the first Betti number. We propose partially ordered Reeb graphs as a model for time consistent phylogenetic networks and propose a certain Hausdorff distance as a metric on these structures.
\textit{Formigrams} are a natural generalization of the notion of \textit{dendrograms}. This notion has recently been proposed as a signature for studying the evolution of clusters in dynamic datasets across different time scales. Although its formulation is set-theoretic, the notion of formigram is deeply related to certain algebraic-topological methods used in \textit{topological data analysis}, such as \textit{Reeb graphs} and \textit{zigzag persistence modules}. In this paper we give a self-contained study of the algebraic structure of formigrams and their interleaving distance. For a finite set $X$, we define a partial order on the collection of all formigrams and we show that every formigram over $X$ has a canonical decomposition into a join of simpler formigrams. This is analogous to the decomposition of persistence modules into direct sums of interval modules. Furthermore, we show that the interleaving distance between formigrams decomposes into a product metric of the interleaving distance between certain pre-cosheaves. This is analogous to the celebrated \textit{interleaving-bottleneck isometry theorem} for persistence modules.
Metrics in computational topology are often either (i) themselves in the form of the interleaving distance $d_{\mathrm{I}}(\mathbf{F},\mathbf{G})$ between certain order-preserving maps $\mathbf{F},\mathbf{G}:(\mathcal{P},\leq)\rightarrow (\mathcal{Q},\leq)$ between posets or (ii) admit $d_{\mathrm{I}}(\mathbf{F},\mathbf{G})$ as a tractable lower bound, where the domain poset $(\mathcal{P},\leq)$ is equipped with a flow. In this paper, assuming that $\mathcal{Q}$ admits a join-dense subset $B$, we propose certain join representations $\mathbf{F}=\bigvee_{b\in B} \mathbf{F}_b$ and $\mathbf{G}=\bigvee_{b\in B} \mathbf{G}_b$ which satisfy $d_{\mathrm{I}}(\mathbf{F},\mathbf{G})=\bigvee_{b\in B} d_{\mathrm{I}}(\mathbf{F}_b,\mathbf{G}_b)$ where each $d_{\mathrm{I}}(\mathbf{F}_b,\mathbf{G}_b)$ is relatively easy to compute. We leverage this result in order to (i) elucidate the structure and computational complexity of the interleaving distance for poset-indexed clusterings (i.e. poset-indexed subpartition-valued functors), (ii) to clarify the relationship between the erosion distance by Patel and the graded rank function by Betthauser, Bubenik, and Edwards, and (iii) to reformulate and generalize the tripod distance by the second author.
There are many metrics available to compare phylogenetic trees since this is a fundamental task in computational biology. In this paper, we focus on one such metric, the ℓ^∞-cophenetic metric introduced by Cardona et al. This metric works by representing a phylogenetic tree with n labeled leaves as a point in ℝ^n(n+1)/2 known as the cophenetic vector, then comparing the two resulting Euclidean points using the ℓ^∞ distance. Meanwhile, the interleaving distance is a formal categorical construction generalized from the definition of Chazal et al., originally introduced to compare persistence modules arising from the field of topological data analysis. We show that the ℓ^∞-cophenetic metric is an example of an interleaving distance. To do this, we define phylogenetic trees as a category of merge trees with some additional structure; namely labelings on the leaves plus a requirement that morphisms respect these labels. Then we can use the definition of a flow on this category to give an interleaving distance. Finally, we show that, because of the additional structure given by the categories defined, the map sending a labeled merge tree to the cophenetic vector is, in fact, an isometric embedding, thus proving that the ℓ^∞-cophenetic metric is, in fact, an interleaving distance.
The interleaving distance was originally defined in the field of Topological Data Analysis (TDA) by Chazal et al. as a metric on the class of persistence modules parametrized over the real line. Bubenik et al. subsequently extended the definition to categories of functors on a poset, the objects in these categories being regarded as `generalized persistence modules'. These metrics typically depend on the choice of a lax semigroup of endomorphisms of the poset. The purpose of the present paper is to develop a more general framework for the notion of interleaving distance using the theory of `actegories'. Specifically, we extend the notion of interleaving distance to arbitrary categories equipped with a flow, i.e. a lax monoidal action by the monoid $[0,\infty)$. In this way, the class of objects in such a category acquires the structure of a Lawvere metric space. Functors that are colax $[0,\infty)$-equivariant yield maps that are $1$-Lipschitz. This leads to concise proofs of various known stability results from TDA, by considering appropriate colax $[0,\infty)$-equivariant functors. Along the way, we show that several common metrics, including the Hausdorff distance and the $L^{\infty}$-norm, can be realized as interleaving distances in this general perspective.