We provide a bottom up construction of torsion generators for weighted homology of a weighted complex over a discrete valuation ring $R=\mathbb{F}[[\pi]]$. This is achieved by starting from a basis for classical homology of the $n$-th skeleton for the underlying complex with coefficients in the residue field $\mathbb{F}$ and then lifting it to a basis for the weighted homology with coefficients in the ring $R$. Using the latter, a bijection is established between $n+1$ and $n$ dimensional simplices whose weight ratios provide the exponents of the $\pi$-monomials that generate each torsion summand in the structure theorem of the weighted homology modules over $R$. We present algorithms that subsume the torsion computation by reducing it to normalization over the residue field of $R$, and describe a Python package we implemented that takes advantage of this reduction and performs the computation efficiently.
We develop a framework for computing the homology of weighted simplicial complexes with coefficients in a discrete valuation ring. A weighted simplicial complex, $(X,v)$, introduced by Dawson [Cah. Topol. Géom. Différ. Catég. 31 (1990), pp. 229--243], is a simplicial complex, $X$, together with an integer-valued function, $v$, assigning weights to simplices, such that the weight of any of faces are monotonously increasing. In addition, weighted homology, $H_n^v(X)$, features a new boundary operator, $\partial_n^v$. In difference to Dawson, our approach is centered at a natural homomorphism $\theta$ of weighted chain complexes. The key object is $H^v_{n}(X/\theta)$, the weighted homology of a quotient of chain complexes induced by $\theta$, appearing in a long exact sequence linking weighted homologies with different weights. We shall construct bases for the kernel and image of the weighted boundary map, identifying $n$-simplices as either $\kappa_n$- or $\mu_n$-vertices. Long exact sequences of weighted homology groups and the bases, allow us to prove a structure theorem for the weighted simplicial homology with coefficients in a ring of formal power series $R=\mathbb{F}[[\pi]]$, where $\mathbb{F}$ is a field. Relative to simplicial homology new torsion arises and we shall show that the torsion modules are connected to a pairing between distinguished $\kappa_n$ and $\mu_{n+1}$ simplices.
We introduce weighted simplicial cohomology with coefficients in Q[pi]. This cohomology is derived from a weighted coboundary operator incorporating simplex weights from Q[pi]. We provide a particular bipartition for the set of n-simplices into mu n and kappa n-simplices and then we establish a four-term exact sequence relating the torsion module of weighted cohomology with regular cohomology having coefficients in Q[pi]. Furthermore, we prove a structure theorem for the torsion, expressing its invariant factors as ratios of weights of distinguished (mu n-1, kappa n)-simplex-pairs. We then employ this result to interpret the long homology sequence arising from a natural map connecting weighted and regular cohomology over Q[pi]. Secondly we leverage weighted homology by a bipartition into mu n-and kappa n-simplices with its torsion expressed via a pairing (kappa n, mu n-1). We show that cohomological torsion is described by a pairing of the form (mu n-1, mu n), which gives rise to an isomorphism between weighted cohomological torsion and Hom(Im partial derivative nv,R).
Mobile genetic elements are key to the global emergence of antibiotic resistance. We successfully reconstructed the complete bacterial genome and plasmid assemblies of isolates sharing the same blaKPC carbapenemase gene to understand evolution over time in six confined hospital drains over five years. From 82 isolates we identified 14 unique strains from 10 species with 113 blaKPC-carrying plasmids across 16 distinct replicon types. To assess dynamic gene movement, we introduced the ‘Composite-Sample Complex’, a novel mathematical approach to using probability to capture the directional movement of antimicrobial resistance genes. The Composite Sample Complex accounts for the co-occurrence of both plasmids and chromosomes within an isolate, and highlights likely gene donors and recipients. From the validated model, we demonstrate frequent transposition events of blaKPC from plasmids to other plasmids, as well as integration into the bacterial chromosome within specific drains. We present a novel approach to estimate the directional movement of antimicrobial resistance via gene mobilization.
Enumerative studies of RNA secondary structures were initiated four decades ago by Waterman and his coworkers. Since then, RNA secondary structures have been explored according to many different structural characteristics, for instance, helices, components and loops by Hofacker, Schuster and Stadler, orders by Nebel, saturated structures by Clote, the 5^'-3^' end distance by Clote, Ponty and Steyaert, and the rainbow spectrum by Li and Reidys. However, the majority of the contributions are asymptotic results, and it is harder to derive explicit formulas. In this paper, we obtain exact formulas counting RNA secondary structures with a given number of helices as well as a given joint size distribution of helices and loops, while some related asymptotic results due to Hofacker, Schuster and Stadler have been known for about twenty years. Our approach is combinatorial, analyzing a recent bijection between RNA secondary structures and plane trees discovered by the first author and proposing a variation of Chen's bijective approach of counting trees by forests of simple trees.
In this article we develop an information-theoretic framework of multiple sequence alignments (MSAs), based on sub-sampling. The key component of this framework is an information-theoretical potential defined on pairs of sites (links) within the MSA. This potential quantifies the expected drop in variation of information between the two constituent sites. The expectation is taken with respect to all possible sub-alignments, obtained by removing a finite, fixed number of rows. We show that the potential is zero for linked sites representing columns, for which symbols are in bijective correspondence and that it is strictly positive, otherwise. It is furthermore shown that the potential assumes its unique minimum for links at which each symbol pair appears with the same multiplicity. We then show that the established drop of the variation of information exceeds finite-size effects inherent to the construction of the potential. Finally, we provide as a proof of concept an application of our results to a specific MSA composed of the inverse fold solutions of three distinguished secondary structures.
A new perspective is introduced regarding the analysis of Multiple Sequence Alignments (MSA), representing aligned data defined over a finite alphabet of symbols. The framework is designed to produce a block decomposition of an MSA, where each block is comprised of sequences exhibiting a certain site-coherence. The key component of this framework is an information theoretical potential defined on pairs of sites (links) within the MSA. This potential quantifies the expected drop in variation of information between the two constituent sites, where the expectation is taken with respect to all possible sub-alignments, obtained by removing a finite, fixed collection of rows. It is proved that the potential is zero for linked sites representing columns, whose symbols are in bijective correspondence and it is strictly positive, otherwise. It is furthermore shown that the potential assumes its unique minimum for links at which each symbol pair appears with the same multiplicity. Finally, an application is presented regarding anomaly detection in an MSA, composed of inverse fold solutions of a fixed tRNA secondary structure, where the anomalies are represented by inverse fold solutions of a different RNA structure.
Mobile genetic elements are key to the global emergence of antibiotic resistance. We successfully reconstructed the complete bacterial genome and plasmid assemblies of isolates sharing the same blaKPC carbapenemase gene to understand evolution over time in six confined hospital drain biofilms over five years. From 82 isolates we identified 14 unique strains from 10 species with 113 blaKPC−carrying plasmids across 16 distinct replicon types. To assess dynamic gene movement, we introduced the 'Composite-Sample Complex', a novel mathematical approach to using probability to capture the directional movement of antimicrobial resistance genes accounting for the co-occurrence of both plasmids and chromosomes within an isolate, and highlighting likely donors and recipients. From the validated model, we demonstrate frequent transposition events of blaKPC from plasmids to other plasmids, as well as integration into the bacterial chromosome within specific drain biofilms. We present a novel approach to estimate the directional movement of antimicrobial resistance via gene mobilization.
We present a novel framework enhancing the prediction of whether novel lineage poses the threat of eventually dominating the viral population. The framework is based purely on genomic sequence data, without requiring prior established biological analysis. Its building blocks are sets of coevolving sites in the alignment (motifs), identified via coevolutionary signals. The collection of such motifs forms a relational structure over the polymorphic sites. Motifs are constructed using distances quantifying the coevolutionary coupling of pairs and manifest as coevolving clusters of sites. We present an approach to genomic surveillance based on this notion of relational structure. Our system will issue an alert regarding a lineage, based on its contribution to drastic changes in the relational structure. We then conduct a comprehensive retrospective analysis of the COVID-19 pandemic based on SARS-CoV-2 genomic sequence data in GISAID from October 2020 to September 2022, across 21 lineages and 27 countries with weekly resolution. We investigate the performance of this surveillance system in terms of its accuracy, timeliness, and robustness. Lastly, we study how well each lineage is classified by such a system.
The study of native motifs of RNA secondary structures helps us better understand the formation and eventually the functions of these molecules. Commonly known structural motifs include helices, hairpin loops, bulges, interior loops, exterior loops and multiloops. However, enumerative results and generating algorithms taking into account the joint distribution of these motifs are sparse. In this paper, we present progress on deriving such distributions employing a tree-bijection of RNA secondary structures obtained by Schmitt and Waterman and a novel rake decomposition of plane trees. The key feature of the latter is that the derived components encode motifs of the RNA secondary structures without pseudoknots associated with the plane trees very well. As an application, we present an algorithm (RakeSamp) generating uniformly random secondary structures without pseudoknots that satisfy fine motif specifications on the length and degree of various types of loops as well as helices.
Mathematical models rooted in network representations are becoming increasingly more common for capturing a broad range of phenomena. Boolean networks (BNs) represent a mathematical abstraction suited for establishing general theory applicable to such systems. A key thread in BN research is developing theory that connects the structure of the network and the local rules to phase space properties or so-called structure-to-function theory. While most theory for BNs has been developed for the synchronous case, the focus of this work is on asynchronously updated BNs (ABNs) which are natural to consider from the point of view of applications to real systems where perfect synchrony is uncommon. A central question in this regard is sensitivity of dynamics of ABNs with respect to perturbations to the asynchronous update scheme. Macauley & Mortveit [Nonlinearity 22, 421-436 (2009)] showed that the periodic orbits are structurally invariant under toric equivalence of the update sequences. In this paper and under the same equivalence of the update scheme, the authors (i) extend that result to the entire phase space, (ii) establish a Lipschitz continuity result for sequences of maximal transient paths, and (iii) establish that within a toric equivalence class the maximal transient length may at most take on two distinct values. In addition, the proofs offer insight into the general asynchronous phase space of Boolean networks.
We propose a novel mathematical paradigm for the study of genetic variation in sequence alignments. This framework originates from extending the notion of pairwise relations, upon which current analysis is based on, to k-ary dissimilarity. This dissimilarity naturally leads to a generalization of simplicial complexes by endowing simplices with weights, compatible with the boundary operator. We introduce the notion of k-stances and dissimilarity complex, the former encapsulating arithmetic as well as topological structure expressing these k-ary relations. We study basic mathematical properties of dissimilarity complexes and show how this approach captures watershed moments of viral dynamics in the context of SARS-CoV-2 and H1N1 flu genomic data.
On the occasion of Dr. Michael Waterman's 80th birthday, we review his major contributions to the field of computational biology and bioinformatics including the famous Smith-Waterman algorithm for sequence alignment, the probability and statistics theory related to sequence alignment, algorithms for sequence assembly, the Lander-Waterman model for genome physical mapping, combinatorics and predictions of ribonucleic acid structures, word counting statistics in molecular sequences, alignment-free sequence comparison, and algorithms for haplotype block partition and tagSNP selection related to the International HapMap Project. His books Introduction to Computational Biology: Maps, Sequences and Genomes for graduate students and Computational Genome Analysis: An Introduction geared toward undergraduate students played key roles in computational biology and bioinformatics education. We also highlight his efforts of building the computational biology and bioinformatics community as the founding editor of the Journal of Computational Biology and a founding member of the International Conference on Research in Computational Molecular Biology (RECOMB).
We present a novel framework facilitating the rapid detection of variants of interest (VOI) and concern (VOC) in a viral multiple sequence alignment (MSA). The framework is purely based on the genomic sequence data, without requiring prior established biological analysis. The framework’s building blocks are sets of co-evolving sites (motifs), identified via co-evolutionary signals within the MSA. Motifs form a weighted simplicial complex, whose vertices are sites that satisfy a certain nucleotide diversity. Higher dimensional simplices are constructed using distances quantifying the co-evolutionary coupling of pairs and in the context of our method maximal motifs manifest as clusters. The framework triggers an alert via a cluster with a significant fraction of newly emerging polymorphic sites. We apply our method to SARS-CoV-2, analyzing all alerts issued from November 2020 through August 2021 with weekly resolution for England, USA, India and South America. Within a week at most a handful of alerts, each of which involving on the order of 10 sites are triggered. Cross referencing alerts with a posteriori knowledge of VOI/VOC-designations and lineages, motif-induced alerts detect VOIs/VOCs rapidly, typically weeks earlier than current methods. We show how motifs provide insight into the organization of the characteristic mutations of a VOI/VOC, organizing them as co-evolving blocks. Finally we study the dependency of the motif reconstruction on metric and clustering method and provide the receiver operating characteristic (ROC) of our alert criterion.
This paper presents a novel virus surveillance framework, completely independent of phylogeny-based methods. The framework issues timely alerts with an accuracy exceeding 85% that are based on the co-evolutionary relations between sites of the viral multiple sequence array (MSA). This set of relations is formalized via a motif complex, whose dynamics contains key information about the emergence of viral threats without the referencing of strain prevalence. Our notion of threat is centered at the emergence of a certain type of critical cluster consisting of key co-evolving sites. We present three case studies, based on GISAID data from UK, US and New York, where we perform our surveillance. We alert on May 16, 2022, based on GISAID data from New York, to a critical cluster of co-evolving sites mapping to the Pango-designation, BA.5. The alert specifies a cluster of seven genomic sites, one of which exhibits D3N on the M (membrane) protein–the distinguishing mutation of BA.5, three encoding ORF6:D61L and the remaining three exhibiting the synonymous mutations C26858T, C27889T and A27259C. New insight is obtained: when projected onto sequences, this cluster splits into two, mutually exclusive blocks of co-evolving sites (m:D3N,nuc:C27889T) linked to the five reverse mutations (nuc:C26858T,nuc:A27259C,ORF6:D61L). We furthermore provide an in depth analysis of all major signaled threats, during which we discover a specific signature concerning linked reverse mutation in the critical cluster.### Competing Interest StatementThe authors have declared no competing interest.### Funding StatementThis work was partially supported by the VDH Grant PV-BII VDH COVID-19 Modeling Program VDH-21-501-0135.### Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesI confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable.YesAll data produced in the present study are available upon reasonable request to the authors
In this paper, we analyze the homology of the simplicial complex induced by a given pair of RNA secondary structures, $$R=(S,T)$$ . Such a pair induces a bi-secondary structure, whose associated loop nerve X is the simplicial complex obtained by loop intersections. We will provide an algebraic proof of the fact that $$H_1(X)=0$$ . We will provide a combinatorial interpretation for the generators of $$H_2(X)$$ in terms of crossing components of the bi-structure and establish that the rank of $$H_2(X)$$ equals the total number of such crossing components. Finally, we shall prove that each crossing component naturally encodes a triangulation of a 2-sphere and provide an analysis of the geometric realization of X.
Recently Yoffe et al. observed that the average distances between 5′-3′ ends of RNA molecules are very small and largely independent of sequence length. This observation is based on numerical computations as well as theoretical arguments maximizing certain entropy functionals. In this paper we compute the exact distribution of 5′-3′ distances of RNA secondary structures for any finite n. We furthermore compute the limit distribution and show that already for n = 30 the exact distribution and the limit distribution are very close. Our results show that the distances of random RNA secondary structures are distinctively lower than those of minimum free energy structures of random RNA sequences. 1
COVID-19 is an infectious disease caused by the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). The viral genome is considered to be relatively stable and the mutations that have been observed and reported thus far are mainly focused on the coding region. This article provides evidence that macrolevel pandemic dynamics, such as social distancing, modulate the genomic evolution of SARS-CoV-2. This view complements the prevalent paradigm that microlevel observables control macrolevel parameters such as death rates and infection patterns. First, we observe differences in mutational signals for geospatially separated populations such as the prevalence of A23404G in CA versus NY and WA. We show that the feedback between macrolevel dynamics and the viral population can be captured employing a transfer entropy framework. Second, we observe complex interactions within mutational clades. Namely, when C14408T first appeared in the viral population, the frequency of A23404G spiked in the subsequent week. Third, we identify a noncoding mutation, G29540A, within the segment between the coding gene of the N protein and the ORF10 gene, which is largely confined to NY (>95%). These observations indicate that macrolevel sociobehavioral measures have an impact on the viral genomics and may be useful for the dashboard-like tracking of its evolution. Finally, despite the fact that SARS-CoV-2 is a genetically robust organism, our findings suggest that we are dealing with a high degree of adaptability. Owing to its ample spread, mutations of unusual form are observed and a high complexity of mutational interaction is exhibited.
An RNA bi-structure is a pair of RNA secondary structures that are considered as arc-diagrams. We present a novel weighted homology theory for RNA bi-structures, which was obtained through the intersections of loops. The weighted homology of the intersection complex X features a new boundary operator and is formulated over a discrete valuation ring, R. We establish basic properties of the weighted complex and show how to deform it in order to eliminate any 3-simplices. We connect the simplicial homology, Hi(X), and weighted homology, Hi,R(X), in two ways: first, via chain maps, and second, via the relative homology. We compute H0,R(X) by means of a recursive contraction procedure on a weighted spanning tree and H1,R(X) via an inflation map, by which the simplicial homology of the 1-skeleton allows us to determine the weighted homology H1,R(X). The homology module H2,R(X) is naturally obtained from H2(X) via chain maps. Furthermore, we show that all weighted homology modules Hi,R(X) are trivial for i>2. The invariant factors of our structure theorems, as well as the weighted Whitehead moves facilitating the removal of filled tetrahedra, are given a combinatorial interpretation. The weighted homology of bi-structures augments the simplicial counterpart by introducing novel torsion submodules and preserving the free submodules that appear in the simplicial homology.
William Y. C. Chen合作论文数Center for Combinatorics|Nankai University2