We consider the Schrödinger bridge problem which, given ensemble measurements of the initial and final configurations of a stochastic dynamical system and some prior knowledge on the dynamics, aims to reconstruct the "most likely" evolution of the system compatible with the data. Most existing literature assume Brownian reference dynamics, and are implicitly limited to modelling systems driven by the gradient of a potential energy. We depart from this regime and consider reference processes described by a multivariate Ornstein-Uhlenbeck process with generic drift matrix $\mathbf{A} \in \mathbb{R}^{d \times d}$. When $\mathbf{A}$ is asymmetric, this corresponds to a non-equilibrium system in which non-gradient forces are at play: this is important for applications to biological systems, which naturally exist out-of-equilibrium. In the case of Gaussian marginals, we derive explicit expressions that characterise exactly the solution of both the static and dynamic Schrödinger bridge. For general marginals, we propose mvOU-OTFM, a simulation-free algorithm based on flow and score matching for learning an approximation to the Schrödinger bridge. In application to a range of problems based on synthetic and real single cell data, we demonstrate that mvOU-OTFM achieves higher accuracy compared to competing methods, whilst being significantly faster to train.
Deep learning methods have revolutionised our ability to predict protein structures, allowing us a glimpse into the entire protein universe. As a result, our understanding of how protein structure drives function is now lagging behind our ability to determine and predict protein structure. Here, we describe how topology, the branch of mathematics concerned with qualitative properties of spatial structures, provides a lens through which we can identify fundamental organising features across the known protein universe. We identify topological determinants that capture global features of the protein universe, such as domain architecture and binding sites. Additionally, our analysis identifies highly specific properties, so-called topological generators, that can be used to provide deeper insights into protein structure-function and evolutionary relationships. We present a practical methodology for mapping the topology of the known protein universe at scale. We then use our approach to determine structural, functional and disease consequences of mutations. Our approach reveals and helps to explain differences in properties of proteins in mesophiles and thermophiles, and the likely structural and functional consequences of polymorphisms in a protein. For eukaryotes we find striking differences between protein topologies in multi-cellular and single-celled organisms.
Cells actively regulate their size during the cell cycle to maintain volume homeostasis across generations. While various mathematical models of cell size regulation have been proposed to explain how this is achieved, relating these models to experimentally observed cell size distributions has proved challenging. In this paper we present a simple formula for the cell size distribution in lineages as observed in e.g. a mother machine, and provide a new derivation for the corresponding result in populations, assuming exponential cell growth. Our results are independent of the underlying cell size control mechanism and explain the characteristic shape underlying experimentally observed cell size distributions. We furthermore derive universal moment identities for these distributions, and show that our predictions agree well with experimental measurements of E. coli cells, both on the distribution and the moment level.
Many cellular processes involve information processing and decision-making. We can probe these processes at increasing molecular detail. The analysis of heterogeneous data remains a challenge that requires new ways of thinking about cells in quantitative, predictive, and mechanistic ways. We discuss the role of mathematical models in the context of cell-fate decision-making systems across the tree of life. Complex multicellular organisms have been a particular focus, but single-celled organisms also have to sense and respond to their environment. We center our discussion around the idea of design principles that we can learn from observations and modeling and exploit in order to (re)-design or guide cellular behavior.
Biological time can be measured in two ways: in generations and in physical (chronological) time. When generations overlap, these two notions diverge, which impedes our ability to relate mathematical models to real populations. In this paper we show that nevertheless, the two clocks can be synchronised in the long run via a simple identity relating generational and physical time. This equivalence allows us to directly translate statements from the generational picture to the physical picture and vice versa. We derive a generalized Euler-Lotka equation linking the basic reproduction number R_0 to the growth rate, and present a simple identity that relates the selection coefficient of a mutation to the history of typical individuals, with applications to epidemiology, population biology and microbial growth.
Cell dynamics and biological function are governed by intricate networks of molecular interactions. Inferring these interactions from data is a notoriously difficult inverse problem. Most existing network inference methods construct population-averaged representations of gene interaction networks, and they do not naturally allow us to infer differences in interaction activity across heterogeneous cell populations. We introduce locaTE, an information theoretic approach that leverages single-cell, dynamical information, together with geometry of the cell-state manifold, to infer cell-specific, causal gene interaction networks in a manner that is agnostic to the topology of the underlying biological trajectory. Through extensive simulation studies and applications to experimental datasets spanning mouse primitive endoderm formation, pancreatic development, and hematopoiesis, we demonstrate superior performance and the generation of additional insights, compared with standard population-averaged inference methods. We find that locaTE provides a powerful network inference method that allows us to distil cell-specific networks from single-cell data. A record of this paper’s transparent peer review process is included in the supplemental information.
Topological data analysis is a powerful tool for describing topological signatures in real world data. An important challenge in topological data analysis is matching significant topological signals across distinct systems. In geometry and probability theory, optimal transport formalises notions of distance and matchings between distributions and structured objects. We propose to combine these approaches, constructing a mathematical framework for optimal transport-based matchings of topological features. Building upon recent advances in the domains of persistent homology and optimal transport for hypergraphs, we develop a transport-based methodology for topological data processing. We define measure topological networks, which integrate both geometric and topological information about a system, introduce a distance on the space of these objects, and study its metric properties, showing that it induces a geodesic metric space of non-negative curvature. The resulting Topological Optimal Transport (TpOT) framework provides a transport model on point clouds that minimises topological distortion while simultaneously yielding a geometrically informed matching between persistent homology cycles.
SummaryTuring patterns1are well-known self-organising systems that can form spots, stripes, or labyrinths. They represent a major theory of patterning in tissue organisation, due to their remarkable similarity to some natural patterns, such as skin pigmentation in zebrafish2, digit spacing3,4, and many others. The involvement of Turing patterns in biology has been debated because of their stringent fine-tuning requirements, where patterns only occur within a small subset of parameters5,6. This has complicated the engineering of a synthetic gene circuit for Turing patterns from first principles, even though natural genetic Turing networks have been successfully identified4,7. Here, we engineered a synthetic genetic reaction-diffusion system where three nodes interact according to a non-classical Turing network with improved parametric robustness6. The system was optimised inE. coliand reproducibly generated stationary, periodic, concentric stripe patterns in growing colonies. The patterns were successfully reproduced with a partial differential equation model, in a parameter regime obtained by fitting to experimental data. Our synthetic Turing system can contribute to novel nanotechnologies, such as patterned biomaterial deposition8,9, and provide insights into developmental patterning programs10.
Here we introduce simple structures for the analysis of complex hypergraphs, hypergraph animals. These structures are designed to describe the local node neighbourhoods of nodes in hypergraphs. We establish their relationships to lattice animals and network motifs and develop their combinatorial properties, for sparse and uncorrelated hypergraphs. Here we can make use of the tight link of hypergraph animals to partition numbers, which opens up a vast mathematical framework for the analysis of hypergraph animals. We then study their abundances in random hypergraphs. Two transferable insights result from this analysis: (i) it establishes the importance of high-cardinality edges in ensembles of random hypergraphs that are inspired by the classical Erdos-Reny\'i random graphs; and (ii) there is a close connection between degree and hyperedge cardinality in random hypergraphs that shapes the animal abundances and spectra profoundly. Together these two findings imply that we need to spend more effort on investigating and developing suitable conditional ensembles of random hypergraph that can capture real-world structures and their complex dependency structures.
Turing patterns are self-organizing systems that can form spots, stripes, or labyrinths. Proposed examples in tissue organization include zebrafish pigmentation, digit spacing, and many others. The theory of Turing patterns in biology has been debated because of their stringent fine-tuning requirements, where patterns only occur within a small subset of parameters. This has complicated the engineering of synthetic Turing gene circuits from first principles, although natural genetic Turing networks have been identified. Here, we engineered a synthetic genetic reaction-diffusion system where three nodes interact according to a non-classical Turing network with improved parametric robustness. The system reproducibly generated stationary, periodic, concentric stripe patterns in growing E. coli colonies. A partial differential equation model reproduced the patterns, with a Turing parameter regime obtained by fitting to experimental data. Our synthetic Turing system can contribute to nanotechnologies, such as patterned biomaterial deposition, and provide insights into developmental patterning programs. A record of this paper's transparent peer review process is included in the supplemental information.
Cells are the fundamental units of life, and like all life forms, they change over time. Changes in cell state are driven by molecular processes; of these many are initiated when molecule numbers reach and exceed specific thresholds, a characteristic that can be described as "digital cellular logic". Here we show how molecular and cellular noise profoundly influence the time to cross a critical threshold-the first-passage time-and map out scenarios in which stochastic dynamics result in shorter or longer average first-passage times compared to noise-less dynamics. We illustrate the dependence of the mean first-passage time on noise for a set of exemplar models of gene expression, auto-regulatory feedback control, and enzyme-mediated catalysis. Our theory provides intuitive insight into the origin of these effects and underscores two important insights: (i) deterministic predictions for cellular event timing can be highly inaccurate when molecule numbers are within the range known for many cells; (ii) molecular noise can significantly shift mean first-passage times, particularly within auto-regulatory genetic feedback circuits. Cells exhibit remarkable temporal precision in regulating their internal states. Here, by solving stochastic first passage time problems for key molecular processes Ham, Coomer et al. shed light on how cells achieve this precision.
Single-cell technologies allow us to gain insights into cellular processes at unprecedented resolution. In stem cell and developmental biology snapshot data allow us to characterize how the transcriptional states of cells change between successive cell types. Here, we show how approximate Bayesian computation (ABC) can be employed to calibrate mathematical models against single-cell data. In our simulation study, we demonstrate the pivotal role of the adequate choice of distance measures appropriate for single-cell data. We show that for good distance measures, notably optimal transport with the Sinkhorn divergence, we can infer parameters for mathematical models from simulated single-cell data. We show that the ABC posteriors can be used (i) to characterize parameter sensitivity and identify dependencies between different parameters and (ii) to construct representations of the Waddington or epigenetic landscape, which forms a popular and interpretable representation of the developmental dynamics. In summary, these results pave the way for fitting mechanistic models of stem cell differentiation to single-cell data.
Single-cell data afford unprecedented insights into molecular processes. But the complexity and size of these data sets have proved challenging and given rise to a large armory of statistical and machine learning approaches. The majority of approaches focuses on either describing features of these data, or making predictions and classifying unlabeled samples. In this study, we introduce repeated decision stumping (ReDX) as a method to distill simple models from single-cell data. We develop decision trees of depth one-hence "stumps"-to identify in an inductive manner, gene products involved in driving cell fate transitions, and in applications to published data we are able to discover the key players involved in these processes in an unbiased manner without prior knowledge. Our algorithm is deliberately targeting the simplest possible candidate hypotheses that can be extracted from complex high-dimensional data. There are three reasons for this: (1) the predictions become straightforwardly testable hypotheses; (2) the identified candidates form the basis for further mechanistic model development, for example, for engineering and synthetic biology interventions; and (3) this approach complements existing descriptive modeling approaches and frameworks. The approach is computationally efficient, has remarkable predictive power, including in simulation studies where the ground truth is known, and yields robust and statistically stable predictors; the same set of candidates is generated by applying the algorithm to different subsamples of experimental data.
Major computational challenges exist in relation to the collection, curation, processing and analysis of large genomic and imaging datasets, as well as the simulation of larger and more realistic models in systems biology. Here we discuss how a relative newcomer among programming languages—Julia—is poised to meet the current and emerging demands in the computational biosciences and beyond. Speed, flexibility, a thriving package ecosystem and readability are major factors that make high-performance computing and data analysis available to an unprecedented degree. We highlight how Julia’s design is already enabling new ways of analyzing biological data and systems, and we provide a list of resources that can facilitate the transition into Julian computing.
Recent studies applying advanced imaging techniques are changing the way we understand bacterial cell surfaces, bringing new knowledge on everything from single-cell heterogeneity in bacterial populations to their drug sensitivity and mechanisms of antimicrobial resistance. In both Gram-positive and Gram-negative bacteria, the outermost surface of the bacterial cell is being imaged at nanoscale; as a result, topographical maps of bacterial cell surfaces can be constructed, revealing distinct zones and specific features that might uniquely identify each cell in a population. Functionally defined assembly precincts for protein insertion into the membrane have been mapped at nanoscale, and equivalent lipid-assembly precincts are suggested from discrete lipopolysaccharide patches. As we review here, particularly for Gram-negative bacteria, the applications of various modalities of nanoscale imaging are reawakening our curiosity about what is conceptually a 3D cell surface landscape: what it looks like, how it is made and how it provides resilience to respond to environmental impacts.
• Edmund Crampin was Professor of Systems Biology at the University of Melbourne. He passed away on 15 May 2021. • A workshop in honour of Edmund was held at the University of Melbourne in November 2022. • 11 of his colleagues gave scientific presentations and reflected on Edmund’s impact in mathematical biology. • A panel discussion explored the future of mathematical biology. • Edmund leaves a legacy of outstanding academic achievements and fostering of a collegial research culture in mathematical biology.
One of the best known ways bacteria cells understand and respond to the environment are through Two-Component Systems (TCS). These signalling systems are highly diverse in function and can detect a range of physical stimuli including molecular concentrations and temperature, with a range of responses including chemotaxis and anaerobic energy production. TCS exhibit a range of different molecular structures and energy costs, and multiple types co-exist in the same cell. TCSs that incur relatively high energy cost are abundant in biology, despite strong evolutionary pressure to efficiently spend energy. We are motivated to discern what benefits, if any, the more energetically expensive variants had for a cell. We seek to answer this question by modelling energy flow through two variants of TCS. This was accomplished using bond graphs, a physics-based modelling framework that accurately models energy transfer through different physical domains. Our analysis demonstrates that energy availability can affect a cell’s signal sensitivity, noise filtering effectiveness, and the stimulus level where cell response is maximal. We also found that these properties are determined not by the molecular parameters themselves, but the reaction rate parameters that govern the reaction systems as a whole. This suggests possible connections between the molecular structure and evolutionary purpose of any two-component system. This opens the door to new synthetic circuit design in systems biology, and we propose new hypotheses about this link between structure and purpose that could be experimentally verified. Author summary Two-component systems are the main way many bacteria sense and respond to their environment. They exist in such well-studied bacteria as E. coli where they have been shown to detect a range of stimuli including nutrients, temperature, acidity, and pressure. Two-component systems are ubiquitous in bacteria yet have a deceptively simple structure. Knowing how they operate and the purpose of variations in signalling structure is helpful to our understanding of cellular biology and the design of synthetic biological circuits. Critical unanswered questions remain about the energy usage and functional benefits of these systems. We sought to improve our understanding of two-component systems by applying a physics-based modelling framework. We found that tracking energy flow through the cell reveals new energy-dependent behaviour in signalling sensitivity, noise filtering, and maximal cell response. We also found that these properties are not strictly dependent on the molecular properties themselves, but from the configuration of the reaction system as a whole.
Network inference is a notoriously challenging problem. Inferred networks are associated with high uncertainty and likely riddled with false positive and false negative interactions. Especially for biological networks we do not have good ways of judging the performance of inference methods against real networks, and instead we often rely solely on the performance against simulated data. Gaining confidence in networks inferred from real data nevertheless thus requires establishing reliable validation methods. Here, we argue that the expectation of mixing patterns in biological networks such as gene regulatory networks offers a reasonable starting point: interactions are more likely to occur between nodes with similar biological functions. We can quantify this behaviour using the assortativity coefficient, and here we show that the resulting heuristic, functional assortativity , offers a reliable and informative route for comparing different inference algorithms.