T-cell activation models describe the way in which antigens and T-cell receptors (TCRs) interact in order to trigger an immune response. A number of mathematical models of TCR-mediated T-cell activation have been proposed in the last decades. Often, the only measurable quantities available for the identification of these models are Emax, the maximum achievable T-cell activation level (over the background level), and EC50, that is, the antigen concentration that results in 50% of Emax. In order to analyse model properties such as identifiability, observability and sensitivity, it is necessary to know how Emax and/or EC50 relate to the model parameters. However, their expressions are only available for a few of the published models of T-cell activation. Here, we present a general method to derive the expressions of Emax and/or EC50, and apply it to those models for which they are yet unknown. Our results expand the range of analyses that can be performed on T-cell activation models, contributing to enable their exploitation to understand and control the immune response.
Structural identifiability and observability are properties that describe the ability to determine the parameters and states of a model from knowledge of its equations, inputs, and outputs. A model is structurally unidentifiable if different parameter sets produce identical outputs. This issue may arise in partially observed systems. In this work, we follow an approach, to which we refer as the SIM algorithm, to analyse identifiability and observability in dynamical systems based on finding its scaling symmetries. The proposed method relies on analysing the invariance of the system under parameter and state scaling transformations and can be systematically applied regardless of the model’s dimensionality. We report the implementation of the SIM method for ODE models in the STRIKE-GOLDD MATLAB toolbox, and extend it to partial differential equation (PDE) models through the newly introduced SIM-PDE algorithm. The approach is straightforward to implement and computationally efficient, making it suitable even for highly nonlinear and large-scale models.
T-cell receptor (TCR)-mediated T-cell activation is a key process in adaptive immune responses. The complexity of this process has led to the development of different mathematical models that seek to describe and predict the conditions of antigen-TCR interactions required for TCR triggering and T-cell activation. These models are characterized by describing different sets of sequential molecular interactions and their kinetics, positing the generation of a final product as a necessary and sufficient condition for T-cell activation. Such modeling could provide an effective tool for simulating antigen recognition by T cells and, consequently, aid in the design of effective therapeutic strategies. However, it is necessary to previously assess the predictive capabilities of the proposed models when fitted to experimental data. As a first step towards this goal, in this work we examine the parameter identifiability and sensitivity of the published models of TCR-based T-cell activation. For each model, we consider different, often experimentally measured, output quantities and show how their availability affects the results. These analyses allow us to determine the ability of each model to correctly describe different experimental situations, and to establish to what extent these models can be applied to reliably predict and control T-cell activation by specific therapeutic targets.
We discuss the use of symmetries for analysing the structural identifiability and observability of control systems. Special emphasis is put on the role of discrete symmetries, in contrast to the more commonly studied continuous or Lie symmetries. We argue that discrete symmetries are the origin of parameters which are structurally locally identifiable, but not globally. We exploit this fact to present a methodology for structural identifiability analysis that detects such parameters and characterizes the symmetries in which they are involved. We demonstrate the use of our methodology by applying it to four case studies.
Synthetic biology is a recent area of biological engineering, whose aim is to provide cells with novel functionalities. A number of important results regarding the development of control circuits in synthetic biology have been achieved during the last decade. A differential geometry approach can be used for the analysis of said systems, which are often nonlinear. Here we demonstrate the application of such tools to analyse the structural identifiability, observability, accessibility, and controllability of several biomolecular systems. We focus on a set of synthetic circuits of current interest, which can perform several tasks, both in open loop and closed loop settings. We analyse their properties with our own methods and tools; further, we describe a new open-source implementation of the techniques.
Mechanistic cancer models can encapsulate beliefs about the main factors influencing tumour growth. In recent decades, many different types of dynamic models have been used for this purpose. The integration of a model’s differential equations yields a simulation of the behaviour of the system over time, thus enabling tumour progression to be predicted. A requisite for the reliability of these quantitative predictions is that the model is structurally identifiable and observable, i.e., that it is theoretically possible to infer the correct values of its parameters and state variables from time course data. In this paper, we show how to analyse these properties of tumour growth models using a well-established methodology, which we implemented previously in an open-source software tool. To this end, we provide an account of 20 published models described by ordinary differential equations, some of which incorporate the effect of interventions including chemotherapy, radiotherapy, and immunotherapy. For each model, we describe its equations and analyse their structural identifiability and observability, discussing how they are affected by the experimental design. We provide computational implementations of these models, which enable readily reproducing results. Our results inform about the possibility of inferring the parameters and state variables of a given model using a specific measurement setup, and, together with the corresponding methodology and implementation, they can be used as a blueprint for analysing other models not included here. Thus, this paper serves as a guide to select the most appropriate model for each application.
Mathematical modelling is one of the pillars of systems biology. In this review, we focus on models that are mechanistic, i.e., they explain the mechanism by which a phenomenon takes place, and dynamic, i.e., they consist of differential equations that simulate the time course of a system. Our aim is to provide an updated state of the art of mechanistic dynamic modelling in systems biology. These models, which are based on first principles, are crucial for obtaining insights about complex physiological processes. They can be used to test hypotheses, predict system behaviour, and explore and optimize intervention strategies. Since biological processes are typically nonlinear, multiscale, and subject to various sources of uncertainty, the task of building and analysing robust and reliable mechanistic models is fraught with difficulties. In this paper, we provide an overview of recent developments in key topics such as model discovery and structure selection, identifiability analysis, parameter estimation, uncertainty quantification, and model reliability. We discuss the challenges and open questions in these areas and outline perspectives for future work.
Determining the accessibility of a nonlinear system is a classical geometric control problem. Accessibility is equivalent to controllability for linear systems at equilibrium points. State observability and controllability are commonly considered as dual properties in a certain sense. Here we investigate the relationship between accessibility and input observability in linear and nonlinear dynamical systems of ordinary differential equations. We start by studying linear systems, building on well known results about their controllability and observability, and then explore ways in which they may be extended to the nonlinear case. Our main finding is that, for a system to be accessible, the rank of the k-th row input observability matrix has to be bounded from below by the rank of the accessibility matrix at the k-th step of the accessibility algorithm. This result provides a non-accessibility test: for systems with exactly one input, local $ n_x $ nx-row observability of the input, assessed by assuming full observation of the state, is a necessary condition for accessibility.
Accessibility and observability are two properties of dynamic models that provide insights into the structural relationships between their input, output, and state variables. They are closely related to controllability and structural local identifiability, respectively. Observability and identifiability determine, respectively, the possibility of inferring the unmeasured state variables and parameters of a model from output measurements; accessibility and controllability describe the possibility of driving its state by changing its input. Analysing these structural properties in nonlinear models of ordinary differential equations can be challenging, particularly when dealing with large systems. Two main approaches are currently used for their study: one based on differential geometry, which uses symbolic computation, and another one based on sensitivity calculations that uses numerical integration. These approaches are implemented in two MATLAB (R2024b) software tools: the differential geometry approach in STRIKE-GOLDD, and the sensitivity-based method in StrucID. These toolboxes differ significantly in their features and capabilities. Until now, their performance had not been thoroughly compared. In this paper we present a comprehensive comparative study of them, elucidating their differences in applicability, computational efficiency, and robustness against computational issues. Our core finding is that StrucID has a substantially lower computational cost than STRIKE-GOLDD; however, it may occasionally yield inconsistent results due to numerical issues.
The concept of identifiability describes the possibility of inferring the parameters of a dynamic model by observing its output. It is common and useful to distinguish between structural and practical identifiability. The former property is fully determined by the model equations, while the latter is also influenced by the characteristics of the available experimental data. Structural identifiability can be determined by means of symbolic computations, which may be performed before collecting experimental data, and are hence sometimes called a priori analyses. Practical identifiability is typically assessed numerically, with methods that require simulations—and often also optimization—and are applied a posteriori. An approach to study structural local identifiability is to consider it as a particular case of observability, which is the possibility of inferring the internal state of a system from its output. Thus, both properties can be analysed jointly, by building a generalized observability matrix and computing its rank. The aim of this paper is to investigate to which extent such observability-based methods can also inform about practical aspects related with the experimental setup, which are usually not approached in this way. To this end, we explore a number of possible extensions of the rank tests, and discuss the purposes for which they can be informative as well as others for which they cannot.
Background: We address the problem of determining the controllability and accessibility of nonlinear biosystems. We consider models described by affine-in-inputs ordinary differential equations, which are adequate for a wide array of biological processes. Roughly speaking, the controllability of a dynamical system determines the possibility of steering it from an initial state to any point in its neighbourhood; accessibility is a weaker form of controllability. Methods: While the methodology for analysing the controllability of linear systems is well established, its generalization to the nonlinear case has proven elusive. Thus, a number of related but different properties - including different versions of accessibility, reachability or weak local controllability - have been defined to approach its study, and several partial results exist in lieu of a general test. Here, leveraging the applicable results from differential geometric control theory, we source sufficient conditions to assess nonlinear controllability, as well as a necessary and sufficient condition for accessibility.Results: We develop an algorithmic procedure to evaluate these conditions efficiently, and we provide its open source implementation. Using this software tool, we analyse the accessibility and controllability of a number of models of biomedical interest. While some of them are fully controllable, we find others that are not, as is the case of some models of EGF and NFX-B signalling networks.Conclusions: The contributions in this paper facilitate the accessibility and controllability analysis of nonlinear models, not only in biomedicine but also in other areas in which they have been rarely performed to date.
Biological communities are populations of various species interacting in a common location. Microbial communities, which are formed by microorganisms, are ubiquitous in nature and are increasingly used in biotechnological and biomedical applications. They are nonlinear systems whose dynamics can be accurately described by models of ordinary differential equations (ODEs). A number of ODE models have been proposed to describe microbial communities. However, the structural identifiability and observability of most of them-that is, the theoretical possibility of inferring their parameters and internal states by observing their output-have not been determined yet. It is important to establish whether a model possesses these properties, because, in their absence, the ability of a model to make reliable predictions may be compromised. Hence, in this paper, we analyse these properties for the main families of microbial community models. We consider several dimensions and measurements; overall, we analyse more than a hundred different configurations. We find that some of them are fully identifiable and observable, but a number of cases are structurally unidentifiable and/or unobservable under typical experimental conditions. Our results help in deciding which modelling frameworks may be used for a given purpose in this emerging area, and which ones should be avoided.
Abstract Motivation Dynamic mechanistic modelling in systems biology has been hampered by the complexity and variability associated with the underlying interactions, and by uncertain and sparse experimental measurements. Ensemble modelling, a concept initially developed in statistical mechanics, has been introduced in biological applications with the aim of mitigating those issues. Ensemble modelling uses a collection of different models compatible with the observed data to describe the phenomena of interest. However, since systems biology models often suffer from a lack of identifiability and observability, ensembles of models are particularly unreliable when predicting non-observable states. Results We present a strategy to assess and improve the reliability of a class of model ensembles. In particular, we consider kinetic models described using ordinary differential equations with a fixed structure. Our approach builds an ensemble with a selection of the parameter vectors found when performing parameter estimation with a global optimization metaheuristic. This technique enforces diversity during the sampling of parameter space and it can quantify the uncertainty in the predictions of state trajectories. We couple this strategy with structural identifiability and observability analysis, and when these tests detect possible prediction issues we obtain model reparameterizations that surmount them. The end result is an ensemble of models with the ability to predict the internal dynamics of a biological process. We demonstrate our approach with models of glucose regulation, cell division, circadian oscillations and the JAK-STAT signalling pathway. Availability and implementation The code that implements the methodology and reproduces the results is available at https://doi.org/10.5281/zenodo.6782638. Supplementary information Supplementary data are available at Bioinformatics online.
Mechanistic dynamical models allow us to study the behavior of complex biological systems. They can provide an objective and quantitative understanding that would be difficult to achieve through other means. However, the systematic development of these models is a non-trivial exercise and an open problem in computational biology. Currently, many research efforts are focused on model discovery, i.e. automating the development of interpretable models from data. One of the main frameworks is sparse regression, where the sparse identification of nonlinear dynamics (SINDy) algorithm and its variants have enjoyed great success. SINDy-PI is an extension which allows the discovery of rational nonlinear terms, thus enabling the identification of kinetic functions common in biochemical networks, such as Michaelis-Menten. SINDy-PI also pays special attention to the recovery of parsimonious models (Occam's razor). Here we focus on biological models composed of sets of deterministic nonlinear ordinary differential equations. We present a methodology that, combined with SINDy-PI, allows the automatic discovery of structurally identifiable and observable models which are also mechanistically interpretable. The lack of structural identifiability and observability makes it impossible to uniquely infer parameter and state variables, which can compromise the usefulness of a model by distorting its mechanistic significance and hampering its ability to produce biological insights. We illustrate the performance of our method with six case studies. We find that, despite enforcing sparsity, SINDy-PI sometimes yields models that are unidentifiable. In these cases we show how our method transforms their equations in order to obtain a structurally identifiable and observable model which is also interpretable.
Structural identifiability determines the possibility of estimating the parameters of a model by observing its output in an ideal experiment. If a parameter is structurally locally identifiable, but not globally (SLING), its true value cannot be uniquely inferred because several equivalent solutions exist. In biological modeling it is sometimes assumed that local identifiability entails global identifiability, which is convenient because local identifiability tests are typically less computationally demanding than global tests. However, this assumption has never been investigated beyond demonstrating the existence of counter-examples. To clarify this matter, in this paper we began by asking how often a structurally locally identifiable parameter is not globally identifiable in systems biology. To answer this question empirically we assembled a collection of 102 mathematical models from the literature, with a total of 763 parameters. We analysed their identifiability, determining that approximately 5% of the parameters are SLING. Next we investigated how the SLING parameters arise, tracing their origin to particular features of the model equations. Finally, we investigated the possibility of obtaining false estimates. Some of the solutions that are mathematically equivalent to the true one involved parameters and/or initial conditions with negative values, which are not biologically meaningful. In other cases the true solution and the equivalent one were in the same range. These results provide insight about a previously unexplored hypothesis, and suggest that in most (albeit not all) systems biology applications it suffices to test for structural local identifiability.
Mechanistic dynamic models allow for a quantitative and systematic interpretation of data and the generation of testable hypotheses. However, these models are often over-parameterized, leading to non-identifiability and non-observability, i.e. the impossibility of inferring their parameters and state variables. The lack of structural identifiability and observability (SIO) compromises a model's ability to make predictions and provide insight. Here we present a methodology, AutoRepar, that corrects SIO deficiencies automatically, yielding reparameterized models that are structurally identifiable and observable. The reparameterization preserves the mechanistic meaning of selected variables, and has the exact same dynamics and input-output mapping as the original model. We implement AutoRepar as an extension of the STRIKE-GOLDD software toolbox for SIO analysis, applying it to several models from the literature to demonstrate its ability to repair their structural deficiencies. AutoRepar increases the applicability of mechanistic models, enabling them to provide reliable information about their parameters and dynamics.
Biological processes are often modelled using ordinary differential equations. The unknown parameters of these models are estimated by optimizing the fit of model simulation and experimental data. The resulting parameter estimates inevitably possess some degree of uncertainty. In practical applications it is important to quantify these parameter uncertainties as well as the resulting prediction uncertainty, which are uncertainties of potentially time-dependent model characteristics. Unfortunately, estimating prediction uncertainties accurately is nontrivial, due to the nonlinear dependence of model characteristics on parameters. While a number of numerical approaches have been proposed for this task, their strengths and weaknesses have not been systematically assessed yet. To fill this knowledge gap, we apply four state of the art methods for uncertainty quantification to four case studies of different computational complexities. This reveals the trade-offs between their applicability and their statistical interpretability. Our results provide guidelines for choosing the most appropriate technique for a given problem and applying it successfully.