In this age of big data and natural language processing, to what extent can we leverage new technologies and new tools to make progress in organizing disparate biomedical data sources? Imagine a system in which one could bring together sequencing data with phenotypes, gene expression data, and clinical information all under the same conceptual heading where applicable. Bio-ontologies seek to carry this out by organizing the relations between concepts and attaching the data to their corresponding concept. However, to accomplish this, we need considerable time and human input. Instead of resorting to human input alone, we describe a novel approach to obtaining the foundation for bio-ontologies: obtaining propositions (links between concepts) from biomedical text so as to fill the ontology. The heart of our approach is applying logic rules from Aristotelian logic and natural logic to biomedical information to derive propositions so that we can have material to organize knowledge bases (ontologies) for biomedical research. We demonstrate this approach by constructing a proof-of-principle bio-ontology for COVID-19 and related diseases.
Formalization of autopoiesis is an ongoing effort among theoretical biologists. In this field, Letelier and co-authors proposed that Robert Rosen's (M,R)-systems theory be used as a formalism for autopoiesis. In (M,R)-systems theory, Rosen proposes that one solve a set of functional closure equations (FCEs) which account for all of the components of the system as coming from within the system itself. A key part of the functional closure equations is the repair of the metabolism component of the system. Rosen's theory gives the organizational closure of the components as well as their products, as found in autopoiesis. However, according to Razeto-Barry (M,R)-systems leaves out some of the messiness and approximation that we find in autopoiesis as he reformulates it. A related problem is that though FCEs have a long history, they are difficult in practice to solve due to their mathematical formulation. In this paper we give a novel exact solution for the FCEs for continuous real vector-valued functions which is nevertheless difficult to compute. In addition we propose an extended form of FCEs which both captures more of the messiness of autopoiesis and also helps to make the FCEs more solvable. Finally, we use our solution for the extended FCEs to give an extended repair function for a metabolism taken from a representative class of biological dynamics for gene expression (the repressilator). More generally we show that one can use our solution for the extended FCEs to get an extended repair function for continuous real vector-valued functions.
Integrated information theory (IIT) starts from consciousness itself and identifies a set of properties (axioms) that are true of every conceivable experience. The axioms are translated into a set of postulates about the substrate of consciousness (called a complex), which are then used to formulate a mathematical framework for assessing both the quality and quantity of experience. The explanatory identity proposed by IIT is that an experience is identical to the cause–effect structure unfolded from a maximally irreducible substrate (a Φ-structure). In this work we introduce a definition for the integrated information of a system (φs) that is based on the existence, intrinsicality, information, and integration postulates of IIT. We explore how notions of determinism, degeneracy, and fault lines in the connectivity impact system-integrated information. We then demonstrate how the proposed measure identifies complexes as systems, the φs of which is greater than the φs of any overlapping candidate systems.
As part of the extended evolutionary synthesis, there has recently been a new emphasis on the effects of biological development on genetic inheritance and variation. The exciting new directions taken by those in the community have by a pre-history filled with related ideas that were never given a rigorous foundation or combined coherently. Part of the historical background of the extended synthesis is the work of James Mark Baldwin on his so-called “Baldwin Effect.” Many variant re-interpretations of his work obscure the original meaning of the Baldwin Effect. This paper emphasizes a new approach to the Baldwin Effect, focusing on his work in developmental psychology and how that would impact evolution. We propose a novel population genetics model of the Baldwin Effect. First, the impact of a kind of learning process motivated by motor babbling, in the developmental psychology literature, on evolution; second, that Information-theoretic phenotype reshaping speeds up evolution compared to populations without this kind of learning. The basic idea behind the model is to allow the organism to apply abstraction to his initial phenotype to situate it within one of a few different classes of phenotypes in the local neighborhood of a fitness maximum. The reshaping of the phenotype space thereby allows the organism to reach a nearby fitness maximum. By so doing, valleys in the fitness landscape are leveled out, making a rugged fitness landscape into a set of mesas and plateaus with increasing height. Using this model we can show the first sizeable speed-up for the Baldwin Effect compared to ordinary population genetics. We also introduce an information-theoretic foundation for the Baldwin Effect, which may be of independent interest.
Universal Semantic Communication (USC) is a theory that models communication among agents without the assumption of a fixed protocol. We demonstrate a connection, via a concept we refer to as process information, between a special case of USC and evolutionary processes. In this context, one agent attempts to interpret a potentially arbitrary signal produced within its environment. Sources of this effective signal can be modeled as a single alternative agent. Given a set of common underlying concepts that may be symbolized differently by different sources in the environment, any given entity must be able to correlate intrinsic information with input it receives from the environment in order to accurately interpret the ambient signal and ultimately coordinate its own actions. This scenario encapsulates a class of USC problems that provides insight into the semantic aspect of a model of evolution proposed by Rivoire and Leibler. Through this connection, we show that evolution corresponds to a means of solving a special class of USC problems, can be viewed as a special case of the Multiplicative Weights Updates algorithm, and that infinite population selection with no mutation and no recombination conforms to the Rivoire-Leibler model. Finally, using process information we show that evolving populations implicitly internalize semantic information about their respective environments.
We study the topology of Boolean functions from the perspective of Simplicial Homology, and characterize Simplicial Homology in turn by using Monotone Boolean functions. In so doing, we analyze to what extent topological invariants of Boolean functions (for instance the Euler characteristic) change under binary operations over Boolean functions and other operations, such as permutations over the input variables. We apply these tools to proving the ∆p2-hardness of calculating the Euler characteristic of general Boolean functions (as defined by Kulkarni and Santha[17]) and the co-NP -hardness of calculating the Euler characteristic for a simplicial complex of arbitrary dimension. We also show that calculating the Betti numbers for Simplicial Homology is co-NP hard. 1998 ACM Subject Classification F.1.3 Complexity Measures and Classes, F.2.2 Nonnumerical Algorithms and Problems, G.2.m Discrete Mathematics (Miscellaneous)
Even the most seasoned students of evolution, starting with Darwin himself, have occasionally expressed amazement that the mechanism of natural selection has produced the whole of Life as we see it around us. There is a computational way to articulate the same amazement: "What algorithm could possibly achieve all this in a mere three and a half billion years?" In this paper we propose an answer: We demonstrate that in the regime of weak selection, the standard equations of population genetics describing natural selection in the presence of sex become identical to those of a repeated game between genes played according to multiplicative weight updates ( MWUA), an algorithm known in computer science to be surprisingly powerful and versatile. MWUA maximizes a tradeoff between cumulative performance and entropy, which suggests a new view on the maintenance of diversity in evolution.
Structural controllability has been proposed as an analytical framework for making predictions regarding the control of complex networks across myriad disciplines in the physical and life sciences (Liu et al., Nature:473(7346):167-173, 2011). Although the integration of control theory and network analysis is important, we argue that the application of the structural controllability framework to most if not all real-world networks leads to the conclusion that a single control input, applied to the power dominating set, is all that is needed for structural controllability. This result is consistent with the well-known fact that controllability and its dual observability are generic properties of systems. We argue that more important than issues of structural controllability are the questions of whether a system is almost uncontrollable, whether it is almost unobservable, and whether it possesses almost pole-zero cancellations.
We study the population genetics of Evolution in the important special case of weak selection, in which all fitness values are assumed to be close to one another. We show that in this regime natural selection is tantamount to the multiplicative updates game dynamics in a coordination game between genes. Importantly, the utility maximized in this game, as well as the amount by which each allele is boosted, is precisely the allele's mixability, or average fitness, a quantity recently proposed in [1] as a novel concept that is crucial in understanding natural selection under sex, thus providing a rigorous demonstration of that insight. We also prove that the equilibria in two-person coordination games can have large supports, and thus genetic diversity does not suffer much at equilibrium. Establishing large supports involves answering through a novel technique the following question: what is the probability that for a random square matrix A both systems Ax = 1 and A^T y = 1 have positive solutions? Both the question and the technique may be of broader interest. [1] A. Livnat, C. Papadimitriou, J. Dushoff, and M.W. Feldman. A mixability theory for the role of sex in evolution. Proceedings of the National Academy of Sciences, 105(50):19803-19808, 2008.
One strategy for winning a coevolutionary struggle is to evolve rapidly. Most of the literature on host-pathogen coevolution focuses on this phenomenon, and looks for consequent evidence of coevolutionary arms races. An alternative strategy, less often considered in the literature, is to deter rapid evolutionary change by the opponent. To study how this can be done, we construct an evolutionary game between a controller that must process information, and an adversary that can tamper with this information processing. In this game, a species can foil its antagonist by processing information in a way that is hard for the antagonist to manipulate. We show that the structure of the information processing system induces a fitness landscape on which the adversary population evolves, and that complex processing logic is required to make that landscape rugged. Drawing on the rich literature concerning rates of evolution on rugged landscapes, we show how a species can slow adaptive evolution in the adversary population. We suggest that this type of defensive complexity on the part of the vertebrate adaptive immune system may be an important element of coevolutionary dynamics between pathogens and their vertebrate hosts.
One strategy for winning a coevolutionary struggle is to evolve rapidly. Most of the literature on host-pathogen coevolution focuses on this phenomenon, and looks for consequent evidence of coevolutionary arms races. An alternative strategy, less often considered in the literature, is to deter rapid evolutionary change by the opponent. To study how this can be done, we construct an evolutionary game between a controller that must process information, and an adversary that can tamper with this information processing. In this game, a species can foil its antagonist by processing information in a way that is hard for the antagonist to manipulate. We show that the structure of the information processing system induces a fitness landscape on which the adversary population evolves. Complex processing logic can carve long, deep fitness valleys that slow adaptive evolution in the adversary population. We suggest that this type of defensive complexity on the part of the vertebrate adaptive immune system may be an important element of coevolutionary dynamics between pathogens and their vertebrate hosts. Furthermore, we cite evidence that the immune control logic is phylogenetically conserved in mammalian lineages. Thus our model of defensive complexity suggests a new hypothesis for the lower rates of evolution for immune control logic compared to other immune structures.
In some decision-making environments, successful solutions are common. If the evaluation of candidate solutions is noisy, however, the challenge is knowing when a “good enough” answer has been found. We formalize this problem as an infinite-armed bandit and provide upper and lower bounds on the number of evaluations or “pulls” needed to identify a solution whose evaluation exceeds a given threshold r0. We present several algorithms and use them to identify reliable strategies for solving screens from the video games Infinite Mario and Pitfall! We show order of magnitude improvements in sample complexity over a natural approach that pulls each arm until a good estimate of its success probability is known.
Noah J. Cowan, Erick J. Chastain, Daril A. Vilhena, James S. Freudenberg, and Carl T. Bergstrom 5 Department of Mechanical Engineering, Johns Hopkins University, Baltimore, MD 21218 Department of Computer Science, Rutgers University, New Brunswick, NJ 08901 Department of Biology, University of Washington, Seattle, WA 98105 Department of Electrical Engineering and Computer Science, University of Michigan, Ann Arbor, MI 48109 Santa Fe Institute, 1399 Hyde Park Rd., Santa Fe, NM 87501 (Dated: July 20, 2011)
Structural controllability has been proposed as an analytical framework for making predictions regarding the control of complex networks across myriad disciplines in the physical and life sciences (Liu et al., Nature:473(7346):167-173, 2011). Although the integration of control theory and network analysis is important, we argue that the application of the structural controllability framework to most if not all real-world networks leads to the conclusion that a single control input, applied to the power dominating set (PDS), is all that is needed for structural controllability. This result is consistent with the well-known fact that controllability and its dual observability are generic properties of systems. We argue that more important than issues of structural controllability are the questions of whether a system is almost uncontrollable, whether it is almost unobservable, and whether it possesses almost pole-zero cancellations.
We propose a new view of active learning algorithms as optimization. We show that many online active learning algorithms can be viewed as stochastic gradient descent on non-convex objective functions. Variations of some of these algorithms and objective functions have been previously proposed without noting this connection. We also point out a connection between the standard min-margin offline active learning algorithm and non-convex losses. Finally, we discuss and show empirically how viewing active learning as non-convex loss minimization helps explain two previously observed phenomena: certain active learning algorithms achieve better generalization error than passive learning algorithms on certain data sets (Schohn and Cohn, 2000; Bordes et al., 2005) and on other data sets many active learning algorithms are prone to local minima (Sch¨ utze et al., 2006). 1 Background We address the active learning problem in this paper. We assume data points X 2 R d and labels Y 2 fi1;1g are drawn from some fixed unknown distribution p(x;y). We wish to choose the classifier f 2 F that minimizes the expected loss EX;Y [l(Y;f;X)] where l is a loss function. In general, the size of F may be uncountable. In standard passive supervised learning we approximately minimize this expected loss by minimizing instead the loss on a set or stream of i.i.d. labeled training examples (xi;yi) » p(x;y). In active learning, we also have access to i.i.d. xi, but we selectively query the labels yi for the examples. In
General navigation requires a spatial map that is not anchored to one environment. The firing fields of the “grid cells” found in the rat dorsolateral medial entorhinal cortex (dMEC) could be such a map. dMEC firing fields are also thought to be modeled well by a regular triangular grid (a grid with equilateral triangles as units). We use computational means to analyze and validate the regularity of the firing fields both quantitatively (using summary statistics for geometric and photometric regularity) and qualitatively (using symmetry group analysis). Upon quantifying the regularity of real dMEC firing fields, we find that there are two types of grid cells. We show rigorously that both are nearest to triangular grids using symmetry analysis. However, type III grid cells are far from regular, both in firing rate (highly non-uniform) and grid geometry. Type III grid cells are also more numerous. We investigate the implications of this for the role of grid cells in path integration.
The firing fields of the ”grid cells” found in the rat dorsocaudal medial entorhinal cortex (dMEC) present a surprising pattern of context-independent regularity. We use computational means to analyze and validate the geometric and algebraic invariant properties of the firing fields, leading to a context invariant spatial map. Our method computes the specific symmetry group implicitly associated with the spatial map, and quantifies the regularity of the firing fields to achieve a symmetry-based clustering of two different types of “grid cells.” This quantified regularity makes spatial mapping more computationally efficient and suggests a way to use the dMEC firing patterns to estimate the probabilty of the rat being located at different points in the room. Finally, general properties of context-independent population codes are suggested. Namely, context-independent population codes gain robustness by introducing uncertainty and ambiguity.