The origin of eukaryotes represents one of the most significant events in evolution since it allowed the posterior emergence of multicellular organisms. Yet, it remains unclear how existing regulatory mechanisms of gene activity were transformed to allow this increase in complexity. Here, we address this question by analyzing the length distribution of proteins and their corresponding genes for 6,519 species across the tree of life. We find a scale-invariant relationship between gene mean length and variance maintained across the entire evolutionary history. Using a simple model, we show that this scale-invariant relationship naturally originates through a simple multiplicative process of gene growth. During the first phase of this process, corresponding to prokaryotes, protein length follows gene growth. At the onset of the eukaryotic cell, however, mean protein length stabilizes around 500 amino acids. While genes continued growing at the same rate as before, this growth primarily involved noncoding sequences that complemented proteins in regulating gene activity. Our analysis indicates that this shift at the origin of the eukaryotic cell was due to an algorithmic phase transition equivalent to that of certain search algorithms triggered by the constraints in finding increasingly larger proteins.
The subjective number-space mapping and, especially, its evolution in young children has been the subject of intense controversy among different competing models. Many studies point out that: (i) young children's innate estimates follow a logarithmic mapping (Weber-Fechner law) and (ii) driven by education, children evolve into a linear mapping. In this paper we show, in consonance with other investigations, that innate numerical intuitions of young children are in fact a linear mapping, and not a logarithmic one. We found that young children extrapolate linearly from a reference point. The apparent logarithmic mapping is an artifact produced by placing a huge range of numbers within a small number line. We show how the appearance of the mapping can be modulated by changing the size and contour conditions of the number lines used in experiments.
Haros graphs have been recently introduced as a set of graphs bijectively related to real numbers in the unit interval. Here we consider the iterated dynamics of a graph operator R over the set of Haros graphs. This operator was previously defined in the realm of graph-theoretical characterization of low-dimensional nonlinear dynamics and has a renormalization group (RG) structure. We find that the dynamics of R over Haros graphs is complex and includes unstable periodic orbits of arbitrary period and nonmixing aperiodic orbits, overall portraiting a chaotic RG flow. We identify a single RG stable fixed point whose basin of attraction is associated with the set of rational numbers, and find periodic RG orbits that relate to (pure) quadratic irrationals and aperiodic RG orbits, related with (nonmixing) families of nonquadratic algebraic irrationals and transcendental numbers. Finally, we show that the graph entropy of Haros graphs is globally decreasing as the RG flows towards its stable fixed point, albeit in a strictly nonmonotonic way, and that such graph entropy remains constant inside the periodic RG orbit associated to a subset of irrationals, the so-called metallic ratios. We discuss the possible physical interpretation of such chaotic RG flow and put results regarding entropy gradients along RG flow in the context of c-theorems.
This paper introduces Haros graphs, a construction which provides a graph-theoretical representation of real numbers in the unit interval reached via paths in the Farey binary tree. We show how the topological structure of Haros graphs yields a natural classification of the reals numbers into a hierarchy of families. To unveil such classification, we introduce an entropic functional on these graphs and show that it can be expressed, thanks to its fractal nature, in terms of a generalised de Rham curve. We show that this entropy reaches a global maximum at the reciprocal of the Golden number and otherwise displays a rich hierarchy of local maxima and minima that relate to specific families of irrationals (noble numbers) and rationals, overall providing an exotic classification and representation of the reals numbers according to entropic principles. We close the paper with a number of conjectures and outline a research programme on Haros graphs.
A lo largo del siglo XX los estudios en lingüística cuantitativa han ido mostrando la aparición de leyes potenciales en las lenguas, primero en textos escritos y posteriormente en el habla. Son leyes que parecen ubicuas y robustas, pero ¿por qué aparecen en el lenguaje? ¿Son resultados espurios debidos a la arbitrariedad de la segmentación de las palabras, o realmente son universales de la comunicación compleja? ¿Podemos investigar la presencia de estas leyes en otros sistemas de comunicación animal de los que no conocemos el código? Los enfoques interdisciplinares y transdisciplinares en la lingüística y el estudio de los sistemas de comunicación se antojan imprescindibles. Se exponen a modo de ejemplo dos estudios recientes realizados sobre corpus acústicos de hasta dieciséis lenguas, mediante un método general de segmentación de señales (método de los umbrales). Exploramos aquí la posibilidad de que las leyes estadísticas que emergen en el lenguaje sean fruto de un sistema crítico auto-organizado, al igual que otros fenómenos presentes en la Naturaleza. El método de los umbrales que se presenta permite analizar cualquier tipo de señal sin necesidad de conocer su codificación o segmentación. Esto abre nuevos caminos en la investigación lingüística permitiendo entre otras cosas realizar estudios comparativos entre el lenguaje humano y otros sistemas de comunicación animal.
This paper discusses the applicability of Visibility Algorithms to detect faults in condition monitoring applications. The general purpose of Visibility Algorithms is to transform time series into graphs and study them through the characterization of their associated network. Degradation of a component results in changes to the network. This technique has been applied using a test rig of an aircraft fuel system to show that there is a correlation between the values of key metrics of visibility graphs and the severity of four failure modes. We compare the results of using Horizontal Visibility algorithms against Natural Visibility algorithms. The results also show how the Kullback-Leibler divergence and statistical entropy can be used to produce condition indicators. Experimental results show that there is little dispersion in the values of condition indicators, leading to a low probability of false positives and false negatives.
Tell Bartolo Luque and Fernando Ballesteros how far the Sun is from the Earth, and they will tell you the size of the Universe.
Additional details on the analysis of Buckeye Corpus and additional results such as: (II.a) Lognormality law for individual speakers, (II.b) Additional representations on the duration distribution of linguistic units, (II.c) Limit distributions of sums of independent Lognormals, (II.d) The error terms at word and BG levels, (III.a) Model selection scheme for Zipf's law, (III.b) Null model for Zipf's law, (III.c) Zipf's law for individual informants, (IV) Herdan's law: additional evidence based on speech rate, (V.a) Binning in scatter plots with high noise in one of the variables, (V.b) Binning in frequencies, (VI.a) Additional fits of Menzerath-Altmann law and (VI.b) Predicting BG time duration distribution using the Menzerath-Altmann law model
Physical manifestations of linguistic units include sources of variability due to factors of speech production which are by definition excluded from counts of linguistic symbols. In this work, we examine whether linguistic laws hold with respect to the physical manifestations of linguistic units in spoken English. The data we analyse come from a phonetically transcribed database of acoustic recordings of spontaneous speech known as the Buckeye Speech corpus. First, we verify with unprecedented accuracy that acoustically transcribed durations of linguistic units at several scales comply with a lognormal distribution, and we quantitatively justify this 'lognormality law' using a stochastic generative model. Second, we explore the four classical linguistic laws (Zipf's Law, Herdan's Law, Brevity Law and Menzerath-Altmann's Law (MAL)) in oral communication, both in physical units and in symbolic units measured in the speech transcriptions, and find that the validity of these laws is typically stronger when using physical units than in their symbolic counterpart. Additional results include (i) coining a Herdan's Law in physical units, (ii) a precise mathematical formulation of Brevity Law, which we show to be connected to optimal compression principles in information theory and allows to formulate and validate yet another law which we call the size-rank law or (iii) a mathematical derivation of MAL which also highlights an additional regime where the law is inverted. Altogether, these results support the hypothesis that statistical laws in language have a physical origin.
Gravitation is one of the main forces of nature and is a fundamental factor for the habitability of the cosmos. Not only it is the main source of structures in a universe that, without it, would be only a faint soup of gases, but it is the main engine for creating stars and worlds, to make these planets inhabitable and to impose structural and functional limits on the organisms that can be developed in them.
We show how the cross-disciplinary transfer of techniques from dynamical systems theory to number theory can be a fruitful avenue for research. We illustrate this idea by exploring from a nonlinear and symbolic dynamics viewpoint certain patterns emerging in some residue sequences generated from the prime number sequence. We show that the sequence formed by the residues of the primes modulo k are maximally chaotic and, while lacking forbidden patterns, unexpectedly display a non-trivial spectrum of Renyi entropies which suggest that every block of size m>1, while admissible, occurs with different probability. This non-uniform distribution of blocks for m>1 contrasts Dirichlet’s theorem that guarantees equiprobability for m=1. We then explore in a similar fashion the sequence of prime gap residues. We numerically find that this sequence is again chaotic (positivity of Kolmogorov–Sinai entropy), however chaos is weaker as forbidden patterns emerge for every block of size m>1. We relate the onset of these forbidden patterns with the divisibility properties of integers, and estimate the densities of gap block residues via Hardy–Littlewood k-tuple conjecture. We use this estimation to argue that the amount of admissible blocks is non-uniformly distributed, what supports the fact that the spectrum of Renyi entropies is again non-trivial in this case. We complete our analysis by applying the chaos game to these symbolic sequences, and comparing the Iterated Function System (IFS) attractors found for the experimental sequences with appropriate null models.