Phylogenetic reconstructions are a primary record of protein evolution. But what other records can attest to the deep history of enzymes, and what tools are needed to decode their meaning? Here, we demonstrate that the history of enzyme discovery and reuse is embedded within the web of interdependencies that constitute contemporary, biosphere-scale metabolism. Using a simple network analysis approach, we reconstruct both the relative temporal ordering of enzyme domain emergence and, where possible, the first reactions that they catalyzed. These network-based histories were found to be broadly concordant with phyletic information, suggesting that the two approaches reflect related generative processes. When enzyme emergence is initiated after the discovery of nucleotide cofactors, a predominantly stepwise trajectory of domain discovery is recovered. We find that the earliest enzyme-mediated metabolisms were dominated by α/β domains, likely due to their high discoverability and functional potential under constraint. Finally, we quantify how the protein universe responded to a major transition, the biological production of molecular oxygen, by preferentially reusing pre-existing enzyme domains. This work presents a self-consistent model of metabolic and enzyme evolution, essential progress towards integrating multiple, independent records into a unified history of protein evolution. ### Competing Interest Statement The authors have declared no competing interest. International Human Frontier Science Program Organization (HFSP), Ref.-No: RGEC29/2025 National Aeronautics and Space Administration (NASA), 80NSSC23K1357, 80NSSC25K7873
One of the puzzles left open by energetic analyses of irreversible stochastic processes is that boundary conditions that prevent the performance of work or the dissipation of heat make no contribution to an entropy-production budget; yet we see ubiquitously in both engineered and living systems that both transient and persistent energy costs are paid to create and maintain such boundaries. We wish to know whether there are inherent limits for the costs of such phenomena, and common units in which those can be traded off against more familiar costs measured in terms of entropy production and heat dissipation. We give this problem a concrete framing in the context of chemical reaction networks (CRNs), for the problem of extracting a topologically restricted pathway from a larger distributed network, through the activation of some reactions and the selective elimination of others. We define a thermodynamic cost function for pathways derived from the large-deviation theory of stochastic CRNs, which decomposes into two components: an ongoing maintenance cost to sustain a non-equilibrium steady state, and a restriction cost, quantifying the ongoing improbability of neutralizing reactions outside the specified pathway. Applying this formalism to detailed-balanced CRNs in the linear response regime, we make use of their formal equivalence to electrical circuits. We prove that the resistance of a CRN decreases as reactions are added that support the throughput current, and that the maintenance cost, the restriction cost, and the thermodynamic cost of nested pathways are bounded below by those of their hosting network. For four- and five-species example CRNs, we show how catalytic and inhibitory mechanisms can drastically alter pathway costs, enabling the unfavorable pathways to become favorable and to approach the cost of the hosting pathway. Our results provide insights into the thermodynamic principles governing open CRNs and offer a foundation for understanding the evolution of metabolic networks.
We consider the problem of evaluating and expanding upon speculative hypotheses about the origin of life. The combination of life's complexity and its potential for historical contingency makes missing knowledge and missing ideas, which we term 'gaps', obstructive for an unusually wide variety of its most basic questions. The methods of scientific empiricism developed to justify beliefs in mature and stable sciences have proved less useful for reasoning through gaps, where criteria of consistency may rarely be met, and weaker criteria such as metaphor serve as motivations in practice. We consider the particular role of scenarios in the justification of speculative hypotheses, as they relate to questions of chance and necessity and the sources of causation. We demonstrate how making causal frameworks explicit may support more systematic reasoning about speculative and fragmentary hypotheses than the scenarios in which they are often framed.This article is part of the theme issue 'Origins of life: the possible and the actual'.
We approach the questions, what part of evolutionary change results from selection, and what is the adaptive information flow into a population undergoing selection, as a problem of quantifying the divergence of typical trajectories realized under selection from the expected dynamics of their counterparts under a null stochastic-process model representing the absence of selection. This approach starts with a formulation of adaptation in terms of information and from that identifies selection from the genetic parameters that generate information flow; it is the reverse of a historical approach that defines selection in terms of fitness, and then identifies adaptive characters as those amplified in relative frequency by fitness. Adaptive information is a relative entropy on distributions of histories computed directly from the generators of stochastic evolutionary population processes, which in large population limits can be approximated by its leading exponential dependence as a large-deviation function. We study a particular class of generators that represent the genetic dependence of explicit transitions around reproductive cycles in terms of stoichiometry, familiar from chemical reaction networks. Following Smith (2023), which showed that partitioning evolutionary events among genetically distinct realizations of lifecycles yields a more consistent causal analysis through the Price equation than the construction from units of selection and fitness, here we show that it likewise yields more complete evolutionary information measures.
We first derive the Hamilton-Jacobi theory underlying continuous-time Markov processes, and then we use the construction to develop a variational algorithm for estimating escape (least improbable or first passage) paths for a generic stochastic chemical reaction network that exhibits multiple fixed points. The design of our algorithm is such that it is independent of the underlying dimensionality of the system, the discretization control parameters are updated toward the continuum limit, and there is an easy-to-calculate measure for the correctness of its solution. We consider several applications of the algorithm and verify them against computationally expensive means such as the shooting method and stochastic simulation. While we employ theoretical techniques from mathematical physics, numerical optimization and chemical reaction network theory, we hope that our work finds practical applications with an inter-disciplinary audience including chemists, biologists, optimal control theorists and game theorists.
We address the problem of defining selection and extracting the adaptive part of evolutionary change, originally formalized by Fisher and Price. Conventionally, selection and adaptation are defined through fitness attributed to genes or genotypes chosen as units of selection. The construction through fitness is known to suffer ambiguities and omissions as a theory of change due to selection. We construct an alternative framing in which units of selection and fitness are replaced as the main abstractions by formal lifecycle models and reproduction rates through genetically distinct lifecycle realizations. Graphical representations of lifecycles express relations among reproductive stages that cannot be assigned to any one unit of selection. The lifecycle partition refines the statistics of overall reproductive success and resolves modes of selection that fitness either excludes or distorts through additive projections. We derive the Price equation in the basis of lifecycle realizations and compare it to the conventional Price equation for additive fitness of organisms. We show how the lifecycle approach recovers fitnesses acting concurrently at multiple levels, or contrasts forms of competition within and between levels that are invisible to additive fitness. Defining selection through lifecycles recasts population genetics from an object-focused to a construction- and process-focused representation.
We have developed the program TwinCons, to detect noisy signals of deep ancestry of proteins or nucleic acids. As input, the program uses a composite alignment containing pre-defined groups, and mathematically determines a ‘cost’ of transforming one group to the other at each position of the alignment. The output distinguishes conserved, variable and signature positions. A signature is conserved within groups but differs between groups. The method automatically detects continuous characteristic stretches (segments) within alignments. TwinCons provides a convenient representation of conserved, variable and signature positions as a single score, enabling the structural mapping and visualization of these characteristics. Structure is more conserved than sequence. TwinCons highlights alternative sequences of conserved structures. Using TwinCons, we detected highly similar segments between proteins from the translation and transcription systems. TwinCons detects conserved residues within regions of high functional importance for the ribosomal RNA (rRNA) and demonstrates that signatures are not confined to specific regions but are distributed across the rRNA structure. The ability to evaluate both nucleic acid and protein alignments allows TwinCons to be used in combined sequence and structural analysis of signatures and conservation in rRNA and in ribosomal proteins (rProteins). TwinCons detects a strong sequence conservation signal between bacterial and archaeal rProteins related by circular permutation. This conserved sequence is structurally colocalized with conserved rRNA, indicated by TwinCons scores of rRNA alignments of bacterial and archaeal groups. This combined analysis revealed deep co-evolution of rRNA and rProtein buried within the deepest branching points in the tree of life.
The objective of this paper is to analyze and characterize high-resolution measurements of geometric imperfections taken from a set of seven slender tapered steel tubes in order to provide insights that can improve methods for predicting buckling behavior. The seven tubes are each similar to 3400 min long with diameters between 800 mm and 1100 mm, diameter to thickness ratios between 300 and 350, and taper angles between 0.67 degrees and 0.86 degrees. The tubes are manufactured from steel plates using an innovative spiral welding process. The geometric imperfections of these tubes are characterized with harmonic analysis of the overall imperfection measurements and with regression analysis of the measured shapes of weld depressions. The results show a consistent imperfection signature caused by the manufacturing process including distinct features attributed to both the rolling and welding processes, i.e. anticlastic deformations and weld depressions. Variability in the imperfection measurements is also analyzed and used to generate a probabilistic scheme capable of generating random fields of geometric imperfections that are consistent with the measurements considered here.
A set of core features is set forth as the essence of a thermodynamic description, which derive from large-deviation properties in systems with hierarchies of timescales, but which are not dependent upon conservation laws or microscopic reversibility in the substrate hosting the process. The most fundamental elements are the concept of a macrostate in relation to the large-deviation entropy, and the decomposition of contributions to irreversibility among interacting subsystems, which is the origin of the dependence on a concept of heat in both classical and stochastic thermodynamics. A natural decomposition that is known to exist, into a relative entropy and a housekeeping entropy rate, is taken here to define respectively the intensive thermodynamics of a system and an extensive thermodynamic vector embedding the system in its context. Both intensive and extensive components are functions of Hartley information of the momentary system stationary state, which is information about the joint effect of system processes on its contribution to irreversibility. Results are derived for stochastic chemical reaction networks, including a Legendre duality for the housekeeping entropy rate to thermodynamically characterize fully-irreversible processes on an equal footing with those at the opposite limit of detailed-balance. The work is meant to encourage development of inherent thermodynamic descriptions for rule-based systems and the living state, which are not conceived as reductive explanations to heat flows.
The objective of this study is to develop and validate a practical finite-element modeling protocol for predicting the flexural strength and collapse behavior of thin-walled spirally welded tapered tubes that can be used as steel wind turbine towers. The overall modeling protocol consists of two parts: (1)a meshing protocol is developed considering the effects shell element type, aspect ratio, inclination angle, and density on the buckling moment relative to theoretical predictions; and (2)two patterns of geometric imperfections (eigenmode-affine and so-called weld depression) scaled to the thresholds of fabrication tolerance quality classes in Eurocode 3 are considered in nonlinear collapse shell finite-element models and the results are compared to a series of eight large-scale flexural tests of spirally welded tubes. The computational results are compared with test results in terms of moment-rotation response, stiffness, and buckling modes and show sufficient agreement to justify the further development of nonlinear analysis methods for the design of steel wind turbine towers made from thin-walled spirally welded tapered tubes.
The increasing availability of large digital corpora of cross-linguistic data is revolutionizing many branches of linguistics. Overall, it has triggered a shift of attention from detailed questions about individual features to more global patterns amenable to rigorous, but statistical, analyses. This engenders an approach based on successive approximations where models with simplified assumptions result in frameworks that can then be systematically refined, always keeping explicit the methodological commitments and the assumed prior knowledge. Therefore, they can resolve disputes between competing frameworks quantitatively by separating the support provided by the data from the underlying assumptions. These methods, though, often appear as a ‘black box’ to traditional practitioners. In fact, the switch to a statistical view complicates comparison of the results from these newer methods with traditional understanding, sometimes leading to misinterpretation and overly broad claims. We describe here this evolving methodological shift, attributed to the advent of big, but often incomplete and poorly curated data, emphasizing the underlying similarity of the newer quantitative to the traditional comparative methods and discussing when and to what extent the former have advantages over the latter. In this review, we cover briefly both randomization tests for detecting patterns in a largely model-independent fashion and phylolinguistic methods for a more model-based analysis of these patterns. We foresee a fruitful division of labor between the ability to computationally process large volumes of data and the trained linguistic insight identifying worthy prior commitments and interesting hypotheses in need of comparison.
The conditions for life in the deep past were different in many important ways from conditions of life on Earth today. Simple sequence-based comparative methods of evolutionary reconstruction, applied to single genes or proteins, can therefore give an incomplete or misleading picture of ancestral life forms. One way to augment sequence-based, single-gene methods to obtain a richer and more reliable picture of the deep past, is to resurrect inferred ancestral protein sequences in living organisms, where their phenotypes can be exposed in a complex molecular-systems context, and to then link consequences of those phenotypes to biosignatures that were preserved in the independent historical repository of the geological record. Good candidates for such ‘revenant gene’ studies are the genes for enzymes involved in carbon-fixation or other core metabolic pathways. Resurrecting ancestral DNA using synthetic-biology methods to engineer modern host bacteria is just at the beginning of its use as a systematic method in evolutionary biology. It has great potential to refine our understanding of the historical era already probed by phylogenetic methods, and even to suggest the forces governing the assembly of living systems reaching back into the “pre-historic” past, before those sequence divergences that have left descendants into the modern era. However, good design of revenant gene studies introduces new and interesting problems in the selection of genes, biosignatures, and modern host organisms, which should be understood as part of the next step in advancement of evolutionary methods.
Both inorganic and hybrid (organo-inorganic) perovskite materials are potential candidates as photocatalysts for use in both photovoltaic (PV) and photocatalytic water splitting applications. Currently, research has been focused on specifically designing perovskite materials so they can harness the broad spectrum of the visible light wavelength. Inorganic perovskites such as titanates, tantalates, niobates, and ferrites show great promise as visible light-driven photocatalysts for water splitting, whereas hybrid perovskites such as methylammonium lead halides reveal unique photovoltaic and charge transport properties. The main objective of this article is to examine the progress on some recent research on perovskite nanomaterials for both solar cell and water splitting applications. This mini review paper summarizes some recent developments of organic and inorganic perovskite materials (PMs) and provides useful insights for their future improvement.
The study of chemical reaction networks (CRN's) is a very active field. Earlier well-known results (Feinberg 1987 Chem. Enc. Sci. 42 2229, Anderson et al 2010 Bull. Math. Biol. 72 1947) identify a topological quantity called deficiency, for any CRN, which, when exactly equal to zero, leads to a unique factorized steady-state for these networks. No results exist however for the steady states of non-zero-deficiency networks. In this paper, we show how to write the full moment-hierarchy for any non-zero-deficiency CRN obeying mass-action kinetics, in terms of equations for the factorial moments. Using these, we can recursively predict values for lower moments from higher moments, reversing the procedure usually used to solve moment hierarchies. We show, for nontrivial examples, that in this manner we can predict any moment of interest, for CRN's with non-zero deficiency and non-factorizable steady states.
The most common wind tower structure, a tapered tubular steel monopole, is currently limited to heights of ~80m due to transportation constraints which arise because tower sections are manufactured at centralized plants and transported to site for assembly. The need to transport the sections imposes a limit on their size, whereby maximum tower diameters are dictated by bridge clearances rather than by structural efficiency. New manufacturing innovations, based on automated spiral welding, may enable on-site production of wind towers, thereby precluding transportation limits and permitting the manufacture of taller towers, which can harvest the steadier, stronger winds at higher elevations. Taller towers, however, are expected to have cross-sections with slenderness that is uncommon in structural engineering (i.e., diameter-to-thickness ratios up to ~500) and much larger than those of conventionally manufactured towers (i.e., diameter-to-thickness ratios up to ~300). Tubular structures with highly slender cross-sections are imperfection-sensitive, and the welding process is known to influence imperfections. To account for this sensitivity, slender tubes are usually designed based on empirical knockdown factors, however there are few experiments of tubes in flexure with slenderness as high as what is expected for spirally welded wind towers, and there are no experiments on tubes within this slenderness range and manufactured with spiral welding. This paper reviews the state-of-the-art for designing spirally welded tubes as wind towers and identifies deficiencies. Relevant experimental and analytical research is summarized and research needs to efficiently design tapered spirally welded steel tubes as wind towers are identified.
Chapter 11 raises the question of what is meant by our usage of “theory”. Different disciplines utilize the word theory differently. Furthermore model and theory appear on occasion to be used interchangeably. Aristotle contrasted theory to practice. Praxis is the Greek term for doing. Mathematical theory is deductive. The sensory or empirical content is implicit in the axioms. The logical consequences of the axioms provide theorems. A semantic view stresses the connection between the axioms and the abstraction of some aspect of reality. We stress that the natural preliminary step before dynamics is to construct process models based on general equilibrium. This can be done utilizing single simultaneous move games. This is sufficient to show the critical roles of money and financial institutions without even having to discuss complication in information and behaviour. The evolution of money and many financial institutions does not even call for the presence of exogenous uncertainty. A single random variable is sufficient to illustrate innovation. We develop a general modeling methodology leading to the construction of models as playable games. Staying with the one move structure leads to describing a manageable number of minimal institutions (below 100). When we consider more moves and information the number of logically feasible and plausible institutions becomes hyperastronomical and we are forced into considering not merely structure but many variants of behaviour even within the simple scope of rational expectations. This problem is taken up in Chapter 12.
This final chapter splits naturally into two parts. The first part presents the basic overview of a theory of money and financial institutions that covers the abstract pre-institutional structure of the price system as an allocation device in a tightly defined structure with no need for money or financial institutions to be specified in the illustration of the efficient equilibrium condition. It stresses that by merely following well defined precepts of basic physics and game theory, at the same level of abstraction process models can be completed and force a discipline on the models where items such as money, default laws, grid size, loose coupling, specification of clearance rules, and time lags are all logical necessities.. These steps convert a static pre-institutional set of models into their natural minimal institution monetary and institutional models. . The utilization of these for application still requires the addition of ad hoc physical facts and relevant understanding of behavior. Our broader goal is directed towards a general theory of organization about which this work is only a narrow part. The second part of this chapter concludes with our Nostrodamus section where we conjecture about the future. This includes the need for supranational organizations especially for weapons control. We also suggest that the fundamental problem of the control of dangerous fluctuations in an enterprise economy is predominately a politico-bureaucratic problem and calls for a stress on the design of flexible coordinating devices within the politico-economic system. A sketch of such a device is presented. In a free society the stress should be less on control and forecasting than on guidance and flexibility.
A new manufacturing process allows for the production of tapered spirally welded steel tubes. This paper describes a series of eight large-scale tests examining the behavior of such tubes in flexure and investigates the impact of imperfections on the flexural strength of the tube. The tests are performed on tapered circular steel tubes with diameters between 0.7 and 1.1m and maximum diameter-to-thickness ratios between 200 and 350. Specimen geometries are selected to provide flexural test data at slenderness ratios not commonly tested in the literature and to be representative of tapered tubes applied as wind turbine towers. The geometries of the specimens are measured with laser scanners before and during testing to characterize initial imperfections and the evolution of local buckling. Results are compared to design strengths per Eurocode EN 1993 1-6. All specimens meeting Eurocode manufacturing quality requirements exceed predicted strengths. The location and orientation of the local buckling region are correlated with the spiral seam welds on the specimens.
In this chapter we introduce our first multi-timescale models of an economy. Longer timescales are associated with commitment to product specialization and to discounting and depreciation of durable goods. We compare an economy mediated by a durable commodity money, such as gold, with one with a fiat money implemented through a bureaucracy. The inefficiency costs associated with asymmetric strategic roles between money-providers and producers of consumption goods are compared with explicit losses of material productivity due to labor costs required to maintain a bureaucracy needed to manage a fiat money system. It is shown how a stable trade system can emerge. We discuss material and institutional capital stock in an economy and note that capital stock acts much like the stock of catalysts used to enable chemical reactions. In a framework such as General Equilibrium, capital stock vanishes from the net input-output relations in an abstract production correspondence.