
Collective intelligence–the ability of groups to solve diverse problems–has been explored using laboratory experiments, computer simulations, and questionnaires. These instruments, however, suffer from limitations, such as external validity in the case of laboratory experiments and self-reporting bias in the case of questionnaires. Here we investigate the exploration-exploitation dynamics of small teams using high-frequency, observational data from escape rooms: a non-interventional yet controlled environment where naturally occurring teams solve exploration and exploitation tasks. We find that more effective teams tend to coordinate throughout problem-solving, exhibit balanced communication patterns, and are more responsive, addressing tasks promptly as they become solvable rather than accumulating them. In contrast, members of less effective teams often work in isolation, participate in problem-solving unequally, and tend to accumulate tasks rather than addressing them as they become solvable. Importantly, we show that no single collaborative structure works for every task: exploitation is faster under team-wide communication and may even benefit from the dominance of key members, while efficient exploration aligns with balanced participation. Additionally, positive exchanges are associated with faster exploitation but slower exploration. These findings expand the external validity of experimental work on collective intelligence to an organic non-interventional setting and highlight the importance of understanding team performance through behavior instead of team demographics.
Modeling emergent, multiscale patterns in complex systems remains a persistent interdisciplinary challenge, particularly when deciphering transient, highly coupled dynamics from severely limited data. To overcome the inherent spectral sparsity of standard finite-dimensional stochastic models, we introduce Multilayered Stochastic Hierarchical Delay Models (MSHDMs). By embedding dynamics within a hierarchical delay structure, MSHDMs leverage an infinite-dimensional phase space to engineer the high spectral density required for generating complex amplitude-frequency modulation (AM-FM) dynamics, avoiding data-heavy neural network parameterizations. To reliably calibrate these sensitive structures from short observational records, we deploy a hybrid offline-online Bayesian optimization framework. Bypassing the failure points of traditional trajectory matching, our algorithm autonomously learns optimal latent coordinates and continuous fractional delays by strictly enforcing spectral consistency against the empirical Global Wavelet Spectrum. Applying this methodology to high-resolution satellite observations of continental cloud fields, the resulting stochastic emulator captures the full spatiotemporal coherence of the turbulent system using just six hyperparameters. The wavelet scalograms demonstrate how the model natively generates emergent wave-packet dynamics and cross-scale energy cascades, recovering semidiurnal, mesoscale, and individual cloud timescales. Supported by rigorous mathematical foundations and robust data-driven calibrations, MSHDMs thus provide a highly compressible, general-purpose tool for resolving latent AM-FM beat patterns across diverse disciplines.
Stochastic simulations underpin computational research in fields from systems biology and epidemiology to finance and physics-informed machine learning, yet their reproducibility is hard to quantify because each run yields different outcomes. Existing practices to improve computational reproducibility focus on sharing code, simulation seeds, data, and software environments but do not ensure independent reproduction of results, limiting scientific rigor. Here we introduce the Empirical Characteristic Function Equality Convergence Test (EFECT), a universal computational framework for quantifying the statistical reproducibility of stochastic simulations based on empirical characteristic functions. EFECT defines a normalized EFECT error that measures distributional differences between two sets of simulation outputs over a model- and scale-independent range, and an EFECT convergence point that specifies the minimum sample size required to achieve a desired reproducibility threshold at a chosen significance level. EFECT applies to any simulation whose results can be represented as bounded, real-valued data, regardless of modeling formalism or source of stochasticity. To facilitate widespread adoption and data exchange, we implemented EFECT in an open-source software library and a compact, machine-readable standardized report. We rigorously evaluated the framework across more than 40 test cases spanning stochastic differential equations, agent-based models, Boolean networks, and partial differential equations. Applying EFECT to published models in pandemic epidemiology, financial stochastic volatility, and physics-informed neural networks reveals that commonly used sample sizes are often underpowered: substantial parameter differences can go undetected unless thousands of replicates are simulated. By providing a domain-agnostic statistical reproducibility metric, reusable software libraries, and a standardized report format, EFECT offers a practical framework for embedding quantitative reproducibility and replicability checks into stochastic simulation workflows across data-intensive disciplines. Thus, EFECT delivers the capability to make scientific inferences and policy decisions on the basis of reliable, reproducible stochastic simulations.
Algorithmic fairness research has largely emphasized model-level metrics, despite evidence that algorithmic systems behave differently once deployed in real-world, complex health systems. This paper reframes fairness as a system-level property and argues that integrating complexity science, implementation research, and sociotechnical perspectives is essential to addressing a central unresolved question: whether fairness-informed algorithms can sustainably narrow health inequalities over time.
Assembly Theory (AT) and its central measure, the assembly index (Ai), provide an opportunity to clarify persistent issues at the interface of computability, compression, and complexity in science. Presented as a novel measure of molecular complexity and as an explanation for selection and evolution, AT can instead be analysed and reproduced using established results from classical information theory, statistical compression, and algorithmic complexity. We show that the latest defence of AT relies on experiments that are incomplete as controls and do not establish the separation between Ai and standard compression-based measures. When extended and compared against established baselines, these experiments reveal substantial overlap in both concept and method with dictionary-based statistical compression algorithms, consistent with our previous formal proof that Ai is mathematically equivalent to the Shannon entropy rate through dictionary-based compression, a result that remains unchallenged. Through theoretical and empirical analysis, we show that Ai offers no new causal or informational insight beyond what existing statistical indices already offer. Rather than establishing the broad explanatory role claimed for AT, the evidence currently supports a more limited interpretation of Assembly Theory as a restricted version of a statistical-compression approach and related measure. We also show that Ai is a special case of an earlier and stronger metric grounded in algorithmic complexity, based on decomposing objects into causal blocks. Finally, we identify multiple technical problems in AT’s computational time-complexity arguments.
Understanding the risk from applications of artificial intelligence (AI) is a critical part of creating AI governance strategies. Building on the idea of studying AI using ecological and evolutionary perspectives, we propose a novel approach for assessing risk from AI using indicators derived from theoretical ecology models. We illustrate our methods by deriving 3 indicators from population and ecosystem models originating from theoretical ecology. We conclude with a discussion of limitations of our analysis and considerations for improving AI governance policy.
How does organized structure arise from dynamics—and how does it, once formed, reshape the dynamics that sustain it? This question lies at the core of complex systems and remains unresolved. While modern theory provides powerful tools to describe equilibrium, fluctuations, and steady processes, it falls short of explaining how structure and dynamics co-evolve in open, evolving systems. Variational approaches—spanning action principles, stochastic path formulations, information-theoretic methods, and thermodynamic extremal principles—offer a unifying perspective based on selection under constraints. Yet these frameworks remain fragmented, and their domains of validity and connections are not fully understood. This special collection brings these perspectives into dialogue, with the aim of advancing a deeper, more integrated understanding of how complex systems organize, adapt, and evolve.
The evolution of novelty is central to understanding the emergence of life. We explored how selective contexts promote innovation, using an empirical model for pre-cellular evolution: in vitro selection of single-stranded DNA with two binding substrates, streptavidin-coated magnetic beads and yeast cells. Evolution proceeded in response to selective pressures for both binding ability and replication efficiency, and variable regimes were incorporated to investigate responses across eco-evolutionary contexts. We found that sequence diversity decreases over eight rounds of selection with beads, corresponding with an enrichment of particular sequences. During selection with cells, diversity was not only maintained but generated, producing long, complex sequences through the recombination of separate genotypes. Low selection stringency facilitated the emergence of these recombinant sequences by allowing diversity to persist, particularly during variable selection. These results demonstrate how the maintenance of diversity provides opportunities for ecological interactions which promote the emergence of innovation, transforming a pre-cellular landscape.
Predicting company growth is a critical yet challenging task because observed dynamics blend an underlying structural growth with volatile fluctuations. Here, we propose a Scaling-Theory-Informed Machine Learning framework (STIML) that integrates a scaling-based model that predicts the mechanism-driven average growth, together with a data-driven forecasting model to learn the residual fluctuations. Using Compustat data of 31,553 North American companies, we extend the growth model to multiple financial indicators, and evaluate STIML against growth model-only and purely data-driven baselines. Across 16 target variables, we show that company growth can be decomposed into trend-driven and fluctuation-driven predictability, whose relative importance varies strongly with company size, while the trend component remains robust across different levels of volatility. Interpretability analyses further show that STIML captures multivariate dependencies beyond autocorrelation, and that macroeconomic variables contribute significantly less to predictive performance on average. Moreover, we find the scaling-based growth model overlooks asymmetric deviations, which instead contain the structured and learnable signals, suggesting a path to refine mechanistic growth models.
The “Theory of Competing Networks”, which integrates network science and game theory, reveals that cooperation is more stable than hierarchy, offering insights into geopolitical restructuring and shifting global alliances. Understanding its implications can help nations navigate today’s rapidly evolving multipolar world.
The influence of network structure on collective problem-solving is a central focus in collective intelligence. However, the causal mechanisms linking structure to collective outcomes remain unexplored. To explore these, we utilize an agent-based model, the Potions Task, which operationalizes problem-solving as a combinatorial process suited to information-theoretic analysis. We examine information-based metrics at the level of agent pairs across networks, analyzing variation over time to determine how they predict problem-solving across network structures. We measure redundancy (or conversely, synergy) of solutions discovered by agents with respect to the network's global knowledge. Building on findings that small-world networks support efficient problem-solving, our results uncover the underlying mechanism by which high-performing networks achieve efficiency, namely by balancing local redundancy with long-range synergy. Furthermore, we find that synergy consistently predicts group performance, including in network structures typically considered inefficient at complex tasks. Synergy in information processing, measured at both local and global levels, therefore mediates the effects of structure and can override them entirely, suggesting that information-processing synergy, rather than network structure itself, is the primary driver of group performance. By formalizing a causal framework for information processing in collectives, our study provides a mechanistic account of collective problem-solving beyond network structure alone.
Abstract How to define ‘life’ is an unresolved question in the philosophy of biology, but has become more urgent as researchers around the world attempt to create synthetic cells in the laboratory, develop intelligent and autonomous robots, and search for signatures of life elsewhere in the galaxy. Here, we discuss the pros and cons of some of the current approaches to defining ‘life’, then propose an alternative approach based on family resemblance. Using a statistical modelling framework, we find that although living and non-living entities can be grouped according to overall similarity, it is difficult to find a single set of criteria that can both define known forms of life and be useful in identifying or characterizing novel forms of life. We hope that the family resemblance approach will prove to be a fruitful alternative to traditional approaches to defining ‘life’.
Heavy-tailed success is common in human competition, but it is unclear when it signals runaway dominance versus fair opportunity for skill to accumulate. We study outcome distributions in three arenas: World War II Luftwaffe fighter aces (victories), U.S. biology and computer science faculty competing for NIH/NSF awards, and U.S. Olympic swimmers and French Olympic fencers (medal totals). For each domain we analyze system-total distributions over the full record, then partition the same data into periods aligned with institutional eras to test whether tail shape is stable or shifts over time. Using a common tail-frontier scan, we fit three discrete upper-tail models—discrete lognormal (dLN), Zipf, and shifted geometric—over varying retained fractions. Where rules are stable and participants enjoy sustained chances to compete, upper tails consistently concentrate around a dLN regime: heavy but sub-power-law, consistent with repeated multiplicative gains under what we term Relative-Fairness, where skill has a fighting chance to accumulate. Time-partitioned analyses probe falsifiability: relaxing selectivity or temporarily doubling resources shifts tails toward a thinner, geometric-like regime, while episodic dominance yields localized Zipf episodes. Stress tests that vary roster size and competition tier under fixed rules show that tail shape distinguishes chance-dominated, relatively fair, and dominance-driven regimes.
As academic knowledge exchange increasingly shifts to online platforms, understanding the dynamics of knowledge-seeking behaviors is essential for fostering inclusive and effective digital communities. This study investigates knowledge-seeking behaviors on academic question-and-answer (Q&A) platforms, focusing on how regional and gender identities shape topic selection and engagement. Findings indicate that gender-based knowledge disparities exhibit both regional-specific characteristics and global patterns. Female researchers are more active than male counterparts in seeking knowledge online, with significant regional disparities in their areas of inquiry. Despite their participation in specific topics, specifically technical topics in Life Science & Biomedicine field, females report lower levels of expertise and receive less assistance in these topics. This study contributes a nuanced view of global and gender-based knowledge disparities, underscoring the need for targeted strategies to enhance support for underrepresented groups within academic knowledge exchange.
Large language models (LLMs) have been increasingly used to simulate human behaviour because of their ability to generate contextually coherent dialogues. Such abilities can enhance the realism of models. However, the pursuit of realism is not necessarily compatible with the epistemic foundation of modelling. We explore when LLM agents can be too ‘human’ to model, i.e., when they are too expressive, detailed and intractable to be consistent with the abstraction, simplification, and interpretability typically demanded by modelling. Through a model-building thought experiment, we uncover five core dilemmas: a temporal resolution mismatch between natural conversation and abstract time steps; the need for intervention in agent conversations without undermining spontaneous outputs; the temptation to introduce rules while maintaining conversational naturalness; the tension between role consistency and role evolution; and the challenge of understanding emergence. These dilemmas lead LLM agents to an “uncanny valley”: more realistic than rule-based agents but recognisably unhuman.
The detection of online influence operations—coordinated campaigns by malicious actors to spread narratives—has traditionally depended on content analysis or network features. These approaches are increasingly brittle as generative models produce convincing text, platforms restrict access to behavioral data, and actors migrate to less-regulated spaces. We introduce a platform-agnostic framework that identifies malicious actors from their behavioral policies by modeling user activity as sequential decision processes. We apply this approach to 12,064 Reddit users, including 99 accounts linked to the Russian Internet Research Agency in Reddit’s 2017 transparency report, analyzing over 38 million activity steps from 2015-2018. Activity-based representations, which model how users act rather than what they post, consistently outperform content models in detecting malicious accounts. When distinguishing trolls—users engaged in coordinated manipulation—from ordinary users, policy-based classifiers achieve a median macro-F1 of 94.9%, compared to 91.2% for text embeddings. Policy features also enable earlier detection from short traces and degrade more gracefully under evasion strategies or data corruption. These findings show that behavioral dynamics encode stable, discriminative signals of manipulation on Reddit’s IRA-linked campaign, and point to resilient detection strategies in the era of synthetic content and limited data access.
The emergence of phototrophy is one of the most significant innovations in the history of life, vastly increasing available metabolic energy. Phototrophy is, however, known to have arisen only twice. This raises a curious question: if phototrophy was accessible enough to evolve twice, why has it never arisen again despite billions of years of subsequent evolution? Through physiological modeling, we demonstrate that chlorophototrophy and retinalophototrophy together saturate the bioenergetic landscape available to light-harvesting systems. They represent opposite solutions to key biophysical trade-offs: maximizing efficiency per photon versus maximizing metabolic flux, specialization versus versatility, and sophistication versus simplicity. Together they create an evolutionary priority effect, blocking any newly-arising phototrophic system from succeeding. By revealing the basis of this competitive exclusion, our work sheds light on a general principle - that early innovations can saturate ecological space such that they constrain future evolutionary possibilities, making apparently ‘easy’ innovations appear as rare events.
Understanding urban mobility requires models that capture how people interact with and navigate the built environment. We present a scalable, generalizable agent-based framework in which daily schedules emerge from the interplay between mandatory (e.g., work, school) and flexible (e.g., errands, food, leisure) activities, driven by evolving individual needs. The results of our model are validated against empirical patterns from the 2017 U.S. National Household Travel Survey, including activity distributions, origin-destination flows, and trip-chain length distributions. We introduce a normalized similarity metric to quantify agreement between simulated and empirical patterns. Most cities achieve scores above 0.80, demonstrating strong alignment without the need for city-specific calibration. The model scales efficiently to over 20 million agents, enabling full-population simulations of large metropolitan areas. This combination of universality and scalability enables scenario analysis for infrastructure stress testing, disaster recovery, innovation diffusion, and disease spread in urban systems.
Understanding how microorganisms adapt to novel physical and chemical environments requires integrating evolutionary, regulatory, and phenotypic perspectives. Here, we examined Streptococcus mutans populations previously evolved for 100 days under simulated microgravity (sMG) or combined microgravity and silver nitrate (sMGAg), generating new transcriptomic and phenotypic datasets and integrating them with prior whole-genome sequencing. These environments model key pressures encountered in enclosed spaceflight habitats, including altered fluid shear, oxidative challenges, and exposure to disinfectants. Populations maintained under normal gravity (NG) largely preserved ancestral metabolic and redox characteristics. In contrast, sMG populations exhibited divergent physiological and transcriptional outcomes that were not predictable from genomic variants alone, including multiple ROS response patterns, broad reductions in carbohydrate metabolism, and consistent retention of trehalose utilization. Populations evolved under sMGAg showed more convergent patterns, characterized by broad activation of oxidoreductase and metal-handling pathways, elevated basal ROS relative to the ancestral strain with reduced inducibility, and a consistent gain in nitrate-reduction capability. These outcomes reflect condition-associated physiological states resolved only through combined genomic, transcriptomic, and phenotype-level data, as no single data type was sufficient to capture the full structure of adaptive responses. Together, these findings illustrate how distinct physical and chemical stress regimes reshape the landscape of accessible evolutionary responses, with microgravity alone permitting a wider range of adaptive trajectories and microgravity combined with silver favoring more uniform physiological states. More broadly, this work demonstrates that integrated multi-level datasets are essential for accurately characterizing adaptive outcomes in extreme or non-terrestrial environments.
Social media platforms frequently prioritize efficiency to maximize ad revenue and user engagement, often sacrificing deliberation, trust, and reflective, purposeful cognitive engagement in the process. This manuscript examines the potential of friction—design choices that intentionally slow user interactions—as an alternate approach. We present a case against efficiency as the dominant paradigm on social media and advocate for a complex systems approach to understanding and analyzing friction. Drawing from interdisciplinary literature, real-world examples, and industry experiments, we highlight the potential for friction to mitigate issues like polarization, disinformation, and toxic content without resorting to censorship. We propose a state space representation of friction to establish a multidimensional framework and language for analyzing the diverse forms and functions through which friction can be implemented. Additionally, we propose several experimental designs to examine the impact of friction on system dynamics, user behavior, and information ecosystems, each designed with complex systems solutions and perspectives in mind. Our case against efficiency underscores the critical role of friction in shaping digital spaces, challenging the relentless pursuit of efficiency and exploring the potential of thoughtful slowing.