Music comprises two core structural components, melody and rhythm, that vary widely across cultures. Whether these components coevolve in a coupled way or follow independent trajectories remains unclear. We introduce a novel computational pipeline to extract vocal melodic pitch-interval and percussive inter-onset timing distributions from 27,628 popular songs across 59 countries, enabling large-scale cross-cultural comparison that bypasses traditional music annotations. Musical similarities between countries aligned with geographic and linguistic relationships, validating our approach. Substantial variation emerged in both melodic and rhythmic structures across countries, yet the diversity of the two components was not significantly correlated, challenging assumptions of coupled evolution. Only rhythmic diversity was significantly associated with ethnic and linguistic heterogeneity, while melodic diversity showed no such association. These findings suggest that melody and rhythm constitute partially independent systems shaped by distinct cultural and evolutionary pressures, rather than components of a single monolithic musical style.
How do people know when it is permissible to break a rule? Sometimes people _universalize_, asking "What if everyone felt at liberty to violate the rule?" While there is mounting evidence that universalization guides rule-breaking judgments, this evidence is limited to participants who are English-speaking, United States residents, leaving open the question of how universal universalization actually is. Moreover, geography and identity have important influences on morality and cultures differ widely in the stringency with which they adhere to rules. In this paper we use a language-agnostic, video-game paradigm to investigate whether universalization guides rule-breaking judgments in 20 countries across the globe (n=2,652 participants) and find that universalization plays an important role in moral judgment in every one. However, the strength of universalization varies. While cultures may vary dramatically in how strongly they adhere to norms, the underlying _logic_ of norm breaking appears remarkably consistent.
As artificial intelligence increasingly mediates public discourse, it becomes important to understand how human-AI collectives shape opinion formation, deliberation, and democratic outcomes. We present a novel experimental method for studying opinion dynamics in hybrid human-AI social networks. Participants, human or AI, were embedded in 5×5 grid lattice networks and iteratively asked to select and revise statements on a given polarizing topic over eight rounds. We compared three conditions: human-only, AI-only, and hybrid networks with equal proportions of human and AI participants. Hybrid human-AI networks achieved the lowest final polarization while, in contrast, human-only networks exhibited higher polarization with lower neighbor agreement. We also ran additional experiments varying Large Language Model (LLM) prompt framing to explore whether instruction design might influence convergence patterns. Although these early findings are preliminary and cannot yet support broad generalizations, they highlight the potential value of experimental social networks for understanding opinion dynamics in human-AI hybrid societies.
Adaptive experiments automatically optimize their design throughout the data collection process, which can bring substantial benefits compared to conventional experimental settings. Potential applications include, among others: computerized adaptive testing (for selecting informative tasks in ability measurements), adaptive treatment assignment (when searching experimental conditions maximizing certain outcomes), and active learning (for choosing optimal training data for machine learning algorithms). However, implementing these techniques in real time poses substantial computational and technical challenges. Additionally, despite their conceptual similarity, the above scenarios are often treated as separate problems with distinct solutions. In this paper, we introduce a practical and unified approach to real-time adaptive experiments that can encompass all of the above scenarios, regardless of the modality of the task (including textual, visual, and audio inputs). Our strategy combines active inference, a Bayesian framework inspired by cognitive neuroscience, with PsyNet, a platform for large-scale online behavioral experiments. While active inference provides a compact, flexible, and principled mathematical framework for adaptive experiments generally, PsyNet is a highly modular Python package that supports social and behavioral experiments with stimuli and responses in arbitrary domains. We illustrate this approach through two concrete examples: (1) an adaptive testing experiment estimating participants' ability by selecting optimal challenges, effectively reducing the amount of trials required by 30–40%; and (2) an adaptive treatment assignment strategy that identifies the optimal treatment up to three times as accurately as a fixed design in our example. We provide detailed instructions to facilitate the adoption of these techniques.
Perceptual systems adapt through individual experience across the lifespan, an ability referred to as plasticity. To understand perceptual plasticity, a promising avenue is to investigate how perception is shaped by cultural experience, as a process deeply embedded within collective practices of cultural production and social learning. The current review synthesizes findings from recent behavioral experiments investigating cross-cultural variation in rhythm perception. Specifically, these studies show that fundamental perceptual processes, such as event timing and rhythm categorization, display shared features but also systematic differences across cultural groups. Critically, these differences correlate with statistically prominent and socially relevant features of cultural production, revealing how perceptual systems are tuned to their music-cultural environments. Yet, how can cross-cultural differences in perception be related back to the collective practices that produce the diversity of cultural environments in the first place? To bridge this gap, we propose perceptual niche construction as an evolutionary approach that positions culture as both a source and a product of perceptual plasticity. That is, cultural experience tunes individual perception, yielding culturally diverse perceptual processes. These processes, in turn, create selection pressures shaping cultural production across nested timescales, resulting in diverse cultural environments. This approach presents implications for research in psychology and neuroscience, notably in proposing to operationalize culture as communities of learning and practice. Moreover, it highlights the relevance of contextually situated research, in view of accounting for the dynamic nature of culture-driven perceptual plasticity.
Real-world creative processes ranging from art to science rely on social feedback-loops between selection and creation. Yet, the effects of popularity feedback on collective creativity remain poorly understood. We investigate how popularity ratings influence cultural dynamics in a large-scale online experiment where participants ($N = 1\,008$) iteratively \textit{select} images from evolving markets and \textit{produce} their own modifications. Results show that exposing the popularity of images reduces cultural diversity and slows innovation, delaying aesthetic improvements. These findings are mediated by alterations of both selection and creation. During selection, popularity information triggers cumulative advantage, with participants preferentially building upon popular images, reducing diversity. During creation, participants make less disruptive changes, and are more likely to expand existing visual patterns. Feedback loops in cultural markets thus not only shape selection, but also, directly or indirectly, the form and direction of cultural innovation.
Across the sciences, autonomous systems are increasingly being used in closed-loop discovery, proposing new theories and designing and running experiments to test them. This approach is yet to be applied in the field of cognitive science, where the central bottleneck is theory-building: the creative step of turning the accumulated failures of existing models into better ones. Theory generation has remained manual even as data collection, modeling, and experiment design have been automated. We present the Automated Cognitive Scientist (AutoCog), a fully autonomous agentic-AI system that closes this loop. Large-language-model agents advocate competing theories, each expressed as an executable cognitive model, design experiments that best discriminate them, collect behavioral data from participants recruited online, score theories against collected data based on their generative performance, diagnose why they fail, and synthesize a better successor. Repeating this cycle allows them to search the space of theories, models, and experiments. In the domain of decision-making, AutoCog recovered known decision-making strategies from simulated behavior, including unconventional ones, showing that its discoveries are ultimately driven by the data rather than strictly bound by the priors of the underlying language models. When run with human participants, it produced theories that outperformed the established theories it was seeded with and generalized to held-out studies in two different experimental settings. It also surfaced a novel theory of multi-cue decision-making in which choices show diminishing sensitivity to feature values. The distinctive predictions of this theory were confirmed in a preregistered study with new participants. AutoCog demonstrates how an automated discovery system can be used to turn cognitive theory-building into an explicit, executable, and cumulative science.
Language influences our thinking and affects many aspects of cognition, from how we perceive the world to how we interact socially. Thus, objectively characterizing linguistic background is crucial for research in many areas, including second language acquisition, psycho-linguistics, and cognitive science. Traditional language proficiency tests, however, are manually composed by experts, limiting their scope for both lab and online settings. Here, we propose a pipeline that automatically derives a language proficiency test from a corpus of text and applies it to create new tests for 1,939 languages. Using this approach, we conducted a large-scale survey examining L1 and L2 proficiency across 34 countries, with participants tested on all 34 languages. Drawing from human ratings from 4,137 participants, our results validate that our test can effectively distinguish native speakers, second-language speakers, and nonspeakers within one minute, making it an effective tool for evaluating linguistic proficiency. We show that participants' linguistic and demographic backgrounds systematically influence both their language proficiency and their self-reported skills, and we map the prevalence of global languages, such as English and Spanish, among online participants. Moreover, we show that our vocabulary tests are strongly correlated with other linguistic competences-such as listening and writing-in a set of typologically varied languages, demonstrating our test is an efficient instrument to assess language proficiency. More broadly, our work offers a significant resource for investigating global variation in language skills and contributes to reducing the overreliance on the English language in the cognitive and social sciences.
Generative AI is increasingly transforming creativity into a hybrid human-artificial process, but its impact on the quality and diversity of creative output remains unclear. We study collective creativity using a controlled word-guessing task that balances open-endedness with an objective measure of task performance. Participants attempt to infer a hidden target word, scored based on the semantic similarity of their guesses to the target, while also observing the best guess from previous players. We compare performance and outcome diversity across human-only, AI-only, and hybrid human-AI groups. Hybrid groups achieve the highest performance while preserving high diversity of guesses. Within hybrid groups, both humans and AI agents systematically adjust their strategies relative to single-agent conditions, suggesting higher-order interaction effects, whereby agents adapt to each other's presence. Although some performance benefits can be reproduced through collaboration between heterogeneous AI systems, human-AI collaboration remains superior, underscoring complementary roles in collective creativity.
Writing code has been one of the most transformative ways for human societies to translate abstract ideas into tangible technologies. Modern AI is transforming this process by enabling experts and non-experts alike to generate code without actually writing code, but instead, through natural language instructions, or "vibe coding". While increasingly popular, the cumulative impact of vibe coding on productivity and collaboration, as well as the role of humans in this process, remains unclear. Here, we introduce a controlled experimental framework for studying collaborative vibe coding and use it to compare human-led, AI-led, and hybrid groups. Across 16 experiments involving 604 human participants, we show that people provide uniquely effective high-level instructions for vibe coding across iterations, whereas AI-provided instructions often result in performance collapse. We further demonstrate that hybrid systems perform best when humans retain directional control (providing the instructions), while evaluation is delegated to AI.
Rapid advances in generative AI are giving rise to hybrid societies in which humans and machines are deeply interconnected. However, understanding how human-AI interactions shape collective outcomes such as intelligence and creativity remains a challenge. We conducted large-scale social network experiments with 1,363 human participants and 7,865 AI bots, where agents iteratively selected, modified, and transmitted stories from their neighbours in the network. We compared networks composed exclusively of humans, AI agents, and hybrid human–AI populations. Results show that AI-only networks initially achieved high creativity and collective diversity but exhibited a pronounced decline in diversity over time. In contrast, hybrid human–AI networks consistently achieved a balance, combining high creativity with the highest levels of collective diversity across iterations. Importantly, humans embedded in hybrid networks became substantially more diverse than when interacting exclusively with other humans, providing empirical evidence that AI influences creativity not only through direct dyadic exchanges but also via emergent network-level processes. Semantic analyses indicate that hybrid networks succeed by balancing human tendency to preserve semantics with the introduction of novel ideas by AI agents. Moreover, partial benefits emerged when different AI models (e.g., GPT and Claude) were combined within the same network, even though each model was limited on its own. Overall, our results show that human–AI misalignment—often viewed as a limitation—can instead be a productive feature of hybrid societies, enhancing individual creativity while preserving collective diversity in hybrid systems.
Understanding cooperation in social systems is challenging because the ever-changing rules that govern societies interact with individual actions, resulting in intricate collective outcomes. In virtual-world experiments, we allowed people to make changes in the systems that they are making decisions within and investigated how they weigh the influence of different rules in decision-making. When choosing between worlds differing in more than one rule, a naive heuristics model predicted participants decisions as well, and in some cases better, than game earnings (utility) or by the subjective quality of single rules. In contrast, when a subset of engaged participants made instantaneous (within-world) decisions, their behavior aligned very closely with objective utility and not with the heuristics model. Findings suggest that, whereas choices between rules may deviate from rational benchmarks, the frequency of real time cooperation decisions to provide feedback can be a reliable indicator of the objective utility of these rules.
Rhythmic ability is a universal aspect of human cultures and sets the basis for musical rhythm perception and synchronization. While humans can synchronize movements to complex rhythms, it is unclear whether this capability extends to our primate ancestors. In this study, we explore whether primates can synchronize to complex rhythms and acquire rhythmic representations with generalizability akin to humans. Using controlled behavioral experiments, we provide evidence that monkeys not only can synchronize to short-long or long-short rhythms but also learn representations that generalize across a wide range of rhythm ratios and total durations. These results indicate ability to flexibly represent ratios within a relative timing framework is not exclusive to humans, but it is also present in monkeys. In addition, the produced intervals show a bias towards rhythmic categories. Notably, in an iterative tapping task, macaques and humans showed large priors for isochrony and integer ratios (2:1, 3:1). These results demonstrate a common biological foundation for rhythm synchronization in primates, extending our understanding of the shared cognitive mechanisms between primates and humans, and highlight the enormous potential of using monkeys to study the neurophysiological basis of complex rhythm perception. ### Competing Interest Statement The authors have declared no competing interest.
The evolution of music, speech, and sociality have been debated since before Darwin. The social bonding hypothesis proposes that these phenomena may be interlinked: musicality may have facilitated the evolution of social bonding beyond the possibilities of spoken language. Although dozens of experimental studies have argued that synchronised rhythms can promote bonding, methodological issues including publication bias, sample bias, experimenter effects, and appropriateness of experimental controls make it unclear whether synchronous singing reliably and generally enhances bonding relative to speaking. Here, we propose a Registered Report to overcome these issues through a global experiment in diverse languages aiming to collect data from 1800 participants across 60 sites. The social bonding hypothesis predicts that bonding will increase more after synchronous singing than after spoken (sequential) conversation or (simultaneous) recitation, while alternative hypotheses predict that song will not increase bonding relative to speech. Regardless of outcome, these results will provide an unprecedented understanding of cross-cultural relationships between music, speech, and sociality.
Understanding how cognitive and social mechanisms shape the evolution of complex artifacts such as songs is central to cultural evolution research. Social network topology (what artifacts are available?), selection (which are chosen?), and reproduction (how are they copied?) have all been proposed as key influencing factors. However, prior research has rarely studied them together due to methodological challenges. We address this gap through a controlled naturalistic paradigm whereby participants (N=2,404) are placed in networks and are asked to iteratively choose and sing back melodies from their neighbors. We show that this setting yields melodies that are more complex and more pleasant than those found in the more-studied linear transmission setting, and exhibits robust differences across topologies. Crucially, these differences are diminished when selection or reproduction bias are eliminated, suggesting an interaction between mechanisms. These findings shed light on the interplay of mechanisms underlying the evolution of cultural artifacts.
Getting a group to adopt cooperative norms is an enduring challenge. But in real-world settings, individuals don't just passively accept static environments, they act both within and upon the social systems that structure their interactions. Should we expect the dynamism of player-driven changes to the "rules of the game" to hinder cooperation – because of the substantial added complexity – or help it, as prosocial agents tweak their environment toward non-zero-sum games? We introduce a laboratory setting to test whether groups can guide themselves to cooperative outcomes by manipulating the environmental parameters that shape their emergent cooperation process. We test for cooperation in a set of economic games that impose different social dilemmas. These games vary independently in the institutional features of stability, efficiency, and fairness. By offering agency over behavior along with second-order agency over the rules of the game, we understand emergent cooperation in naturalistic settings in which the rules of the game are themselves dynamic and subject to choice. The literature on transfer learning in games suggests that interactions between features are important and might aid or hinder the transfer of cooperative learning to new settings.
Humans organize semantic knowledge into complex networks that encode relations between concepts. The structure of those networks has broad implications for human cognitive processes, and for theories of semantic development. Evidence from large lexical networks such as those derived from word associations suggest that semantic networks are characterized by high sparsity and clustering while maintaining short average paths between concepts, a phenomenon known as a 'small-world' network. It has also been argued that those networks are 'scale-free', meaning that the number of connections (or degree) between concepts follows a power-law distribution whereby most concepts have few connections while a few have many. However, the scale-free property is still debated, and the extent to which the lexical evidence reflects the naturally occurring semantic regularities of the environment has not been investigated systematically. To address this, we collected and analyzed semantic descriptors, human evaluations, and similarity judgments from four large datasets of naturalistic stimuli across three modalities (visual, auditory, and audio-visual) comprising 7,916 stimuli and 610,841 human responses. By connecting concepts that co-occur as descriptors of the same stimuli, we construct 'grounded' semantic networks. We show that these networks exhibit a clear small-world structure with a degree distribution that is best captured by a truncated power law (i.e., the most-connected concepts are less common than predicted by a perfect power law). We further show that these networks are predictive of human sensory judgments on these domains, as well as reaction times in an independent lexical decision task. Finally, we show that grounded networks also share overlapping themes with previously analyzed lexical networks, which upon a more rigorous re-analysis are revealed to be truncated too. Our findings shed new light on the origins of the structure of semantic networks by grounding it in the semantic regularities of the environment.
Humans across cultures show an outstanding capacity to perceive, learn, and produce musical rhythms. These skills rely on mapping the infinite space of possible rhythmic sensory inputs onto a finite set of internal rhythm categories. What is the nature of the brain processes underlying rhythm categorization? We used electroencephalography to measure brain activity as human participants listened to a continuum of rhythmic sequences characterized by repeating patterns of two interonset intervals. Using frequency and representational similarity analyses, we show that brain activity does not merely track the temporal structure of rhythmic inputs but, instead, produces categorical representation of rhythms. These neural rhythm categories arise automatically, independent of any motor- or timing-related tasks, yet exhibit strong similarity with categorization observed in overt behavior. Together, these results and methodological advances constitute a critical step toward understanding the biological roots and diversity of musical behaviors across cultures.