Large language models (LLMs) are growing increasingly capable, prompting recent interest in LLM teams. Yet, despite increased deployment of LLM teams at scale, we lack a principled framework for addressing key questions such as when a team is helpful, how many agents to use, how structure impacts performance – and whether a team is better than a single agent. Rather than designing and testing these possibilities through trial-and-error, we propose using distributed systems as a principled foundation for creating and evaluating LLM teams. We find that many of the fundamental advantages and challenges studied in distributed computing also arise in LLM teams, highlighting the rich practical insights that can come from the cross-talk of these two fields of study.
Large Language Models (LLMs) have emerged as integral tools for reasoning, planning, and decision-making, drawing upon their extensive world knowledge and proficiency in language-related tasks. LLMs thus hold tremendous potential for natural language interaction within multi-agent systems to foster cooperation. However, LLM agents tend to over-report and comply with any instruction, which may result in information redundancy and confusion in multi-agent cooperation. Inspired by human organizations, this paper introduces a framework that imposes prompt-based organization structures on LLM agents to mitigate these problems. Through a series of experiments with embodied LLM agents and human-agent collaboration, our results highlight the impact of designated leadership on team efficiency, shedding light on the leadership qualities displayed by LLM agents and their spontaneous cooperative behaviors. Further, we harness the potential of LLMs to propose enhanced organizational prompts, via a Criticize-Reflect process, resulting in novel organization structures that reduce communication costs and enhance team efficiency.
A hallmark of effective teaching is that it grants learners not just a collection of facts about the world, but also a toolkit of abstractions that can be applied to solve new problems. How do humans teach abstractions from examples? Here, we applied Bayesian models of pedagogy to a necklace-building task where teachers create necklaces to teach a learner "motifs" that can be flexibly recombined to create new necklaces. In Experiment 1 (N = 151), we find that human teachers produce necklaces that are simpler (i.e., have lower algorithmic complexity) than would be expected by chance, as indexed by a model that samples uniformly from all necklaces that contain the target motifs. This tendency to select simpler examples is captured by a pedagogical sampling model that tries to maximize the learner's belief in the true motifs by prioritizing examples that have fewer alternative interpretations. In Experiment 2 (N = 295), we find that simplicity is beneficial. Human learners recover the underlying motifs better when teachers produce simpler sequences, as predicted by the pedagogical sampling model. However, humans learn best from human teachers rather than from model-generated examples, which suggests that human teachers have additional expectations about how learners will interpret examples that are not captured by standard models of teaching. Our work provides a principled framework to understand when and why teachers use simple examples to convey abstract knowledge.
Individual contributors to a collaborative task are often rewarded for going above and beyond-salespeople earn commissions, athletes earn performance bonuses, and companies award special parking spots to their employee of the month. How do we decide when to reward collaborators, and are these decisions closely aligned with how responsible they were for the outcome of a collaboration? In Experiments 1a and 1b (N = 360), we tested how participants give bonuses, using stimuli and an experiment design that has previously been used to elicit responsibility judgments (Xiang et al., 2023a). Past work has found that responsibility judgments are driven both by how much effort people actually contributed and how much they could have contributed (Xiang et al., 2023a). In contrast, here we found that participants allocated bonuses based only on how much effort agents actually contributed. In Experiments 2a and 2b (N = 358), we introduced agents who were instructed to exert a particular level of effort; participants still rewarded effort, but their rewards were more sensitive to the precise level of effort exerted when the agents decided how much effort to exert. Together, these findings suggest that people reward collaborators based on their willingness to exert effort, and point to a difference between decisions about how to assign responsibility to collaborators and how to incentivize them. One possible explanation for this difference is that responsibility judgments may reflect causal inference about past collaborations, whereas providing incentives may motivate collaborators to keep exerting effort in the future. Our work sheds light on the cognitive capacities that underlie collaboration.
Humans are remarkably adept at collaboration, able to infer the strengths and weaknesses of new partners in order to work successfully towards shared goals. To build AI systems with this capability, we must first understand its building blocks: does such flexibility require explicit, dedicated mechanisms for modelling others—or can it emerge spontaneously from the pressures of open-ended cooperative interaction? To investigate this question, we train simple model-free RNN agents to collaborate with a population of diverse partners. Using the 'Overcooked-AI' environment, we collect data from thousands of collaborative teams, and analyse agents' internal hidden states. Despite a lack of additional architectural features, inductive biases, or auxiliary objectives, the agents nevertheless develop structured internal representations of their partners' task abilities, enabling rapid adaptation and generalisation to novel collaborators. We investigated these internal models through probing techniques, and large-scale behavioural analysis. Notably, we find that structured partner modelling emerges when agents can influence partner behaviour by controlling task allocation. Our results show that partner modelling can arise spontaneously in model-free agents—but only under environmental conditions that impose the right kind of social pressure.
When should we encourage specialization in multi-agent systems versus train generalists that perform the entire task independently? We propose that specialization largely depends on task parallelizability: the potential for multiple agents to execute task components concurrently. Drawing inspiration from Amdahl's Law in distributed systems, we present a closed-form bound that predicts when specialization improves performance, depending only on task concurrency and team size. We validate our model on two standard MARL benchmarks that represent opposite regimes – StarCraft Multi-Agent Challenge (SMAC, unlimited concurrency) and Multi-Particle Environment (MPE, unit-capacity bottlenecks) – and observe close alignment between the bound at each extreme and an empirical measure of specialization. Three follow-up experiments in Overcooked-AI demonstrate that the model works in environments with more complex spatial and resource bottlenecks that allow for a range of strategies. Beyond prediction, the bound also serves as a diagnostic tool, highlighting biases in MARL training algorithms that cause sub-optimal convergence to specialist strategies with larger state spaces.
Humans collaborate to improve productivity, but when is it acceptable for a collaborator to remain idle? Theories from distributed computer systems suggest that, depending on the task structure, division of labor leads to diminishing returns in efficiency as group size increases. We examine whether people are aware of these limitations to collaboration, and how considerations of task efficiency may affect the perceived acceptability of idleness, the withholding of effort during collaborative tasks. Across four experiments (N=1,124), participants saw scenarios where a single collaborator remained idle while other group members washed dishes, prepared a salad, or created flashcards. We manipulated task structure by varying the number of guests (group size), the amount of work to be done (workload), and the number of tools available to do it (environmental bottlenecks), which each constrain how much faster the group could have finished the task if the idle agent had contributed. Participants judged idleness as more acceptable when the idle agent's contributions would have a smaller effect on task efficiency. These judgments were best captured by a variant of Amdahl's Law, a theory from distributed systems that predicts the idle agent's potential impact by integrating group size, workload, and bottlenecks, compared to simpler heuristic models that consider a subset of these factors. Together, our findings lay the groundwork to study human collaborations as natural distributed systems.
Humans have developed technologies to adapt to virtually every habitat on Earth. But why do some communities develop thriving technological repertoires while others stagnate? We address this question by analyzing player behavior in One Hour One Life (OHOL), a multiplayer online game where players can build technologically advanced communities from scratch (N = 22,011 players, 2,700 communities, 428,255 playthroughs). Players are randomly assigned to a community in each playthrough and can contribute to it for up to one hour. Over time, through many players' contributions, communities can survive for weeks and amass rich technological repertoires. Thus, this dataset provides a unique quasi-experiment into how the composition of communities affects their growth and decline. Using this approach, we find that technological developments are the product of interactions between individuals and the communities where they are placed. Individuals take on jobs that align with those of their closest peers, and they selectively contribute new technologies in areas of their expertise that diverge from the rest of the community. In the aggregate, these two processes—aligning to form specialized communities, or diverging to form diverse ones—have opposing effects on the size and stability of a community's technological repertoire, and imbalances in specialization serve as an early indicator of population collapse. Our results suggest that, to survive, communities must balance between diversifying to develop new technologies, while specializing to maintain the ones they already have. Our approach provides a testbed for theories of large-scale social phenomena that would otherwise be difficult to test against real-world data or traditional laboratory experiments.
Human learning does not stop at solving a single problem. Instead, we seek new challenges, define new goals, and come up with new ideas. Unlike the classic explore-exploit trade-off between known and unknown options, making new tools or generating new ideas is not about collecting data from existing unknown options, but rather about create new options out of what is currently available. We introduce a discovery game designed to study how rational agents make decisions about pursuing innovations, where discovering new ideas is a process of combining existing ideas in an open-ended compositional space. We derive optimal policies of this decision problem formalized as a Markov decision process, and compare people's behaviors to the model predictions in an online behavioral experiment. We found evidence that people both innovate rationally, guided by potential returns in this discovery game, and under- and over-explore systematically in different settings.
A hallmark of effective teaching is that it grants learners not just a collection of facts about the world, but also a toolkit of abstractions that can be applied to solve new problems. How do humans transmit and acquire generalizable abstractions from examples? Here, we applied Bayesian models of pedagogy to a necklace-building task where teachers create necklaces to teach a learner ``motifs'' that can be flexibly recombined to create new necklaces. In Experiment 1 (N = 151), we find that human teachers produce necklaces that are simpler (i.e., have lower algorithmic complexity) than would be expected by chance, as indexed by a model that samples uniformly from all necklaces that contain the target motifs. This tendency to select simpler examples is partially captured by a pedagogical sampling model that tries to maximize the learner's belief in the underlying motifs. In Experiment 2 (N = 295), we find that simplicity is beneficial. Human learners recover the underlying motifs better when teachers produce simpler sequences, and they learn best from human teachers rather than from model-generated examples. Our results suggest that the computational principles that underlie effective communication and teaching may also provide a first step towards understanding the transmission of culturally-specific abstractions.
How do teachers learn about what learners already know? How do learners aid teachers by providing them with information about their background knowledge and what they find confusing? We formalize this collaborative reasoning process using a hierarchical Bayesian model of pedagogy. We then evaluate this model in two online behavioral experiments (N = 312 adults). In Experiment 1, we show that teachers select examples that account for learners' background knowledge, and adjust their examples based on learners' feedback. In Experiment 2, we show that learners strategically provide more feedback when teachers' examples deviate from their background knowledge. These findings provide a foundation for extending computational accounts of pedagogy to richer interactive settings.
Board, card, or video games have been played by virtually every individual in the world population, with both children and adults participating. Games are popular because they are intuitive and fun. These distinctive qualities of games also make them ideal as a platform for studying the mind. By being intuitive, games provide a unique vantage point for understanding the inductive biases that support behavior in more complex, ecological settings than traditional lab experiments. By being fun, games allow researchers to study new questions in cognition such as the meaning of "play'' and intrinsic motivation, while also supporting more extensive and diverse data collection by attracting many more participants. We describe both the advantages and drawbacks of using games relative to standard lab-based experiments and lay out a set of recommendations on how to gain the most from using games to study cognition. We hope this article will lead to a wider use of games as experimental paradigms, elevating the ecological validity, scale, and robustness of research on the mind.
In order to efficiently divide labor with others, it is important to understand what our collaborators can do (i.e., their competence ). However, competence is not static-people get better at particular jobs the more often they perform them. This plasticity of competence creates a challenge for collaboration: For example, is it better to assign tasks to whoever is most competent now, or to the person who can be trained most efficiently "on -the -job"? We conducted four experiments ( N = 396 ) that examine how people make decisions about whom to train (Experiments 1 and 3) and whom to recruit (Experiments 2 and 4) to a collaborative task, based on the simulated collaborators' starting expertise, the training opportunities available, and the goal of the task. We found that participants' decisions were best captured by a planning model that attempts to maximize the returns from collaboration while minimizing the costs of hiring and training individual collaborators. This planning model outperformed alternative models that based these decisions on the agents' current competence, or on how much agents stood to improve in a single training step, without considering whether this training would enable agents to succeed at the task in the long run. Our findings suggest that people do not recruit and train collaborators based solely on their current competence, nor solely on the opportunities for their collaborators to improve. Instead, people use an intuitive theory of competence to balance the costs of hiring and training others against the benefits to the collaboration.
Human learning does not stop at solving a single problem. Instead, we seek new challenges, define new goals, and come up with new ideas. What drives people to disrupt the existing conceptual landscape and create new things? Here, we examine the decision to create new things under different levels of potential returns. We formalize innovation as stochastically recombining existing ideas, where successful and more complex combinations generate higher returns. This formalization allows us to cast innovation-seeking as a Markov decision process, and derive optimal policies under different settings. Data collected through an online behavioral experiment confirm our prediction that people should invest more time and effort in seeking innovations when they know the chances of success are high and the potential new ideas would be rewarding. However, people also deviate from being optimal, both innovating more and less than they should in different settings.
With the advent of multivariate pattern analysis (MVPA) as an important analytic approach to fMRI, new insights into the functional organization of the brain have emerged. Several software packages have been developed to perform MVPA analysis, but deploying them comes with the cost of adjusting data to individual idiosyncrasies associated with each package. Here we describe PyMVPA BIDS-App, a fast and robust pipeline based on the data organization of the BIDS standard that performs multivariate analyses using powerful functionality of PyMVPA. The app runs flexibly with blocked and event-related fMRI experimental designs, is capable of performing classification as well as representational similarity analysis, and works both within regions of interest or on the whole brain through searchlights. In addition, the app accepts as input both volumetric and surface-based data. Inspections into the intermediate stages of the analyses are available and the readability of final results are facilitated through visualizations. The PyMVPA BIDS-App is designed to be accessible to novice users, while also offering more control to experts through command-line arguments in a highly reproducible environment.
By collaborating with others, humans can pool their limited knowledge, skills, and resources to achieve goals that outstrip the abilities of any one person. What cognitive capacities make human collaboration possible? Here, we propose that collaboration is grounded in an intuitive understanding of how others think and of what they can do-in other words, of their mental states and competence. We present a belief-desire-competence framework that formalizes this proposal by extending existing models of commonsense psychological reasoning. Our framework predicts that agents recursively reason how much effort they and their partner will allocate to a task, based on the rewards at stake and on their own and their collaborator's competence. Across three experiments (N = 249), we show that the belief-desire-competence framework captures human judgments in a variety of contexts that are critical to collaboration, including predicting whether a joint activity will succeed (Experiment 1), selecting incentives for collaborators (Experiment 2), and choosing which individuals to recruit for a collaborative task (Experiment 3). Our work provides a theoretical framework for understanding how commonsense psychological reasoning contributes to collaborative achievements.