nbgrader is a flexible tool for creating and grading assignments in the Jupyter Notebook (Kluyver et al., 2016). nbgrader allows instructors to create a single, master copy of an assignment, including tests and canonical solutions. From the master copy, a student version is generated without the solutions, thus obviating the need to maintain two separate versions. nbgrader also automatically grades submitted assignments by executing the notebooks and storing the results of the tests in a database. After auto-grading, instructors can manually grade free responses and provide partial credit using the formgrader Jupyter Notebook extension. Finally, instructors can use nbgrader to leave personalized feedback for each student’s submission, including comments as well as detailed error information.
Mining scientific articles is hard when many of them are inaccessible behind paywalls.The Public Library of Science (PLOS) is a non-
Binder is an open source web service that lets users create sharable, interactive, reproducible environments in the cloud. It is powered by other core projects in the open source ecosystem, including JupyterHub and Kubernetes for managing cloud resources. Binder works with pre-existing workflows in the analytics community, aiming to create interactive versions of repositories that exist on sites like GitHub with minimal extra effort needed. This paper details several of the design decisions and goals that went into the development of the current generation of Binder.
When evaluating causal explanations, simpler explanations are widely regarded as better explanations. However, little is known about how people assess simplicity in causal explanations or what the consequences of such a preference are. We contrast 2 candidate metrics for simplicity in causal explanations: node simplicity (the number of causes invoked in an explanation) and root simplicity (the number of unexplained causes invoked in an explanation). Across 4 experiments, we find that explanatory preferences track root simplicity, not node simplicity; that a preference for root simplicity is tempered (but not eliminated) by probabilistic evidence favoring a more complex explanation; that committing to a less likely but simpler explanation distorts memory for past observations; and that a preference for root simplicity is greater when the root cause is strongly linked to its effects. We suggest that a preference for root-simpler explanations follows from the role of explanations in highlighting and efficiently representing and communicating information that supports future predictions and interventions.
The experimental study of cultural evolution, social learning, cooperation, and collective decision-making asks fundamental questions about our capacities to learn, decide, and communicate in a world that is shared with other people. Experimental studies of cultural evolution have revealed a wealth of findings, including how structured forms of communication emerge from individual learning and decisionmaking (Verhoef, Kirby, & Padden, 2011; Claidière, Smith, Kirby, & Fagot, 2014), the inductive biases underlying human decision making (Griffiths et al., 2008), how innovations accumulate in populations to produce technologies that go beyond what any one individual could create (Caldwell, & Millen, 2008; Derex & Boyd, 2015), and how the mode of communication affects transmission and acquisition of new skills (Morgan et al., 2015). However, in-laboratory experiments of this kind are resource intensive and logistically complex, requiring recruitment and coordination of participants to perform tasks sequentially and in concert, with enough space and time to isolate and control their interactions. These requirements drive experimental designs towards simple network structures (such as the transmission chain), small groups, and limited interaction between participants. Even where such experiments can be carried out using computers, the complexity of each experiment often means that existing software is unsuitable, leading each researcher to build bespoke software for their particular experiment. In addition to slowing the rate at which such experiments can be carried out, this also makes it hard to share code, to replicate other's experiments, and to build off the work of others. To address these issues, we created a software-based tool for orchestrating cultural transmission using online crowdsourcing. Our tool, named Wallace, builds on psiTurk (Gureckis et al., 2015) to provide efficient high-throughput automation for running behavioral experiments involving cultural transmission. Wallace recruits participants, obtains their informed consent, arranges them into a network, coordinates their communication, records the data they produce, pays them, and validates and manages the resulting data. Wallace runs on commodity hardware and cloud platforms, uses a custom API, and uses widely supported languages and markup languages such as Python, HTML5, JavaScript, and CSS. It is released as open-source software under the permissive MIT license. Wallace is modular and includes a library of components that can be used to quickly create new experiments. Prepackaged network structures include linear chains (e.g., Bartlett, 1932), scale-free networks (e.g., Bednarik et al., 2014), star and burst formations, micro-society (e.g. McElreath et al., 2005), and the discrete generational structure of the Wright–Fisher model from population genetics (Wright, 1931; Fisher, 1930), among others. Prepackaged behavioral tasks include story recall, category learning, function learning, magnitude estimation, a public goods game, stimulus–response mapping, and numerosity judgment. Nonetheless, experiments can also use custom network structures, processes, and tasks, which can be built by modifying the provided templates, allowing experimental designs of arbitrary complexity.
A successful design accounts for the structure of the problem it is aimed at solving. When it is a human-directed design, this includes the expectations of its users. How do we arrive at such a design? One approach starts from first principles (e.g., simplicity, unity, symmetry, balance) to evaluate the quality of proposed designs. Here, we introduce design from zeroth principles, a form of human-in-the-loop computation that synthesizes a design that conforms to its users’ expectations. The technique begins by constructing a transmission chain seeded with a random design. Each user in the chain is exposed to the design and then recreates it, passing along their recreation to the next user, who does the same. Through this iterative process, the users’ perceptual, inductive, and reconstructive biases directly transform the initial design into one that is better fit to human cognition. Such designs are easier to learn and harder to forget. We evaluated the approach in three domains — stimulus–response mappings, vanity phone numbers, and letter placement in typeset words — and show that it produces a good design in each.
The craft of writing is hard despite the abundance of thoughtful advice available in usage guides and other sources. This is partly a problem of medium: amassing advice is not enough to improve writing. Writing would thus benefit if our collective knowledge about best practices in writing were extracted and transformed into a medium that makes the knowledge more accessible to authors.
Past research has shown that people use temporal information to detect and discriminate between different causal relationships and that timing-based causal inferences are modulated by explicit information and domain-appropriate expectations. Many of these past results suggest that learners make inferences about hidden causes from timing information, but there have been no systematic studies of the ways in which subtle changes in temporal information can shape inferences about the presence and nature of hidden causes. We present new results showing that people make nuanced causal inferences when faced with streams of events, using temporal information to infer the presence of simple generative relationships, independent and common hidden causes, and causal cycles. Interpreted in a Bayesian framework, these results shed light on the cues and tacit temporal expectations that people use to make efficient use of temporal information.
Data continuously stream into our minds, guiding our learning and inference with no trial delimiters to parse our experience. These data can take on a variety of forms, but research on causal learning has emphasized discrete contingency data over continuous sequences of events. We present a formal framework for modeling causal inferences about sequences of point events, based on Bayesian inference over nonhomogeneous Poisson processes (NHPPs). We show how to apply this framework to successfully model data from an experiment by Lagnado and Speekenbrink (2010) which examined human learning from sequences of point events.
Cultural Evolution with Sparse Testimony: when does the Cultural Ratchet Slip? Andrew Whalen (aczw@st-andrews.ac.uk) Department of Biology, University of St Andrews, St Andrews, Fife KY16 9TH UK Luke Maurits (luke.maurits@berkeley.edu) Michael Pacer (mpacer@berkeley.edu) Thomas L. Griffiths (tom griffiths@berkeley.edu) Department of Psychology, University of California, Berkeley, Berkeley, CA 94720 USA Abstract Humans have accumulated a wealth of knowledge over the course of many generations, implementing a kind of “cultural ratchet”. Past work has used models and experiments in the iterated learning paradigm to understand how knowledge is ac- quired and changed over generations. However, this work has assumed that learners receive extremely rich testimony from their teacher: the teacher’s entire posterior distribution over possible states of the world. We relax this assumption and show that much sparser testimony may still be sufficient for learners to improve over time, although with limits on the con- cepts that can be learned. We experimentally demonstrate this result by running an iterated learning experiment based on a classic category learning task. Introduction The sciences are impressive; humanity can be proud. How- ever the work of science can hardly be conceived of, let alone realized, in a single generation. Data needs to accrue over time to shine light on different theories; with new in- formation, the landscape of theories changes, making some plausible while rendering others unthinkable. In addition to science, humanity has accumulated a vast body of practical knowledge and technology which has permitted our adapta- tion to nearly every environment on Earth (Boyd & Richer- son, 1988). Our species is distinguished by this accumulation of knowledge, but our understanding of the process underly- ing this “cultural ratchet” (Tomasello, 1994, 1999) is in its early stages. One aspect of the cultural ratchet’s operation that may have significant consequences for cultural evolution is the amount and kind of information that is passed between generations. Beppu and Griffiths (2009) investigated this aspect of the cul- tural ratchet in the context of an iterated function learning task. When participants in this task provided testimony to fu- ture participants which consisted only of demonstrated data (assumed to be undifferentiated from data observed “in the wild”), groups performed no better than single learners who received only data from the world — no ratcheting effect oc- curred. However, when participants provided their entire the- ory about how the world works — in Bayesian learning terms, their complete posterior beliefs — then the groups eventually learned the correct function. We know that when copious and rich information is passed from generation to generation, the cultural ratchet works flawlessly; given a steady stream of data from the world, sci- ence marches forward. We also know that under conditions of much poorer information passing, the cultural ratchet “slips”; a constant stream of data does not guarantee progress. Be- tween these two extremes are a range of possibilities, which may lie closer to actual human information passing than ei- ther extreme. What forms of testimony passing are needed for the cultural ratchet to catch more often than slip? We take a first step toward addressing this question in this paper. We construct a simple model of iterated learning with “sparse” testimony and consider three forms of evaluating social testimony. We apply this model to a category learn- ing task and find that limited social testimony may lead to iterative improvements across generations; however we also find that these improvements may not allow learners to find the correct hypothesis. We find that in hard category learning tasks, with limited personal data, learners may not perfectly learn the category, although they perform better than the ini- tial learners in the chain. These predictions are confirmed by an iterated learning experiment using a similar category learning task. In the experiment, we find that participant ac- curacy improved across generations, however in most of the conditions the amount of improvement is limited depending on the difficulty of the task and the amount of private data re- ceived. These results suggest that while passing limited testi- mony can still be sufficient to improve the accuracy of groups compared to receiving no testimony, it may not be enough to learn particularly challenging tasks with limited data. Iterated Learning and Cultural Evolution Iterated learning is a widely used computational and exper- imental paradigm for understanding how inductive biases might shape linguistic preferences over the course of multiple generations and influence how languages develop and change (Kirby, 2000, 2001; Perfors & Navarro, 2011). It has since been generalized beyond this setting. Griffiths and Kalish (2007) showed that if learners receive only social testimony then the long term distribution of beliefs of the population will be the same as the prior beliefs of each learner. Much of human learning does not take place purely on the basis of socially transmitted information. Human learn- ers also receive data directly from the world: be they scien- tists measuring the behavior of particles in a laboratory or hunter-gatherers testing new tools in an unfamiliar environ- ment. When learners receive outside data as well as testi- mony, the convergence to the prior shown by Griffiths and Kalish does not hold, and learners’ long-term behavior de-
Caching Algorithms and Rational Models of Memory Avi Press (avipress@berkeley.edu) Michael Pacer (mpacer@berkeley.edu) Thomas L. Griffiths (tom griffiths@berkeley.edu) Department of Psychology, University of California, Berkeley, Berkeley, CA 94720 USA Brian Christian (brian.christian@berkeley.edu) Institute of Cognitive and Brain Sciences, University of California, Berkeley, Berkeley, CA 94720 USA Abstract People face a problem similar to that faced by algorithms that manage the memory of computers: trying to organize informa- tion to maximize the chance it will be available when needed in the future. In computer science, this problem is known as “caching”. Inspired by this analogy, we compared the prop- erties of a model of human memory proposed by Anderson and Schooler (1991) and caching algorithms used in computer science. We tested each algorithm on a dataset relevant to hu- man cognition: headlines from the New York Times. In addi- tion to overall performance, we investigated whether the algo- rithms from computer science replicated the well-documented effects of recency, practice, and spacing on human memory. Anderson and Schooler’s model performed comparably to the worst caching algorithms, but was the only model that captured the spacing effects seen in human memory data. All models showed similar effects of recency and practice. Keywords: memory, caching algorithms, rational analysis Introduction Between our own experience forgetting things and the vol- umes of literature describing how fallible our memory is, it is easy to be critical of human memory. On the other hand, those with hyperthymestic syndrome (who can’t help but remember excessive details about every day of their lives) struggle to deal with their inability to forget useless information (Parker, Cahill, & McGaugh, 2006). How does our brain know what information should be kept and what should be forgotten? An analogue of the problem faced by human memory – as pointed out by Anderson and Milson (1989) – is a library try- ing to determine which books to keep in its collection. With finite shelf space, the library needs to decide which books are most likely to be needed in the future, relegating the others to long-term storage. Anderson and Milson used this obser- vation as inspiration for a rational model of human memory, which prioritizes stored information by how likely it is to be needed in the future. Anderson and Milson showed that this approach captured several phenomena of human memory. However, the rational model proposed by Anderson and Milson deviates from the library analogy in allowing in- finitely many items to be stored in memory, with retrieval failure being the result of the need probability of a target item being below a threshold determined by the cost of searching. But an alternative construal of the problem is much closer to the original analogy: What if human memory really did have finite capacity? How should we choose what to forget? The problem of choosing what to forget is an instance of what computer scientists term caching. A cache is a small yet fast block of memory (normally due to its hardware de- sign and proximity to the processor), where a limited set of items can be stored. Every time the computer needs data that is not in the cache it must fetch it from somewhere that will take far more time to access, such as a hard disk. Whenever an item is added to the cache, the computer must use an al- gorithm to decide which other item to evict. The computer is thus constantly deciding what to forget, managing its limited memory resources to maximize the probability of the cache containing the items most likely to be needed. In this paper, we explore the consequences of a rational analysis of human memory that assumes finite, rather than infinite, capacity. By looking at what happens when an envi- ronment is filtered through a finite cache, we can determine whether the statistical patterns corresponding to practice, re- cency, and spacing effects are relevant to successfully man- aging a finite memory. This is also potentially valuable for computer science, as we can see whether a caching scheme based on human memory improves on existing algorithms. The plan of the paper is as follows. First, we summarize the memory phenomena that have been used to evaluate ra- tional models, and describe the models themselves. Next, we define the problem of caching and introduce a set of caching algorithms. These algorithms are then evaluated in a set of four simulations. The first assesses overall performance. The others explore practice, recency, and spacing effects in turn. Human Memory Following Anderson and Schooler (1991), we will focus on three properties of human memory: practice, recency, and spacing. We review these phenomena, then turn to how they have been explained using rational models of memory. Behavioral Phenomena Practice The first two phenomena – practice and recency – are based on data collected by Ebbinghaus (1885/1913). Through rigorous self-experimentation, Ebbinghaus was able to discover some of the most basic aspects of human memory. The practice effect is simply that more times an item has been encountered, the more likely it can be recalled. Subsequent work has attempted to identify the form of the relationship between practice and retention, and found that a power-law best captures this relationship (Newell & Rosenbloom, 1981). Recency Ebbinghaus also noted that the more recently an item was encountered, the more likely it can be recalled.
We evaluate four computational models of explanation in Bayesian networks by comparing model predictions to human judgments. In two experiments, we present human participants with causal structures for which the models make divergent predictions and either solicit the best explanation for an observed event (Experiment 1) or have participants rate provided explanations for an observed event (Experiment 2). Across two versions of two causal structures and across both experiments, we find that the Causal Explanation Tree and Most Relevant Explanation models provide better fits to human data than either Most Probable Explanation or Explanation Tree models. We identify strengths and shortcomings of these models and what they can reveal about human explanation. We conclude by suggesting the value of pursuing computational and psychological investigations of explanation in parallel.
Elements of a rational framework for continuous-time causal induction Michael Pacer (mpacer@berkeley.edu) Thomas L. Griffiths (tom griffiths@berkeley.edu) Department of Psychology, University of California, Berkeley, Berkeley, CA 94720 Abstract Temporal information plays a major role in human causal in- ference. We present a rational framework for causal induc- tion from events that take place in continuous time. We de- fine a set of desiderata for such a framework and outline a strategy for satisfying these desiderata using continuous-time stochastic processes. We develop two specific models within this framework, illustrating how it can be used to capture both generative and preventative causal relationships as well as de- lays between cause and effect. We evaluate one model through a new behavioral experiment, and the other through a compar- ison to existing data. We begin with a brief overview of previous work on the role of time in human causal inference, focusing on Griffiths and Tenenbaum (2005) and Greville and Buehner (2007). We then lay out a set of desiderata for a computational frame- work for human continuous-time causal inference. We go on to describe formally how we implement these desiderata in our proposed framework. Following this, we apply the frame- work to our two case studies, evaluating models that use pre- ventative causes and delay distributions. Finally, we conclude and suggest directions for future work. Introduction Continuous-time causal induction Causal induction plays a key role in human cognition, allow- ing people to identify the causal relationships that structure their environment. Recent work in cognitive science has re- sulted in many successful models of how people infer causal relationships from contingency data (Anderson, 1990; An- derson & Sheu, 1995; Cheng, 1997; Griffiths & Tenenbaum, 2005) and events that unfold in discrete time (Wasserman, 1990; Greville & Buehner, 2007). However, relatively few models have explored events that occur in continuous time. And yet, people regularly and easily reason about causal phe- nomena that evolve in continuous time (Michotte, 1963; Grif- fiths & Tenenbaum, 2009). Our understanding of causal in- ference would thus benefit from a framework capable of ex- plaining human continuous-time causal inferences. In this paper, we address this challenge by undertaking a rational analysis of continuous-time causal induction, in the spirit of Anderson (1990) and Marr (1982). We formalize the abstract problem posed by continuous-time causal inference, identifying a set of desiderata that a solution to this problem needs to incorporate. We then outline a framework that sat- isfies these desiderata, based on rational statistical inference over continuous-time stochastic processes. Our framework makes it possible to define both generative and preventative causes that unfold in continuous time, and to take into account delays between causes and effects. With this framework in hand, we present two case stud- ies from experimental psychology on human continuous-time causal inference. The first case study involves a novel ex- periment based on an experiment conducted by Griffiths and Tenenbaum (2005), allowing us to show how our framework can be used to infer whether a cause prevents events from occurring. The second case study is a re-analysis of an ex- periment on the effects of temporal information on human causal inference that was originally conducted by Greville and Buehner (2007). This second case study demonstrates the value of being able to use delay distributions to character- ize how the effect of a cause changes over time. Studying the role of time in causal induction has a long his- tory in cognitive science. One of the earliest established findings in the study of human causal inference is our abil- ity to perceive causal relations in collisions, which is highly dependent on precise timing (Michotte, 1963). More re- cently, Buehner and colleagues have been very active in studying causal inference as it interacts with temporal infor- mation (e.g., Buehner & May, 2003; Greville & Buehner, 2007). However, while studies of causal induction often present events to participants in continuous time, they are typ- ically analyzed using the discrete trial structure which the re- searchers used to design the stimuli (e.g., Anderson & Sheu, 1995; Wasserman, 1990). Nonetheless, it may be enlightening to treat events as if they occurred over continuous time. Models considered by Griffiths and Tenenbaum (2005, 2009) take this approach, treating events that occur in continuous time as existing in a continuous dimension, or analyzing summaries of events as if they had occurred during a continuous interval. Even stim- uli that are explicitly designed to convey information in dis- crete time (e.g., Greville & Buehner, 2007) can be analyzed in terms of continuous time by integrating over the time in- tervals. In the remainder of this section we summarize results from two studies on causal induction with temporal informa- tion, providing context for our later analyses. Causal induction from rates Griffiths and Tenenbaum (2005) showed that people are ca- pable of reasoning about causes that increase the rate at which events occur over continuous time, and their judgments are in close accordance with the predictions of a computa- tional model engaging in continuous-time causal inference. In their experiments, participants observed a series of results that they were told came from physics experiments studying whether different electrical fields cause different radioactive compounds to emit more or fewer particles (the compound always released particles at some rate). For each “experi-
Grounding spatial language in non-linguistic cognition: Evidence for universal and relative spatial semantics in thought Michael Pacer* (mpacer@berkeley.edu) Alexandra Carstensen* (abc@berkeley.edu) Department of Psychology, University of California at Berkeley Berkeley, CA 94720 USA Terry Regier (terry.regier@berkeley.edu) Department of Linguistics and Cognitive Science Program, University of California at Berkeley Berkeley, CA 94720 USA similar to those expressed by in and on in English. She demonstrated that these primitives could be conjoined to form the linguistic spatial categories observed in diverse languages, accounting for both universal similarities and variation across languages. Xu and Kemp (2010) have since expanded on Feist‟s attributes, constructing their own set of universal primitives and demonstrating that conjunctions of these primitives can be used to describe a wide range of variation in the ways that languages partition the semantic space of spatial relations. Primitive-based accounts have been demonstrated to characterize spatial terms across languages, capturing distinctions in both the intensional meanings and extensional uses of spatial words. However, the central issue in this debate is about cognition—not language. It is unclear whether these primitives accurately characterize the structure of thought as well. Despite the assumption in the literature that semantic primitives are universal components of spatial cognition, these proposed units of thought have been developed and tested exclusively on the basis of linguistic data. Primitives of spatial cognition are typically inferred top-down from observations of cross-language variation in spatial terms, and evaluated on their ability to explain that variation. Importantly, this is generally done without direct reference to non-linguistic cognition. Consequently, we know little about how the semantic primitives of spatial language relate to spatial cognition. While the primitives derived from language could plausibly account for variation in spatial cognition, they clearly suggest a subsequent inquiry: would it be possible to infer spatial primitives more directly from measures of nonlinguistic cognition? If so, would these cognitive primitives similarly account for the varying semantic systems across languages? Would they reflect language- specific influences as well? Abstract The categories named by spatial terms vary considerably across languages. It is often proposed that underlying this variation is a universal set of primitive spatial concepts that are combined differently in different languages. Despite the inherently cognitive assumptions of this proposal, such spatial primitives have generally been inferred in a top-down manner from linguistic data. Here we show that comparable spatial primitives can be inferred bottom-up from non-linguistic pile- sorting of spatial stimuli by speakers of English, Dutch, and Chichewa. We demonstrate that primitives obtained in this fashion explain meaningful cross-linguistic variation in spatial categories better than primitives designed by hand for that purpose, and reflect both universal and language-specific spatial semantics. Keywords: Language and thought; spatial cognition; semantic primitives; semantic universals; linguistic relativity. Spatial language and semantic primitives 1 Languages categorize spatial relations differently, and the significance of this cross-language variation for spatial cognition is a topic of ongoing debate (e.g. Bowerman & Pederson, 1992; Feist, 2000; Hespos & Spelke, 2004; Khetarpal et al., 2010; Levinson & Meira, 2003; Majid et al., 2004; see also e.g. Boroditsky & Gaby, 2010 on spatial structuring of non-spatial domains). This debate has traditionally pitted two views against each other. On the one hand, some have argued (e.g. Majid et al., 2004) that language structures spatial cognition, such that cross- language differences in spatial categorization cause underlying cognitive differences in speakers of those languages. On the other hand, others have suggested (e.g. Levinson & Meira, 2003) that the cross-language variation may reflect different partitions of a universal underlying conceptual representation. An influential version of the universalist view holds that a set of semantic primitives (e.g. Wierzbicka, 1996) is universally available to human cognition, and that spatial categories in different languages can be obtained by composing such spatial semantic primitives in different ways. Feist (2000) proposed a set of spatial attributes characterizing cross-linguistic uses of spatial relations Spatial primitives in language and cognition To determine whether proposals of semantic primitives are supported by direct evidence from nonlinguistic cognition, we would ideally want to obtain both cognitive and linguistic data from speakers of differing languages, extract primitives from the cognitive data of each group of *The first two authors contributed equally to this work.
Previous work showed that people‟s causal judgments are modeled better as estimates of the probability that a causal relationship exists (a qualitative inference) than as estimates of the strength of that relationship (a quantitative inference). Here, using a novel task, we present experimental evidence in support of the importance of qualitative causal inference. Our findings cannot be explained through the use of parameter estimation and related quantitative inference. These findings suggest the role of qualitative inference in causal reasoning has been understudied despite its unique role in cognition. Further, we suggest these findings open interesting questions about the role of qualitative inference in many domains.
Rational models of causal induction have been successful in accounting for people's judgments about causal relationships. However, these models have focused on explaining inferences from discrete data of the kind that can be summarized in a 2 × 2 contingency table. This severely limits the scope of these models, since the world often provides non-binary data. We develop a new rational model of causal induction using continuous dimensions, which aims to diminish the gap between empirical and theoretical approaches and real-world causal induction. This model successfully predicts human judgments from previous studies better than models of discrete causal inference, and outperforms several other plausible models of causal induction with continuous causes in accounting for people's inferences in a new experiment.
What Ockham’s Razor Cuts: Quantifying simplicity in explanation choice Michael Pacer University of California, Berkeley Tania Lombrozo University of California, Berkeley Abstract: Observations from everyday life, the history of science, and well-controlled laboratory experiments suggest that when it comes to choosing between competing causal explanations, simplicity is an important factor. Less examined is the metric or metrics by which simplicity is quantified. Generally, it is assumed that simplicity can be described by counting the number of elements in an explanation, e.g. the number of causes, and this approach has been taken by fields as diverse as philosophy, psychology, and statistics. One alternative is that in the case of causal reasoning, one might consider the simpler explanation to be the one that includes the fewest unexplained causes, i.e. the fewest root nodes in the language of Bayesian causal-nets. We present two experiments supporting the hypothesis that this metric for simplicity is sometimes used in choosing between explanations, and can outweigh the total number of causes invoked in an explanation.