The explosion of data generated during human interactions online presents an opportunity for psychologists to evaluate cognitive models outside the confines of the laboratory. Moreover, the size of these online data sets can allow researchers to construct far richer models than would be feasible with smaller in-lab behavioral data. In the current article, we illustrate this potential by evaluating 3 popular psychological models of generalization on 2 web-scale online data sets typically used to build automated recommendation systems. We show that each psychological model can be efficiently implemented at scale and in certain cases can capture trends in human judgments that standard recommendation systems from machine learning miss. We use these results to illustrate the opportunity Internet-scale data sets offer to psychologists and to underscore the importance of using insights from cognitive modeling to supplement the standard predictive-analytic approach taken by many existing machine learning approaches. (PsycInfo Database Record (c) 2020 APA, all rights reserved).
The explosion of data generated during human interactions online presents an opportunity for psychologists to evaluate cognitive models outside the confines of the laboratory. Moreover, the size of these online datasets can allow researchers to construct far richer models than would be feasible with smaller in-lab behavioral data. In the current paper we illustrate this potential by evaluating three popular psychological models of generalization on two web-scale online datasets typically used to build automated recommendation systems. We show that each psychological model can be e � ciently implemented at scale and in certain cases can capture trends in human judgments which standard recommendation systems from machine learning miss. We use these results to illustrate the opportunity internet-scale datasets o � er to psychologists, as well as to underscore the importance of using insights from cognitive modeling to supplement the standard predictive-analytic approach taken by many existing machine learning approaches.
Decades of psychological research have been aimed at modeling how people learn features and categories. The empirical validation of these theories is often based on artificial stimuli with simple representations. Recently, deep neural networks have reached or surpassed human accuracy on tasks such as identifying objects in natural images. These networks learn representations of real-world stimuli that can potentially be leveraged to capture psychological representations. We find that state-of-the-art object classification networks provide surprisingly accurate predictions of human similarity judgments for natural images, but they fail to capture some of the structure represented by people. We show that a simple transformation that corrects these discrepancies can be obtained through convex optimization. We use the resulting representations to predict the difficulty of learning novel categories of natural images. Our results extend the scope of psychological experiments and computational modeling by enabling tractable use of large natural stimulus sets.
The explosion of data generated during human interactions online presents an opportunity for cognitive scientists to evaluate their models on popular real-world tasks outside the confines of the laboratory. We demonstrate this approach by evaluating two cognitive models of generalization against two machine learning approaches to recommendation on an online dataset of over 100K human playlist selections. Across two experiments we demonstrate that a model from cognitive science can both be efficiently implemented at scale and can capture generalization trends in human recommendation judgments which neither machine learning model is capable of replicating. We use these results to illustrate the opportunity internet-scale datasets offer to cognitive scientists, as well as to underscore the importance of using insights from cognitive modeling to supplement the standard predictive-analytic approach taken by many existing machine learning approaches.
Deep neural networks have become increasingly successful at solving classic perception problems (e.g., recognizing objects), often reaching or surpassing human-level accuracy. In this abridged report of Peterson et al. [2016], we examine the relationship between the image representations learned by these networks and those of humans. We find that deep features learned in service of object classification account for a significant amount of the variance in human similarity judgments for a set of animal images. However, these features do not appear to capture some key qualitative aspects of human representations. To close this gap, we present a method for adapting deep features to align with human similarity judgments, resulting in image representations that can potentially be used to extend the scope of psychological experiments and inform human-centric AI.
Artificial neural networks have seen a recent surge in popularity for their ability to solve complex problems as well as or better than humans. In computer vision, deep convolutional neural networks have become the standard for object classification and image understanding due to their ability to learn efficient representations of high-dimensional data. However, the relationship between these representations and human psychological representations has remained unclear. Here we evaluate the quantitative and qualitative nature of this correspondence. We find that state-of-the-art object classification networks provide a reasonable first approximation to human similarity judgments, but fail to capture some of the structure of psychological representations. We show that a simple transformation that corrects these discrepancies can be obtained through convex optimization. Such representations provide a tool that can be used to study human performance on complex tasks with naturalistic stimuli, such as predicting the difficulty of learning novel categories. Our results extend the scope of psychological experiments and computational modeling of cognition by enabling tractable use of large natural stimulus sets.
The biological basis of the commonality in color lexicons across languages has been hotly debated for decades. Prior evidence that infants categorize color could provide support for the hypothesis that color categorization systems are not purely constructed by communication and culture. Here, we investigate the relationship between infants' categorization of color and the commonality across color lexicons, and the potential biological origin of infant color categories. We systematically mapped infants' categorical recognition memory for hue onto a stimulus array used previously to document the color lexicons of 110 nonindustrialized languages. Following familiarization to a given hue, infants' response to a novel hue indicated that their recognition memory parses the hue continuum into red, yellow, green, blue, and purple categories. Infants' categorical distinctions aligned with common distinctions in color lexicons and are organized around hues that are commonly central to lexical categories across languages. The boundaries between infants' categorical distinctions also aligned, relative to the adaptation point, with the cardinal axes that describe the early stages of color representation in retinogeniculate pathways, indicating that infant color categorization may be partly organized by biological mechanisms of color vision. The findings suggest that color categorization in language and thought is partially biologically constrained and have implications for broader debate on how biology, culture, and communication interact in human cognition.
Focal colors, or best examples of color terms, have traditionally been viewed as either the underlying source of cross-language color-naming universals or derived from category boundaries that vary widely across languages. Existing data partially support and partially challenge each of these views. Here, we advance a position that synthesizes aspects of these two traditionally opposed positions and accounts for existing data. We do so by linking this debate to more general principles. We show that best examples of named color categories across 112 languages are well-predicted from category extensions by a statistical model of how representative a sample is of a distribution, independently shown to account for patterns of human inference. This model accounts for both universal tendencies and variation in focal colors across languages. We conclude that categorization in the contested semantic domain of color may be governed by principles that apply more broadly in cognition and that these principles clarify the interplay of universal and language-specific forces in color naming.
How does cognition organize sparse and ambiguous input from the environment into useful representations and concepts for understanding the world? The work in this dissertation explores how people learn and reason with abstract knowledge, focusing on the kinds of processes and representations used in semantic memory. In particular, I present three case studies, each investigating different assumptions for the semantic representations and algorithms used to model cognition. The first chapter introduces this work situated in the framework of probabilistic models of cognition and outlines the goals of each case study. The second chapter focuses on the distinction be- tween process and representation in semantic memory search. The simulations and analyses in this chapter show that behavioral results on a semantic fluency task previously explained as a strategic search process can also be produced by a non-strategic search process, depending on the structure of representation used for semantic memory. The third chapter investigates semantic representations as a means to explore universals and variation in cognition across cultures. In this chapter, I present a simple computational model operating over an irregularly shaped perceptual color space which accounts both for universal tendencies and for variation in focal colors, or best examples of color terms, across the world’s languages. The fourth chapter explores the challenges of develop- ing models that can learn to appropriately apply new labels to concepts from only a few example observations, like people. Building upon a successful Bayesian word learning model, I propose adapting large-scale knowledge representations, typically used in machine learning and computer vision, to automatically construct hypothesis spaces for generalization models that account for these challenges. Finally, in the fifth chapter, I discuss the theoretical and practical implications for this body of work as a whole, and suggest a variety of future directions. Taken together, this research suggests general principles of computation over structured knowledge representations illuminates how people make sense of the world around them, and may lead to developing machines that think more like people do.
Most cognitive psychology experiments evaluate models of human cognition using a relatively small, well-controlled set of stimuli. This approach stands in contrast to current work in neuroscience, perception, and computer vision, which have begun to focus on using large databases of natural images. We argue that natural images provide a powerful tool for characterizing the statistical environment in which people operate, for better evaluating psychological theories, and for bringing the insights of cognitive science closer to real applications. We discuss how some of the challenges of using natural images as stimuli in experiments can be addressed through increased sample sizes, using representations from computer vision, and developing new experimental methods. Finally, we illustrate these points by summarizing recent work using large image databases to explore questions about human cognition in four different domains: modeling subjective randomness, defining a quantitative measure of representativeness, identifying prior knowledge used in word learning, and determining the structure of natural categories.
When people are asked to retrieve members of a category from memory, clusters of semantically related items tend to be retrieved together. A recent article by Hills, Jones, and Todd (2012) argued that this pattern reflects a process similar to optimal strategies for foraging for food in patchy spatial environments, with an individual making a strategic decision to switch away from a cluster of related information as it becomes depleted. We demonstrate that similar behavioral phenomena also emerge from a random walk on a semantic network derived from human word-association data. Random walks provide an alternative account of how people search their memories, postulating an undirected rather than a strategic search process. We show that results resembling optimal foraging are produced by random walks when related items are close together in the semantic network. These findings are reminiscent of arguments from the debate on mental imagery, showing how different processes can produce similar results when operating on different representations.
Empirical Evidence for Markov Chain Monte Carlo in Memory Search David D. Bourgin (ddbourgin@berkeley.edu) Joshua T. Abbott (joshua.abbott@berkeley.edu) Thomas L. Griffiths (tom griffiths@berkeley.edu) Department of Psychology University of California Berkeley Berkeley, CA 94720 USA Abstract Kevin A. Smith (k2smith@ucsd.edu) Edward Vul (evul@ucsd.edu) Department of Psychology University of California San Diego La Jolla, CA 92093 USA mote Associates Test (RAT; S. A. Mednick, 1962), a classic multiply-constrained search task. Building on recent work using random walks on semantic networks to model semantic fluency (Griffiths, Steyvers, & Firl, 2007) and singly-constrained search in memory (Abbott, Austerweil, & Griffiths, 2012), we consider a range of MCMC-inspired models for search on the RAT. We find that human response patterns are well-matched by a model closely related to the Metropolis-Hastings (M-H) algorithm, provid- ing empirical support for MCMC as a candidate for modeling creative ideation. The plan for the paper is as follows. We first review pre- vious work on human response patterns on the Remote Asso- ciates Test and discuss existing theories of search on this task. We then introduce Markov chain Monte Carlo methods with an emphasis on their application to creative search. We end by formally introducing the model used in the current study, and offer an analysis of its behavior in relation to human data collected by Smith, Huber, and Vul (2013). Previous theoretical work has proposed the use of Markov chain Monte Carlo as a model of exploratory search in mem- ory. In the current study we introduce such a model and eval- uate it on a semantic network against human performance on the Remote Associates Test (RAT), a commonly used creativ- ity metric. We find that a family of search models closely re- sembling the Metropolis-Hastings algorithm is capable of re- producing many of the response patterns evident when human participants are asked to report their intermediate guesses on a RAT problem. In particular we find that when run our model produces the same response clustering patterns, local depen- dencies, undirected search trajectories, and low associative hi- erarchies witnessed in human responses. Keywords: Creativity; Remote Associates Test; Information retrieval; Semantic networks; Markov chain Monte Carlo. Introduction Exploratory search in memory is a major component of cre- ative problem solving. Often, the demands of a task impose constraints on the solution space: when searching for a word to rhyme with A that also means B, a would-be poet is un- likely to include items that mean T or that rhyme with Q (as- suming that Q does not rhyme with A and T is not a synonym of B) in her search. Similarly, an inventor would be remiss to consider inventions that do not satisfy certain requirements (e.g., novelty, usefulness) if she plans to pursue a patent. The current study proposes a formal model for the process by which people perform this type of exploratory search under multiple constraints. Locating relevant pieces of information in memory re- quires a strategy for quickly traversing the space of poten- tial solutions. Markov chain Monte Carlo (MCMC) methods have been a particularly useful in this regard, offering effi- cient statistical techniques for exploring spaces that would otherwise prove computationally intractable to traverse fully (Brooks, 1998). MCMC methods have been employed to model phenomena as diverse as theory change (Ullman, Goodman, & Tenenbaum, 2012), perceptual multistability (Gershman, Vul, & Tenenbaum, 2012), and conceptual flu- idity (Gabora, 2000), making them a popular tool within sta- tistical and computational cognitive modeling. MCMC meth- ods have also been proposed within the creativity literature to model search behavior in memory (e.g., Martindale, 1995; Paulus, Levine, Brown, Minai, & Doboli, 2010), although to date there exists little empirical evidence on which to evaluate these proposals. In the current study we aim to fill this gap by evaluating a formal model of MCMC search on the Re- Background Stage models of problem solving propose that individuals first search through memory to identify potential answers to a problem and then test those candidate answers against the constraints of the problem to determine acceptability (Gruenewald & Lockhead, 1980; Raaijmakers & Shiffrin, 1981). In the current paper it is assumed that these search and test processes constitute qualitatively distinct cognitive tasks. Thus in reviewing recent work on the RAT and mem- ory search we emphasize accounts that focus not only on the identification of correct answers but also on the path by which these answers are selected. The Remote Associates Test In the Remote Associates Test participants are shown a set of three cue words (e.g., ‘surprise,’ ‘line,’ ‘birthday’) and in- structed to generate a fourth word to relate them (in this case, ‘party’). Although originally introduced to test for individual differences creative ability (S. A. Mednick, 1962), early work established that performance on the RAT correlated with IQ (M. T. Mednick & Andrews, 1967) and originality during brainstorming (Forbach & Evans, 1981). More recently the RAT has also been used to measure the effects of manipu- lations related to creative ability, including intuition and in- cubation (Bowers, Regehr, Balthazard, & Parker, 1990; Vul & Pashler, 2007), the role of affect during problem solving
Learning a visual concept from a small number of positive examples is a significant challenge for machine learning algorithms. Current methods typically fail to find the appropriate level of generalization in a concept hierarchy for a given set of visual examples. Recent work in cognitive science on Bayesian models of generalization addresses this challenge, but prior results assumed that objects were perfectly recognized. We present an algorithm for learning visual concepts directly from images, using probabilistic predictions generated by visual classifiers as the input to a Bayesian generalization model. As no existing challenge data tests this paradigm, we collect and make available a new, large-scale dataset for visual concept learning using the ImageNet hierarchy as the source of possible concepts, with human annotators to provide ground truth labels as to whether a new image is an instance of each concept using a paradigm similar to that used in experiments studying word learning in children. We compare the performance of our system to several baseline algorithms, and show a significant advantage results from combining visual classifiers with the ability to identify an appropriate level of abstraction using Bayesian generalization.
Approximating Bayesian inference with a sparse distributed memory system Joshua T. Abbott (joshua.abbott@berkeley.edu) Jessica B. Hamrick (jhamrick@berkeley.edu) Thomas L. Griffiths (tom griffiths@berkeley.edu) Department of Psychology, University of California, Berkeley, CA 94720 USA Abstract we take on this challenge by showing that an associative memory using sparse distributed representations can be used to approximate Bayesian inference, producing behavior consistent with a structured statistical model while using distributed representations of the kind normally associated with artificial neural networks. Probabilistic models of cognition have enjoyed recent success in explaining how people make inductive inferences. Yet, the difficult computations over structured representations that are often required by these models seem incompatible with the continuous and distributed nature of human minds. To recon- cile this issue, and to understand the implications of constraints on probabilistic models, we take the approach of formalizing the mechanisms by which cognitive and neural processes could approximate Bayesian inference. Specifically, we show that an associative memory system using sparse, distributed represen- tations can be reinterpreted as an importance sampler, a Monte Carlo method of approximating Bayesian inference. This ca- pacity is illustrated through two case studies: a simple letter reconstruction task, and the classic problem of property induc- tion. Broadly, our work demonstrates that probabilistic mod- els can be implemented in a practical, distributed manner, and helps bridge the gap between algorithmic- and computational- level models of cognition. Keywords: Bayesian inference, importance sampling, rational process models, associative memory models, sparse distributed memory Introduction Probabilistic models of cognition can be used to explain the complex inductive inferences people make every day, such as identifying the content of images or learning new con- cepts from limited evidence (Griffiths, Chater, Kemp, Per- fors, & Tenenbaum, 2010; Tenenbaum, Kemp, Griffiths, & Goodman, 2011). However, these models are typically for- mulated at what Marr (1982) called the computational level, focusing on the abstract problems people have to solve and their ideal solutions. As a result, they explain why people behave the way they do, rather than how cognitive and neu- ral processes support these behaviors. This approach is thus quite different from previous work on modeling human cog- nition, which focused on Marr’s algorithmic and implemen- tation levels, and has been criticized because it seems to im- ply that human minds and brains need to solve intractable computational problems and use structured representations (Gigerenzer & Todd, 1999; McClelland et al., 2010). Understanding the actual commitments that computational-level accounts of human cognition based on probabilistic models make at the algorithmic and im- plementation level requires considering how these levels of analysis could be bridged (Griffiths, Vul, & Sanborn, 2012). Identifying specific cognitive algorithms and neural architectures that can approximate Bayesian inference is a key step towards knowing whether it really poses an intractable problem for human minds, or whether structured representations need to be used to implement models that involve structured probability distributions. In this paper, The associative memory that we use to approximate Bayesian inference implements a Monte Carlo algorithm known as importance sampling. Previous work has shown that this algorithm can be implemented in a common psycho- logical process model – an exemplar model (Shi, Griffiths, Feldman, & Sanborn, 2010). Shi and Griffiths (2009) further demonstrated that importance sampling can be implemented with a radial basis function neural network. However, this neural network used a localist representation, in which each hypothesis considered by the model had to be represented with a single neuron – a “grandmother cell.” While this might be plausible for modeling aspects of perception in which a wide range of neurons prefer specific stimuli, it becomes less plausible for modeling complex cognitive tasks in which hy- potheses correspond to structured representations. For exam- ple, having separate neurons for each concept or causal struc- ture we consider seems implausible. We demonstrate that an associative memory that uses distributed representations – specifically, sparse distributed memory (SDM) (Kanerva, 1988, 1993) – can be used to ap- proximate Bayesian inference through importance sampling. The underlying idea is simple: we use the associative mem- ory to store and retrieve exemplars, allowing us to build on the equivalence between exemplar models and importance sam- pling. The critical advance is that this is done using dis- tributed representations, meaning that arbitrary hypotheses can be represented, and arbitrary distributions of exemplars encoded by a single architecture. We show that the SDM nat- urally implements one class of Bayesian models, and explain how to generalize it to implement a broader range of models. The plan of the paper is as follows. First, we give a brief overview of performing Bayesian inference with importance sampling and summarize how the sparse distributed memory system is implemented. Next, we formalize how importance sampling can be performed using a SDM. We provide two case studies drawn from existing literature in which we use the SDM to approximate existing Bayesian models. The first case study is a simple example involving reconstructing En- glish letters from noisy inputs, and the second is a more so- phisticated model of property induction. We conclude the pa- per with a discussion of implications and future directions.
The human mind has a remarkable ability to store a vast amount of information in memory, and an even more remarkable ability to retrieve these experiences when needed. Understanding the representations and algorithms that underlie human memory search could potentially be useful in other information retrieval settings, including internet search. Psychological studies have revealed clear regularities in how people search their memory, with clusters of semantically related items tending to be retrieved together. These findings have recently been taken as evidence that human memory search is similar to animals foraging for food in patchy environments, with people making a rational decision to switch away from a cluster of related information as it becomes depleted. We demonstrate that the results that were taken as evidence for this account also emerge from a random walk on a semantic network, much like the random web surfer model used in internet search engines. This offers a simpler and more unified account of how people search their memory, postulating a single process rather than one process for exploring a cluster and one process for switching between clusters.
Constructing a hypothesis space from the Web for large-scale Bayesian word learning Joshua T. Abbott (joshua.abbott@berkeley.edu) Joseph L. Austerweil (joseph.austerweil@gmail.com) Thomas L. Griffiths (tom griffiths@berkeley.edu) Department of Psychology, University of California, Berkeley, CA 94720 USA Abstract Bayesian generalization model. In this paper, we use this approach to show how a hypothesis space and prior can be constructed automatically from a large online database, mak- ing it possible to apply the Bayesian generalization frame- work to a wide range of naturalistic stimuli. We focus on one specific generalization problem, word learning, where peo- ple learn new words from observing a few objects that can be labeled with that word. Given that the number of possible ex- tensions of a word is essentially infinite, learning the objects referred to by a word is a very difficult inductive problem (Quine, 1975). Xu and Tenenbaum (2007) showed how the Bayesian generalization framework could be used to explain how people learn new words. However, to construct the hy- pothesis space of their Bayesian model, Xu and Tenenbaum (2007) elicited approximately 400 similarity judgments from their participants. Clearly this is not practical to extend into every domain where people learn words. Thus, word learn- ing is an appropriate setting for exploring novel methods of constructing hypothesis spaces and prior distributions. The Bayesian generalization framework has been successful in explaining how people generalize a property from a few observed stimuli to novel stimuli, across several different domains. To create a successful Bayesian generalization model, modelers typically specify a hypothesis space and prior probability distribution for each specific domain. How- ever, this raises two problems: the models do not scale beyond the (typically small-scale) domain that they were designed for, and the explanatory power of the models is reduced by their reliance on a hand-coded hypothesis space and prior. To solve these two problems, we propose a method for deriving hypothesis spaces and priors from large online databases. We evaluate our method by constructing a hypothesis space and prior for a Bayesian word learning model from WordNet, a large online database that encodes the semantic relationships between words as a network. After validating our approach by replicating a previous word learning study, we apply the same model to a new experiment featuring three additional taxonomic domains (clothing, containers, and seats). In both experiments, we found that the same automatically constructed hypothesis space explains the complex pattern of generalization behavior, producing accurate predictions across a total of six different domains. We propose a method for automatically constructing the hypothesis space and prior distribution of a Bayesian word learning model using freely available online resources. In particular, we use WordNet (Fellbaum, 2010; Miller, 1995) as an initial source for automatically creating the hypothesis space, and ImageNet (Deng et al., 2009) as a source of natu- ralistic images that can be used as stimuli to test the resulting model in behavioral experiments. WordNet is a popular lexi- cal database of English comprised of over 100,000 relational sets of synonyms. ImageNet is a large ontology of images conforming to the hierarchical structure of WordNet, with the aim of providing over 500 high-quality images per noun in WordNet. These resources allow us to construct hypothesis spaces and prior distributions for word learning without elic- iting a single judgment from participants and test the result- ing model on a much larger scale than was previously pos- sible. We demonstrate that the Bayesian model formulated from WordNet captures participant judgments in two behav- ioral experiments, addressing the practical and theoretical is- sues with Bayesian models discussed earlier. Keywords: generalization; concept learning; word learning; Bayesian modeling; online databases Introduction Many problems solved by the mind conform to the same ab- stract computational formulation: How should a property be generalized to novel stimuli from a set of stimuli observed to have the property? As there are many ways to extend the property that are consistent with some observed evidence, these are problems of induction, where the evidence con- strains, but does not determine, the solution to a problem. The Bayesian generalization framework (Shepard, 1987; Tenen- baum & Griffiths, 2001) has been remarkably successful at explaining human generalization behavior in a wide range of domains. However, its success is largely dependent on the choice of a hypothesis space and a prior probability distribu- tion on hypotheses, which are usually hand constructed by the researcher for each specific problem. This is unsatisfy- ing practically, because the models do not scale beyond the originally modeled problem, and theoretically, as it is unclear whether their success is due to the cleverness of the modeler and not because of a deep mathematical property of the com- putational problem that people solve. One possible solution is to use existing sources of infor- mation about the organization of a domain as the basis for specifying a hypothesis space and prior. This helps address both the practical and the theoretical concerns raised by the The plan of the rest of the paper is as follows. In the next sections we review the Bayesian generalization model and then examine how Xu and Tenenbaum (2007) constructed the hypothesis space for their Bayesian word learning model. We then show how to build a hypothesis space from WordNet that can be used to evaluate word learning models on a large scale. Afterwards, we present two experiments utilizing this hypoth-
Predicting focal colors with a rational model of representativeness Joshua T. Abbott (joshua.abbott@berkeley.edu) Department of Psychology, University of California, Berkeley, CA 94720 USA Terry Regier (terry.regier@berkeley.edu) Department of Linguistics, Cognitive Science Program, University of California, Berkeley, CA 94720 USA Thomas L. Griffiths (tom griffiths@berkeley.edu) Department of Psychology, University of California, Berkeley, CA 94720 USA Abstract They argued that color categories are not defined around uni- versal foci, but are instead defined at their boundaries by local linguistic convention, which varies across languages. They proposed: “Once a category has been delineated at the bound- aries, exposure to exemplars may lead to the abstraction of a central tendency so that observers behave as if their categories have prototypes” (p. 395). On this view, best examples do not reflect a universal cognitive or perceptual substrate, but are merely an after-effect of category construction by language: best examples are derived from language-specific boundaries, rather than boundaries from universal best examples. A proposal by Jameson and D’Andrade (1997) has the po- tential to reconcile these two opposed stances. This proposal holds that there are genuine universals of color naming, but they do not stem from a small set of focal colors. Instead, universals of color naming may stem from irregularities in the overall shape of perceptual color space, which is partitioned into categories by language in a near-optimally informative way. This proposal has been shown to explain universal ten- dencies in the boundaries of color categories (Regier et al., 2007). However it has not yet provided an account of best examples of these categories, which lie at the heart of the de- bate. Here, we address this open issue, completing the reconcil- iation of the two standardly opposed views. We suggest that best examples are largely universal (in line with the universal- foci view), but nonetheless derived from category boundaries (in line with the relativist view). Specifically, given the in- dependent explanation of category boundaries in terms of the shape of color space, we propose that universal tendencies of best examples are derived from those of boundaries, rather than the other way around as has been traditionally assumed. Moreover, we propose that best examples are derived from category boundaries in an optimal manner, echoing the opti- mal or near-optimal partition of color space into categories. To pursue this idea, we draw on previous work on a rational model of representativeness, and ask whether the best exam- ples of color categories can be well-predicted as those colors that are most representative of a given category. The remainder of the paper proceeds as follows. In the next section, we discuss the previous work on representative- ness on which we draw, and contrast it with other approaches to that problem. We then describe the color naming data we consider, and a set of competing models that predict the foci Best examples of categories lie at the heart of two major de- bates in cognitive science, one concerning universal focal col- ors across languages, and the other concerning the role of rep- resentativeness in inference. Here we link these two debates. We show that best examples of named color categories across 110 languages are well-predicted by a rational model of repre- sentativeness, and that this model outperforms several natural competitors. We conclude that categorization in the contested semantic domain of color may be governed by general princi- ples that apply more broadly in cognition, and that these prin- ciples clarify the interplay of universal and language-specific forces in color naming. Keywords: Language and perception; semantic universals; color naming; representativeness; Bayesian inference. Introduction Do the world’s languages reflect a universal repertoire of cog- nitive and perceptual categories? Or do different languages partition the experienced world in fundamentally different ways? These questions have been pursued in depth in the do- main of color naming and cognition (e.g. Berlin & Kay, 1969; Kay & McDaniel, 1978; Lindsey & Brown, 2006; Rober- son, Davidoff, Davies, & Shapiro, 2005; Roberson, Davies, & Davidoff, 2000), and current findings suggest an inter- estingly mixed picture. There are clear universal tendencies of color naming across languages, but there is also substan- tial cross-language variation (e.g. Regier, Kay, & Khetarpal, 2007), more than is suggested by traditional universalist ac- counts. At the center of this debate is the disputed role of focal colors, or the best examples of named color categories. It has long been claimed that color naming across lan- guages is constrained by six universal privileged points, or foci, in color space, corresponding to the best examples of what would be described in English as white, black, red, yel- low, green, and blue. This view has received empirical sup- port: the best examples of color terms across languages tend to cluster near these six points (Berlin & Kay, 1969; Regier, Kay, & Cook, 2005), and these colors have also been found to be cognitively privileged (Heider, 1972; but see Roberson et al., 2000). A natural and influential proposal (Kay & Mc- Daniel, 1978) is that these privileged colors constitute a uni- versal foundation for color naming, such that languages dif- fer in their color naming systems primarily by grouping these universal foci together into categories in different ways. Roberson et al. (2000) advanced a diametrically opposed view of color naming, and of the role of best examples in it.
Learning the meaning of a novel noun from a few labeled objects is one of the simplest aspects of learning a language, but approximating human performance on this task is still a significant challenge for current machine learning systems. Current methods typically fail to find the appropriate level of generalization in a concept hierarchy for a given visual stimulus. Recent work in cognitive science on Bayesian models of word learning partially addresses this challenge, but it assumes that the labels of objects are given (hence no object recognition) and it has only been evaluated in small domains. We present a system for learning nouns directly from images, using probabilistic predictions generated by visual classifiers as the input to Bayesian word learning, and compare this system to human performance in an automated, large-scale experiment. The system captures a significant proportion of the variance in human responses. Combining the uncertain outputs of the visual classifiers with the ability to identify an appropriate level of abstraction that comes from Bayesian word learning allows the system to outperform alternatives that either cannot deal with visual stimuli or use a more conventional computer vision approach.