As in biological evolution, multiple forces are involved in cultural evolution. One force is analogous to selection, and acts on differences in the fitness of aspects of culture by influencing who people choose to learn from. Another force is analogous to mutation, and influences how culture changes over time owing to errors in learning and the effects of cognitive biases. Which of these forces need to be appealed to in explaining any particular aspect of human cultures is an open question. We present a study that explores this question empirically, examining the role that the cognitive biases that influence cultural transmission might play in universals of colour naming. In a large-scale laboratory experiment, participants were shown labelled examples from novel artificial systems of colour terms and were asked to classify other colours on the basis of those examples. The responses of each participant were used to generate the examples seen by subsequent participants. By simulating cultural transmission in the laboratory, we were able to isolate a single evolutionary force-the effects of cognitive biases, analogous to mutation-and examine its consequences. Our results show that this process produces convergence towards systems of colour terms similar to those seen across human languages, providing support for the conclusion that the effects of cognitive biases, brought out through cultural transmission, can account for universals in colour naming.
In 1969, Berlin and Kay proposed that there exist cross-cultural universals in the form of basic color terms. To test this hypothesis, the World Color Survey (WCS) collected color naming data from 110 non-industrial societies, identifying regularities in the structure of languages with different numbers of terms. This leaves us with the question of where these universals come from. We use a simple model of cultural evolution known as "iterated learning" to explore the hypothesis that universals emerge from human perceptual and learning biases. We conducted an experiment simulating the process of cultural transmission in the laboratory, and compared the results to the systems of color terms that appear in the WCS data. Our results show that cultural evolution results in convergence of systems of color terms towards a form consistent with the WCS, supporting the hypothesis that universals are the result of perceptual and learning biases.
Berlin and Kay (1969) proposed that there exist cross-cultural universals in the form of basic color terms. To test this hypothesis, the World Color Survey (WCS) collected color-naming data from 110 non-industrial societies, identifying regularities in the structure of languages with different numbers of terms. This leaves us with the question of where these universals come from. We use a simple model of cultural evolution known as “iterated learning” (Kirby, 2001) to explore the hypothesis that universals emerge from human perceptual and learning biases. We conducted an experiment simulating the process of cultural transmission in the laboratory, and compared the results to the systems of color terms that appear in the WCS data. Our results show that cultural evolution results in convergence of systems of color terms towards a form consistent with the WCS, supporting the hypothesis that universals are the result of perceptual and learning biases.
Replicating Color Term Universals through Human Iterated Learning Jing Xu (jing.xu@berkeley.edu) Thomas L. Griffiths (tom griffiths@berkeley.edu) Department of Psychology, 3210 Tolman Hall Berkeley, CA 94720 USA Mike Dowman (Mike@ImageScope.net) ImageScope Abstract in an attempt to resolve the debate over the universality of color naming. For example, Kay and Regier (2003; Regier, Kay, & Cook, 2005) showed that the focal colors in the WCS data largely fall in similar regions to those seen in English; in another study they defined a statistical measure of “well- formedness”, and used this measure to show that observed systems of color terms correspond to a near-optimal partition of color space (Regier, Kay, & Khetarpal, 2007). The consistent cross-linguistic structure highlighted by the WCS raises a new question: Where do these universals come from? They may be a result of cultural universals that may arise from the homogeneity of biological traits and evolu- tionary paths across cultures that constrain people to consider only a limited range of color categories when learning lan- guage, thus forcing color term systems to conform to a lim- ited range of universal types (Hawkins, 1988). However, if we view language as a system culturally transmitted from generation to generation, a simpler hypothesis is that these universals may arise directly from biases that cause learners to prefer some color categorizations over others, but that do not place absolute constraints on the types of color categories that are learnable. One way to explore this hypothesis is using the iterated learning model, a simple model of cultural trans- mission in which a sequence of agents each learns from the behavior of the previous agent in the sequence (Kirby, 2001). In an iterated learning model of the transmission of systems of color terms, each agent learns a system of color terms from examples provided by another agent, and then generates ex- amples which are provided to the next agent in the sequence. Mathematical analyses of iterated learning show that as this process continues, the information being transmitted gradu- ally changes to become consistent with the learning biases of the agents involved (Griffiths & Kalish, 2007; Kirby, Dow- man, & Griffiths, 2007). If systems of color terms similar to those seen in the WCS emerge from a process of cultural In 1969, Berlin and Kay proposed that there exist cross- cultural universals in the form of basic color terms. To test this hypothesis, the World Color Survey (WCS) collected color naming data from 110 non-industrial societies, identifying reg- ularities in the structure of languages with different numbers of terms. This leaves us with the question of where these uni- versals come from. We use a simple model of cultural evo- lution known as “iterated learning” to explore the hypothesis that universals emerge from human perceptual and learning bi- ases. We conducted an experiment simulating the process of cultural transmission in the laboratory, and compared the re- sults to the systems of color terms that appear in the WCS data. Our results show that cultural evolution results in convergence of systems of color terms towards a form consistent with the WCS, supporting the hypothesis that universals are the result of perceptual and learning biases. Keywords: basic color terms; iterated learning model; color term universals; cultural evolution; Bayesian inference. Introduction Linguistic universals – properties that seem to hold across all human languages – have the potential to provide unique in- sight into the nature of human cognition. Universals in sys- tems of color terms are among the best documented of these properties. Berlin and Kay (1969) proposed that color naming systems across different cultures are based on one or more of eleven focal colors corresponding to the English color terms black, white, red, green, yellow, blue, brown, purple, pink, or- ange, and gray. Kay and McDaniel (1978) and Kay and Maffi (1999) later refined this model to emphasize the six Hering primary colors (black, white, red, green, yellow, and blue) (Hering, 1964), and to characterize the process by which so- cieties might transition from one system of color terms to an- other as new terms are introduced. The World Color Survey (WCS) was initiated in the late 1970’s to provide a more comprehensive empirical test of the univerality hypothesis (Kay, Berlin, & Merrifield, 1991; Kay, Berlin, Maffi, & Merrifield, 1997). In the WCS, a total of 330 color chips, comprised of 40 equally spaced Munsell hues at 8 levels of lightness and achromatic chips at 10 levels of light- ness (see Figure 1), were presented to speakers of 110 dif- ferent languages in non-industrial societies. Those speakers were asked to name each color chip, and also to point out the most representative chip for each color term. Later analysis of the WCS data showed that the universality hypothesis was by and large confirmed (Kay et al., 1997). Recently, several sta- tistical analyses of the WCS data have also been conducted Figure 1: The World Color Survey stimulus array.
There is an ongoing debate as to whether the words in early presyntactic forms of human language had simple atomic meanings like modern words, or whether they were holophrastic. Simulations were conducted using an iterated learning model in which the agents were able to associate words with meanings, but in which they were not able to use syntactic rules to combine words into phrases or sentences. In some of these simulations words emerged that had neither holophrastic nor atomic meanings, demonstrating the possibility of protolanguages intermediate between these two extremes. Further simulations show how increases in cognitive or articulatory capacity would have produced changes in the type of words that was dominant in protolanguages. It is likely that at some point in time humans spoke a protolanguage in which most words had neither holophrastic nor atomic meanings.
An ongoing debate concerns whether the words in protolanguages expressed single atomic concepts (Bickerton, 1990, Tallerman, 2007), or whether they were holophrastic (Wray, 1998; Arbib, 2005). Here we suggest that there is no clear distinction between holophrastic and atomic meanings, as there is no clear definition of what level of conceptualization is atomic. We show that there is a continuum between holophrastic words and words denoting single concepts, depending on how narrow a range of meanings each word denotes. Using a computer model, we show that the type of words occurring in protolanguages could have changed over time, and that protolanguages could have contained a mixture of words of differing degrees of holophrasticity. We must therefore take into account these alternative possibilities when considering the nature of protolanguage. Holophrastic words convey complex meanings comprised of several constituent concepts, while words in modern languages are said to express single concepts. However, when we compare different languages we often find that words for some domain have much narrower and more specific denotations in one language than in another. For example, while in English we have the word brother, Japanese has separate words for younger brother (otouto) and older brother (mi) , while German has a single word meaning brother or sister (geschwister). This suggests that the English and Japanese words are in fact multi-concept holophrases (MALE-SIBLING and YOUNGEWOLDER-MALESIBLING respectively). A similar situation is seen within languages when one word expresses a more specific meaning than another. Consider for example English die, kill, murder and strangle, where each successive word conveys somewhat more information. Is strangle therefore a holophrase for ‘Illegally cause to die by choking’, or are both DIE and STRANGLE atomic concepts with overlapping denotations? Furthermore, some of the holophrases that have
In order to determine the points at which meeting discourse changes from one topic to another, probabilistic models were used to approximate the process through which meeting transcripts were produced. Gibbs sampling was used to estimate the values of random variables in the models, including the locations of topic boundaries. This paper shows how discourse features were integrated into the Bayesian model and reports empirical evaluations of the benefit obtained through the inclusion of each feature and of the suitability of alternative models of the placement of topic boundaries. It demonstrates how multiple cues to segmentation can be combined in a principled way, and empirical tests show a clear improvement over previous work.
There is an ongoing debate about whether the words in the first languages spoken by humans expressed single concepts or complex holophrases. A computer model was used to investigate the nature of the protolanguages that would arise if speakers could associate words and meanings, but lacked any productive ability beyond saying the word whose past uses most closely matched the meaning that they wished to express. It was found that both words expressing single concepts, and holophrastic words could arise, depending on the conceptual and articulatory abilities of the agents. However, most words were of an intermediate type, as they expressed more than a single concept but less than a holophrase. The model therefore demonstrates that protolanguages may have been of types that are not usually considered in the debate over the nature of the first human languages.
Human language arises from biological evolution, individual learning, and cultural transmission, but the interaction of these three processes has not been widely studied. We set out a formal framework for analyzing cultural transmission, which allows us to investigate how innate learning biases are related to universal properties of language. We show that cultural transmission can magnify weak biases into strong linguistic universals, undermining one of the arguments for strong innate constraints on language learning. As a consequence, the strength of innate biases can be shielded from natural selection, allowing these genes to drift. Furthermore, even when there is no natural selection, cultural transmission can produce apparent adaptations. Cultural transmission thus provides an alternative to traditional nativist and adaptationist explanations for the properties of human languages.
An expression-induction model was used to simulate the evolution of basic color terms to test Berlin and Kay's (1969) hypothesis that the typological patterns observed in basic color term systems are produced by a process of cultural evolution under the influence of biases resulting from the special properties of universal focal colors. Ten agents were simulated, each of which could learn color term denotations by generalizing from examples using Bayesian inference, and for which universal focal red, yellow, green, and blue were especially salient, but unevenly spaced in the perceptual color space. Conversations between these agents, in which agents would learn from one another, were simulated over several generations, and the languages emerging at the end of each simulation were investigated. The proportion of color terms of each type correlated closely with the equivalent frequencies found in the World Color Survey, and most of the emergent languages could be placed on one of the evolutionary trajectories proposed by Kay and Maffi (1999). The simulation therefore demonstrates how typological patterns can emerge as a result of learning biases acting over a period of time.
There is an ongoing debate as to whether the words in early pre-syntactic forms of human language had simple atomic meanings like modern words (Bickerton, 1990, 1996), or whether they were holophrastic (Wray, 1998, 2000). Simulations were conducted using an iterated learning model in which the agents were able to associate words with meanings, but in which they were not able to use syntactic rules to combine words into phrases or sentences. In some of these simulations words emerged which had neither holophrastic nor atomic meanings, demonstrating the possibility of protolanguages intermediate between these two extremes. Further simulations show how increases in cognitive or articulatory capacity would have produced changes in the type of words that were dominant in protolanguages. It is likely that at some point in time humans spoke a protolanguage in which most words had neither holophrastic nor atomic meanings.
The effect of adding noise to an expression-induction model of language evolution was investigated. The model consisted of a number of artificial people who were able to infer the denotation of basic colour terms from examples of colours which the words had been used to identify, using a Bayesian inference procedure. The artificial people would express colours to one-another, so producing data from which other people could learn. Occasionally they would be creative, which allowed new words to enter the language. When certain points in the colour space were made especially salient, so that the artificial people were more likely to remember colours at these points, the languages emerging over a number of generations in evolutionary simulations replicated the typological patterns seen in the 110 languages of the world colour survey. It was found that if random noise was added to the data from which the artificial people learned, this had no major effect on the emergent languages, demonstrating that the Bayesian inference procedure is able to learn effectively despite the presence of random noise, even when placed in an evolutionary context
Rich News, a system that augments news broadcasts with textual content, is described. The system identifies individual stories in news broadcasts, and annotates them with related content from the World Wide Web. The web content is subsequently semantically analysed, and used to produce summary information for each news story. This content can then be delivered to users as part of an interactive television broadcast, or used to create semantically enhanced electronic programme guides. It also enables sophisticated search and browsing of news stories via a web interface. Rich News could be deployed either by broadcasters, or on digital video recorders in viewers’ homes, and allows the creation of new personalized media experiences that integrate television and web content into one unified viewing experience.
The Rich News system, that can automatically annotate radio and television news with the aid of resources retrieved from the World Wide Web, is described. Automatic speech recognition gives a temporally precise but conceptually inaccurate annotation model. Information extraction from related web news sites gives the opposite: conceptual accuracy but no temporal data. Our approach combines the two for temporally accurate conceptual semantic annotation of broadcast news. First low quality transcripts of the broadcasts are produced using speech recognition, and these are then automatically divided into sections corresponding to individual news stories. A key phrases extraction component finds key phrases for each story and uses these to search for web pages reporting the same event. The text and meta-data of the web pages is then used to create index documents for the stories in the original broadcasts, which are semantically annotated using the KIM knowledge management platform. A web interface then allows conceptual search and browsing of news stories, and playing of the parts of the media files corresponding to each news story. The use of material from the World Wide Web allows much higher quality textual descriptions and semantic annotations to be produced than would have been possible using the ASR transcript directly. The semantic annotations can form a part of the Semantic Web, and an evaluation shows that the system operates with high precision, and with a moderate level of recall.
Valentin Tablan合作论文数Ieso Digital Health3
Hamish Cunningham合作论文数Computer Science,University of Sheffield3