The discovery that listeners more accurately identify words repeated in the same voice than in a different voice has had an enormous influence on models of representation and speech perception. Widely replicated in English, we understand little about whether and how this effect generalizes across languages. In a continuous recognition memory study with Hindi speakers and listeners (N = 178), we replicated the talker-specificity effect for accuracy-based measures (hit rate and D′), and found the latency advantage to be marginal (p = 0.06). These data help us better understand talker-specificity effects cross-linguistically and highlight the importance of expanding work to less studied languages.
Given any feasible amount of time, a talker would never be able to produce the same word twice in an identical manner. Yet recognition memory experiments have consistently used identical tokens to demonstrate that listeners recognize a word more quickly and accurately when it is repeated by the same talker than by a different talker. These talker-specificity effects have served as the foundation of decades of research in speech perception, but the use of identical tokens introduces a confound: Is it the talker or the physical stimulus that drives these effects? And consequently, to what extent do listeners encode the high-level acoustic characteristics of a talker's voice? We investigate the roles of token and talker repetition in two continuous recognition memory experiments. In Exp. 1, listeners heard the voice of one talker, with either Identical or Novel repeated tokens. In Exp. 2, listeners heard two demographically matched talkers, with same-voice repetitions being either Identical or Novel. Classic talker-specificity effects were replicated in both Identical and Novel tokens, but recognition of Identical tokens was in some cases stronger than recognition of Novel tokens. In addition, recognition memory varied across demographically matched talkers, suggesting stronger episodic encoding for one talker than for the other. We argue that novel tokens should serve as the default design for similar studies and that consideration of talker variation can advance our understanding of encoding and memory differences more broadly.
In this study, we replicated and extended Experiment 1 of Palmeri et al. (1993) in two experiments. Using the continuous recognition memory paradigm, we investigated effects of a demographically heterogeneous set of talkers varying across race, gender, and regional accent (Exp. 1) and effects of two demographically homogeneous sets of talkers (8 identifiably white male or 8 identifiably Black male talkers) across two listener populations (white and Black listeners) (Exp. 2). Words repeated in the same voice were recognized more quickly and accurately than words repeated in a different voice in both experiments, as found in the original study. This pattern is extremely robust. However, we also found differences across talker conditions, number of voices, lag, false alarms, and d’ that differ from the original study (Exp. 1). In addition, we found effects of talker, talker context, and listener population suggesting that social ideologies and experiences greatly influence the encoding of and memory for spoken words (Exp. 2).
The assignment of phrasal prominence has been variously attributed to syntactic structure, part of speech, predictability, informativity, and speaker's intent. A recent account asserts that prominence is memorized on a by-word basis as Accent Ratio (AR), the likelihood that a word is accented (Nenkova et al. 2007). We examined whether AR outperforms the traditional predictors, in particular syntax and informativity, and if not, whether the traditional predictors shed light on the variance left unexplained by AR. We used a corpus of spoken American English consisting of the first inaugural addresses of six recent American presidents, hand-annotated for stress by two native English speakers. Regression models fitted to the data revealed that AR, syntax, and informativity all independently matter. Dividing the data into high-prominence and low-prominence tokens further revealed that AR and informativity are significant among low-prominence words, but only syntax is significant among high-prominence words. We conclude that although AR is a highly successful predictor, certain aspects of phrasal prominence require reference to syntax and informativity.