Regular incidental exposure to a language that someone does not speak can allow a person to build a proto-lexicon of the language - a set of implicitly stored word forms without semantic knowledge. Previous research shows that non-speakers of M & amacr;ori in New Zealand (NZ) have well-developed phonotactic and proto-lexical knowledge of M & amacr;ori. However, we still do not know whether exposure in early childhood is completely necessary for the formation of the proto-lexicon, nor how much proto-lexical knowledge develops from exposure in childhood versus exposure in adulthood. It is also unclear how robust these implicit memories are. To address these gaps, we examine the extent to which childhood exposure and/or ongoing consistent adult exposure are necessary conditions for the creation and/or maintenance of a proto-lexicon by comparing four groups (expats, migrants, local adolescents and local adults). The results reveal that migrants' proto-lexicon is greater than that of local adolescents, indicating that early childhood exposure to the language is not necessary. In addition, expats retained robust memory for their proto-lexical knowledge, acquired before moving overseas. Thus, implicit learning in adulthood plays a significant role and a proto-lexicon can be built and maintained over the course of the entire lifespan.
Previous work has demonstrated that New Zealanders who do not speak Māori but are regularly exposed to the language develop implicit knowledge of it. The core of this knowledge, it has been argued, is the 'proto-lexicon'-a set of stored word-forms, without associated meaning, which yields subsequent Māori phonotactic and morphological knowledge. Previous research shows that having a proto-lexicon gives learners a head start in learning Māori word meanings in formal education. We investigate experimentally whether the proto-lexicon confers an advantage for attaching meanings to words. In Experiment 1, non-Māori-speaking New Zealanders were tested on their ability to identify meanings of Māori words in a forced-choice definition task, and they did this relatively well. Then, words with low accuracy were selected for Experiment 2, where non-Māori-speaking New Zealanders and non-New Zealanders were asked to learn meanings for Māori words and nonwords. New Zealanders performed better, indicating that familiarity with Māori word shapes confers an advantage. However, they showed no greater advantage for real words over nonwords. If these words are definitely in the proto-lexicon, then this would suggest that knowledge of individual word-forms does not, in fact, confer an advantage. In Experiment 3, we therefore explore whether the words in Experiment 2 are actually robustly in the participants' proto-lexicon, by running a word identification task with the same participants. These words were not robustly distinguished from nonwords. By selecting words for their lack of semantic knowledge, we also inadvertently selected words that do not appear to be in the proto-lexicon. Together, our results indicate that different levels of semantic knowledge exist for different words, even when we consider only words that cannot confidently be said to be in a full lexicon. The results suggest that the claim of previous studies that the proto-lexicon is 'without semantics' may be oversimplified.
This study examines the production of the te reo Māori opening vowel sequences /ia/ /ea/ /oa/ and /ua/, which have been described as hiatuses. We re-examine this classification through acoustic phonetic analysis, using data from three generations of speakers in the MAONZE corpus. We propose a novel bottom-up approach which reveals that the opening sequences are not uniformly hiatus-like as described. Rather, there are structured patterns of variation among duration, formant and intensity trajectories, which form a clear continuum of variation between hiatus- and diphthong-like realisations. All generations of speakers show structured variation in their production of the sequences. We also find that there are word position and morpheme boundary effects on the production of the sequences, and that there is evidence that they have become decoupled from the monophthongs over time. These findings have several implications for how the phonology of Māori is described. The paper also makes a broader methodological contribution,as the methods used do not require presupposing categories like 'hiatus' and `diphthong'. Instead, they allow us to examine patterns of variation between trajectories (formants, intensity) and other measures like duration, without pre-supposing what kind of variation is present in the data.
This study examines the vowel sequences /ai ae ao au ei oi ou/ in te reo Māori, the indigenous language of Aotearoa New Zealand. Descriptions state that these closing sequences are comprised of separate phonemes, which are realised as surface diphthongs within morphemes, but as hiatuses when they straddle a morpheme boundary. We examine whether this description is borne out through acoustic analysis of data from the MAONZE corpus, which contains recordings from three generations of Māori speakers. Using the approach developed in Culhane et al. (In press. Variation in the production of te reo Māori opening vowel sequences. Laboratory Phonology), we find that for all sequences, there is systematic covariation between duration, intensity, and formant trajectories, which form a spectrum from hiatus-like to diphthong-like realisations. There is some evidence for the morpheme boundary effect described, but also for effects of stress and word position. We find that production of closing sequences has changed over time, with increasingly diphthong-like productions. Our analysis reveals that these patterns of change are likely systematic, whereby the change towards more diphthongised sequences reflects a change in the phonetics-phonology system more broadly.
Psycholinguistic research has traditionally relied on human ratings for stimulus norming, but whether large language models (LLMs) can reliably replace human ratings remains uncertain. This study compares human participants and three LLMs-one proprietary model (ChatGPT-4o) and two open-source models (LLaMA-3.3-70B and DeepSeek-V3)-with respect to their statistical knowledge of English binomials. For each binomial, we obtained ratings of frequency, dispersion, forward association strength, and backward association strength from 34 human participants and from 30 output samples per LLM. We examined rating-to-corpus consistency (consistencyrating2corpus), the sensitivity of statistical ratings to corpus data, and the influence of other psycholinguistic factors on ratings. All LLMs' statistical knowledge broadly mirrored that of humans. Ratings from both groups were sensitive to corpus data but not fully consistent with it. Frequency showed the highest consistencyrating2corpus, whereas dispersion showed the lowest consistencyrating2corpus and the weakest sensitivity. LLM ratings were also influenced by word-level cues. Nonetheless, LLM ratings showed greater consistencyrating2corpus, heightened sensitivity, and stronger reliance on other psycholinguistic cues than human ratings. Overall, while LLMs' performance generally aligned with that of humans, their internal statistical representations differed significantly from human cognition. The three LLMs also showed variation in their rating behavior. Thus, although multi-LLM ratings can aid pilot studies in psycholinguistics, they should not replace human ratings in formal experiments.
Statistical regularities can be acquired from usage. To examine language speakers' statistical metacognition about multiword expressions (MWEs), we collected ratings for frequency, dispersion, and directional association strength of English binomials from L1, advanced and intermediate L2 speakers. Mixed-effects modeling showed all speakers had limited speaker-to-corpus consistency but significant sensitivity to statistical regularities of language, supporting usage-based (Gries & Ellis, 2015) and statistical learning theories (Christiansen, 2019). Their statistical metacognition was also shaped by word-level cues, consistent with dual-route model (Carrol & Conklin, 2014). Despite similarities, frequency metacognition showed the strongest speaker-to-corpus consistency, while dispersion metacognition was the hardest to develop. Advanced L2 speakers showed the greatest speaker-to-corpus consistency and sensitivity, while lower-proficiency speakers relied more on word-level cues in metacognitive judgments, supporting the shallow-structure hypothesis (Clahsen & Felser, 2006). Overall, L1 and L2 speakers develop diverse statistical metacognition, with L2 speakers not necessarily inferior, suggesting that statistical metacognition is not solely shaped by usage-based experience.
Recent findings show adult New Zealanders who do not speak te reo M & amacr;ori (the M & amacr;ori language, the indigenous language of New Zealand) nonetheless have impressive implicit lexical and phonotactic knowledge of the language. These findings have been interpreted as showing that regular ambient exposure to a non-native language develops an implicit "proto-lexicon", a memory store of lexical forms in that language, without any meaning. However, what is not known is the timeframe over which this knowledge is acquired. Does the knowledge stem exclusively from implicit learning during childhood, or does it continue to grow based on exposure during adulthood? To investigate this question, we directly compare non-M & amacr;ori-speaking school-aged adolescents and adults in New Zealand and explore how age affects the degree of observed knowledge. Our results show that ambient exposure leads to implicit knowledge both in childhood and adulthood, and that continuing exposure throughout the lifespan leads to increased knowledge.
Previous studies report that exposure to theMaori language on a regular basis allows New Zealand adults who cannot speak Maori to build a proto-lexicon of Maori-an implicit memory of word forms without detailed knowledge of meaning. How might this knowledge feed into explicit language learning? Is it possible to "awaken" the proto-lexicon in the context of overt language learning? We investigate whether implicit linguistic knowledge represented in a proto-lexicon gives any advantages for intentional language learning in a tertiary educational environment. We conducted a three-task experiment which: (a) assessed participants' Maori proto-lexicon, (b) assessed their phonotactic knowledge, and (c) tested them on Maori vocabulary that they had been exposed to during the course at two time points. The results show that students with larger Maori proto-lexicons learn more words in a classroom setting. This study shows that proto-lexicon acquired from ambient exposure can lead to significant benefits in language learning.
Previous research has shown that non-Māori Speaking New Zealanders have extensive latent knowledge of Māori, despite not being able to speak it. This knowledge plausibly derives from a memory store of Māori forms (Oh et al., 2020; Panther et al., 2023). Modelling suggests that this ‘proto-lexicon’ includes not only Māori words, but also word-parts; however, this suggestion has not yet been tested experimentally. We present the results of a new experiment in which non-Māori speaking New Zealanders and non-New Zealanders were asked to segment a range of Māori words into parts. We show that the degree to which segmentations of non-Māori speakers correlate to the segmentations of two fluent speakers of Māori is stronger among New Zealanders than non-New Zealanders. This research adds to the growing evidence that even in a largely ‘monolingual’ population, there is evidence of latent bilingualism through long-term exposure to a second language.
Non-M\=aori-speaking New Zealanders (NMS)are able to segment M\=aori words in a highlysimilar way to fluent speakers (Panther et al.,2024). This ability is assumed to derive through the identification and extraction of statistically recurrent forms. We examine this assumption by asking how NMS segmentations compare to those produced by Morfessor, an unsupervised machine learning model that operates based on statistical recurrence, across words formed by a variety of morphological processes. Both NMS and Morfessor succeed in segmenting words formed by concatenative processes (compounding and affixation without allomorphy), but NMS also succeed for words that invoke templates (reduplication and allomorphy) and other cues to morphological structure, implying that their learning process is sensitive to more than just statistical recurrence.
Most non-Māori-speaking New Zealanders are regularly exposed to Māori throughout their lives without seeming to build any extensive Māori lexicon; at best, they know a small number of words which are frequently used and sometimes borrowed into English. Here, we ask how many Māori words non-Māori-speaking New Zealanders know, in two ways: how many can they identify as real Māori words, and how many can they actively define? We show that non-Māori-speaking New Zealanders can readily identify many more Māori words than they can define, and that the number of words they can reliably define is quite small. This result adds crucial support to the idea presented in earlier work that non-Māori-speaking New Zealanders have implicit form-based (proto-lexical) knowledge of many Māori words, but explicit semantic (lexical) knowledge of few. Building on this distinction, we further ask how different levels of word knowledge modulate effects of phonotactic probability on the accessing of that knowledge, across both tasks and participants. We show that participants’ implicit word knowledge leads to effects of phonotactic probability–and related effects of neighbourhood density–in a word/non-word discrimination task, but not in a more explicit task that requires the active definition of words. Similarly, we show that the effects of phonotactic probability on word/non-word discrimination are strong among participants who appear to lack explicit word knowledge, as indicated by their weak discrimination performance, but absent among participants who appear to have explicit word knowledge, as indicated by their strong discrimination performance. Together, these results suggest that phonotactic probability plays its strongest roles in the absence of explicit semantic knowledge.
Most people in New Zealand are exposed to the Māori language on a regular basis, but do not speak it. It has recently been claimed that this exposure leads them to create a large proto-lexicon, consisting of implicit memories of words and word parts, without semantic knowledge. This yields sophisticated phonotactic knowledge (Oh et al., 2020). This claim was supported by two tasks in which Non-Māori-Speaking New Zealanders: (i) Distinguished real words from phonotactically matched non-words, suggesting lexical knowledge; (ii) Gave wellformedness ratings of non-words almost indistinguishable from those of fluent Māori speakers, demonstrating phonotactic knowledge.Oh et al. (2020) ran these tasks on separate participants. While they hypothesised that phonotactic and lexical knowledge derived from the proto-lexicon, they did not establish a direct link between them. We replicate the two tasks, with improved stimuli, on the same set of participants. We find a statistically significant link between the tasks: Participants with a larger proto-lexicon (evidenced by performance in the Word Identification Task) show greater sensitivity to phonotactics in the Wellformedness Rating Task. This extends the previously reported results, increasing the evidence that exposure to a language you do not speak can lead to large-scale implicit knowledge about that language.
Given any feasible amount of time, a talker would never be able to produce the same word twice in an identical manner. Yet recognition memory experiments have consistently used identical tokens to demonstrate that listeners recognize a word more quickly and accurately when it is repeated by the same talker than by a different talker. These talker-specificity effects have served as the foundation of decades of research in speech perception, but the use of identical tokens introduces a confound: Is it the talker or the physical stimulus that drives these effects? And consequently, to what extent do listeners encode the high-level acoustic characteristics of a talker's voice? We investigate the roles of token and talker repetition in two continuous recognition memory experiments. In Exp. 1, listeners heard the voice of one talker, with either Identical or Novel repeated tokens. In Exp. 2, listeners heard two demographically matched talkers, with same-voice repetitions being either Identical or Novel. Classic talker-specificity effects were replicated in both Identical and Novel tokens, but recognition of Identical tokens was in some cases stronger than recognition of Novel tokens. In addition, recognition memory varied across demographically matched talkers, suggesting stronger episodic encoding for one talker than for the other. We argue that novel tokens should serve as the default design for similar studies and that consideration of talker variation can advance our understanding of encoding and memory differences more broadly.
We develop and probe a model for detecting the boundaries of prosodic chunks in untranscribed conversational English speech. The model is obtained by fine-tuning a Transformer-based speech-to-text (STT) model to integrate the identification of Intonation Unit (IU) boundaries with the STT task. The model shows robust performance, both on held-out data and on out-of-distribution data representing different dialects and transcription protocols. By evaluating the model on degraded speech data, and comparing it with alternatives, we establish that it relies heavily on lexico-syntactic information inferred from audio, and not solely on acoustic information typically understood to cue prosodic structure. We release our model as both a transcription tool and a baseline for further improvements in prosodic segmentation.
Recent work shows that ambient exposure in everyday situations can yield implicit knowledge of a language that an observer does not speak. We replicate and extend this work in the context of Spanish in California and Texas. In Word Identification and Wellformedness Rating experiments, non-Spanish-speaking Californians and Texans show implicit lexical and phonotactic knowledge of Spanish, which may be affected by both language structure and attitudes. Their knowledge of Spanish appears to be weaker than New Zealanders' knowledge of Māori established in recent work, consistent with structural differences between Spanish and Māori. Additionally, the strength of a participant's knowledge increases with the value they place on Spanish and its speakers in their state. These results showcase the power and generality of statistical learning of language in adults, while also highlighting how it cannot be divorced from the structural and attitudinal factors that shape the context in which it occurs.
This study investigates the clustering of words into Part-of-Speech (POS) classes in Kolyma Yukaghir. In grammatical descriptions, lexical items are assigned to POS classes based on their morphological paradigms. Discursively, however, these classes share a fair amount of morphology. In this study, we turn to POS induction to evaluate if classes based on quantification of the distributions in which roots and affixes are used can be useful for language description purposes, and, if so, what those classes might be. We qualitatively compare clusters of roots and affixes based on four different definitions of their distributions. The results show that clustering is more reliable for words that typically bear more morphology. Additionally, the results suggest that the number of POS classes in Kolyma Yukaghir might be smaller than stated in current descriptions. This study thus demonstrates how unsupervised learning methods can provide insights for language description, particularly for highly inflectional languages.
We present an extension of the Morfessor Baseline model of unsupervised morphological segmentation (Creutz and Lagus, 2007) that incorporates abstract templates for reduplication, a typologically common but computationally underaddressed process. Through a detailed investigation that applies the model to Maori, the ̄ Indigenous language of Aotearoa New Zealand, we show that incorporating templates improves Morfessor's ability to identify instances of reduplication, and does so most when there are multiple minimally-overlapping templates. We present an error analysis that reveals important factors to consider when applying the extended model and suggests useful future directions.
We investigate implicit vocabulary learning by adults who are exposed to a language in their ambient environment. Most New Zealanders do not speak Māori, yet are exposed to it throughout their lifetime. We show that this exposure leads to a large proto-lexicon – implicit knowledge of the existence of words and sub-word units without any associated meaning. Despite not explicitly knowing many Māori words, non-Māori-speaking New Zealanders are able to access this proto-lexicon to distinguish Māori words from Māori-like nonwords. What's more, they are able to generalize over the proto-lexicon to generate sophisticated phonotactic knowledge, which lets them evaluate the well-formedness of Māori-like nonwords just as well as fluent Māori speakers.
Empirically-observed word frequency effects in regular sound change present a puzzle: how can high-frequency words change faster than low-frequency words in some cases, slower in other cases, and at the same rate in yet other cases? We argue that this puzzle can be answered by giving substantial weight to the role of the listener. We present an exemplar-based computational model of regular sound change in which the listener plays a large role, and we demonstrate that it generates sound changes with properties and word frequency effects seen in corpora. In particular, we consider the experimentally-supported assumption that high-frequency words may be more robustly recognized than low-frequency words in the face of acoustic ambiguity. We show that this assumption allows high-frequency words to change at the same rate as low-frequency words when a phoneme category moves without encroaching on the acoustic space of another, faster than low-frequency words when it moves toward another, and slower than low-frequency words when it moves away from another. We discuss how these predicted word frequency effects apply to different types of sound changes that have been observed in the literature. Importantly, these frequency effects follow from assumptions regarding processes in perception, not production. Frequency-based asymmetries in perception predict different frequency effects for different kinds of sound change.