Most Japanese dialects, including the standard Tokyo variety, employ a lexical tonal system, whereas certain dialects, referred to as accentless dialects, lack such a system. Although dialectological research has frequently suggested that younger speakers in accentless regions have acquired the standard tonal system through dialect standardization, evidence from perception studies indicates that these distinctions may not yet be fully established. The present study investigates whether the acquisition of the standard tonal system among speakers of accentless dialects is complete. We employed both evaluations of accentless dialect speakers' speech production by naive listeners and a sequence recall task designed to assess phonological contrasts in speech perception. Our findings demonstrate that the speech production of accentless dialect speakers closely resembled that of standard Japanese speakers and was often indistinguishable to naïve listeners, consistent with previous claims of complete standardization. However, the perception experiment revealed significantly poorer performance among accentless dialect speakers in discriminating lexical pitch contrasts. These results suggest that the standardization of accentless dialects remains incomplete at the perceptual level, despite the apparent convergence in production. Moreover, the sequence recall task proved to be an effective tool for identifying subtle cases of incomplete acquisition in the standardization of dialects.
The prosodic characteristics of a native language greatly influence early language acquisition. Yet, Japanese mothers are known to use a specific prosodic structure in infant-directed vocabulary (IDV)-specifically, three-mora, two-syllable words with a heavy-light pattern-which, crucially, differs from the standard prosodic rhythm of adult vocabulary. This study used near-infrared spectroscopy to examine hemodynamic responses to the Japanese IDV form in 5-month-old (n = 31) and 9-month-old (n = 34) Japanese infants, targeting the period before and during the emergence of this preference. The results revealed that oxygenated hemoglobin was greater for the IDV form than for the non-IDV form in the left superior temporal gyrus (STG) for both age groups, consistent with the advantage of the IDV form observed in previous behavioral studies. Furthermore, this effect was localized to the left middle and left posterior STG in 5- and 9-month-old infants, respectively, highlighting early sensitivity to its prosodic structure followed by the emergence of phonological representation. This cortical shift, along with an observed trend toward adult-like patterns, may suggest a broader transition from perceptually accessible IDV structures to the more diverse patterns of standard adult vocabulary. Although 5-month-old infants who have not yet exhibited a preference for the IDV form may not have developed specific phonological representations, their brains' ability to process its prosodic pattern could serve as a foundation for subsequent learning. These findings demonstrate that the specific structure of Japanese IDV acts as a foundational scaffold, guiding the transition from initial prosodic tuning to mature word-level processing. SUMMARY: Early sensitivity to prosodic structure was observed in 5-month-old Japanese infants. Emerging phonological representation was observed in 9-month-old Japanese infants. Japanese infant-directed vocabulary form serves as a prosodic template.
Advancing generalizability, replicability, and public trust in developmental science requires testing theories across diverse contexts and disseminating findings widely. Yet researchers based outside the USA or studying non-USA samples (non-USA based researchers) often face obstacles during peer-review in mainstream psychology journals. Moreover, USA-based research on developmental science benefits from a global perspective. To better understand these challenges, we surveyed 229 non-USA based global developmental scientists about their peer-review experiences. We separately assessed how often participants expected and were asked to make changes that devalued the global aspects of their research, such as providing excessive justification for the sample or its generalizability, stratifying the sample, and collecting a new comparison sample. Analyses of free responses revealed a high incidence of comments received during peer review that questioned, minimized, or sought to alter the globally-relevant aspects of the research. Notably, participants felt the need to alter or downplay these aspects more frequently than they were asked to, suggesting anticipated or internalized devaluation of one's research among global developmental scientists. Qualitative responses reinforced these findings and offered recommendations for improving peer-review practices. Overall, the study highlights unique challenges that non-USA based researchers encounter during peer review. Such challenges may discourage the pursuit and publication of global developmental research, limiting overall advancement of replicability and generalizability of developmental science. Addressing these issues could strengthen developmental science by integrating insights from many contexts, ultimately enriching developmental research both within and beyond the USA. SUMMARY: Online survey conducted revealed the challenges encountered during peer-review by global developmental scientists (based outside of USA or studying non-USA samples). Various challenges during peer review that questioned, devalued or sought to alter globally relevant aspects of developmental research were frequently reported. Participants reported feeling the need to alter or downplay relevant aspects of their research more often than being requested to do so. Global researchers may face pervasive peer-review challenges, despite the importance of their work for generalizability and replicability of developmental science within and beyond USA.
It has been proposed that the world's languages fall into three rhythmic classes according to the units that determine their rhythmic structure: stress-based (e.g., English), syllable-based (e.g., French), and mora-based (e.g., Japanese). These rhythmic differences modulate language-specific processing in adults and young infants. It has also been proposed that the syllable is the default rhythmic unit across languages, more prominent in acquisition and processing. This study investigated the development of these language-specific versus language-general effects on the neural processing of speech, focusing on cortical tracking at frequency bands potentially corresponding to three rhythmic units (2.5 Hz/foot, 5 Hz/syllable, 10 Hz/mora). We assessed cortical tracking in 119 infants belonging to one of three native-language groups: English, French, and Japanese, and two age groups: 4 and 8 months. Infants heard CVN (consonant-vowel-nasal) syllable strings synthesized using a native-language voice in their respective native language, a non-native voice (Polish), and a non-speech (vocoded) version of the non-native voice stimuli. In all conditions, the stimuli consisted of 100-ms morae, combining into 200-ms syllables, combining into 400-ms bisyllabic feet. In all three language groups and at both ages, results showed (1) efficient cortical tracking at 5 and 10 Hz, but inconsistent tracking at 2.5 Hz; (2) stronger tracking at 5 Hz, and (3) differential tracking of speech and non-speech stimuli. These findings establish that language-general rhythmic processing guides cortical tracking, that the cross-linguistic syllabic level is dominant until at least 8 months, and that speech stimuli are processed differently than non-speech stimuli by at least 4 months.
Young infants' remarkable ability to discriminate non-native phoneme contrasts played a critical role in shaping the tenets of the perceptual narrowing hypothesis: early on, infants are sensitive to most phoneme categories, including those not used in their native language, but lose this sensitivity as they attune to their language. However, supporting evidence was derived from limited geographical regions and languages, particularly on early sensitivity, requiring further studies to specify the extent of early sensitivity and reassess the dominant developmental pattern. This study aimed to fill this gap by examining discrimination patterns for three-way Thai stop contrasts by two other Asian language learners (Korean and Japanese) at age 4-6 months. The three stop categories in Thai are distinct along the voice onset time (VOT) dimension, encompassing both negative and positive values. Thai pre-voiced and voiceless (i.e., short lag) stops are similar to stop categories used in languages such as French, Dutch, and Spanish. Thai voiceless and voiceless aspirated (i.e., long lag) stops are similar to those in English, Chinese, and German. Therefore, Thai stop categories provide an ideal test continuum for confirming early universal sensitivities to two supposedly language-general VOT boundaries (-30 ms, +30 ms). We presented two Thai phoneme pairs (pre-voiced vs. voiceless, voiceless vs. voiceless aspirated) to Korean and Japanese infants aged 4-6 months and observed their discrimination patterns using a visual habituation paradigm. The results showed divergent discrimination between the two language learners. Korean infants showed sensitivity to the pre-voiced-voiceless pair, whereas Japanese infants did not. By contrast, only Japanese infants showed some sensitivity to the voiceless-voiceless aspirated pair with some directionality effect, whereas Korean infants did not. These results demonstrate systematic cross-linguistic differences reflecting input influence in early perceptual sensitivity and suggest the ambient language environment may influence consonant perception much earlier than has been considered by the perceptual narrowing theory, calling for further refinement of the extent of initial perceptual state in the theory.
Increasing geographical and cultural diversity in research participation has been a key priority for psychological researchers. In this article, we track changes in participant diversity in developmental science over the past decade. These analyses reveal surprisingly modest shifts in global diversity of research participants over time, calling into question the generalizability of our empirical foundation. We provide examples from the study of early child development of the significant epistemic and ethical costs of a lack of geographical and cultural diversity to demonstrate why greater diversification is essential to a generalizable science of human development. We also discuss strategies for diversification that could be implemented throughout the research ecosystem in the service of a culturally anchored, generalizable, and replicable science.
Perceptual narrowing of speech perception supposes that young infants can discriminate most speech sounds early in life. During the second half of the first year, infants' phonetic sensitivity is attuned to their native phonology. However, supporting evidence for this pattern comes primarily from learners from a limited number of regions and languages. Very little evidence has accumulated on infants learning languages spoken in Asia, which accounts for most of the world's population. The present study examined the developmental trajectory of Korean-learning infants' sensitivity to a native stop contrast during the first year of life. The Korean language utilizes unusual voiceless three-way stop categories, requiring target categories to be derived from tight phonetic space. Further, two of these categories-lenis and aspirated-have undergone a diachronic change in recent decades as the primary acoustic cue for distinction has shifted among modern speakers. Consequently, the input distributions of these categories are mixed across speakers and speech styles, requiring learners to build flexible representations of target categories along these variations. The results showed that among the three age groups-4-6 months, 7-9 months, and 10-12 months-we tested, only 10-12-month-olds showed weak sensitivity to the two categories, suggesting that robust discrimination is not in place by the end of the first year. The study adds scarcely represented data, lending additional support for the lack of early sensitivity and prolonged emergence of native phonology that are inconsistent with learners of predominant studies and calls for more diverse samples to verify the generality of the typical perceptual narrowing pattern. RESEARCH HIGHLIGHTSWe investigated Korean-learning infants' developmental trajectory of native phoneme categories and whether they show the typical perceptual narrowing pattern.Robust discrimination did not appear until 12 months, suggesting that Korean infants' native phonology is not stabilized by the end of the first year.The prolonged emergence of sensitivity could be due to restricted phonetic space and input variations but suggests the possibility of a different developmental trajectory.The current study contributes scarcely represented Korean-learning infants' phonetic discrimination data to the speech development field.
Chapter 13 Developmental changes in the interpretation of an ambiguous structure and an ambiguous prosodic cue in Japanese was published in Volume 2 Interaction Between Linguistic and Nonlinguistic Factors on page 255.
Recent research shows that adults' neural oscillations track the rhythm of the speech signal. However, the extent to which this tracking is driven by the acoustics of the signal, or by language-specific processing remains unknown. Here adult native listeners of three rhythmically different languages (English, French, Japanese) were compared on their cortical tracking of speech envelopes synthesized in their three native languages, which allowed for coding at each of the three language's dominant rhythmic unit, respectively the foot (2.5 Hz), syllable (5 Hz), or mora (10 Hz) level. The three language groups were also tested with a sequence in a non-native language, Polish, and a non-speech vocoded equivalent, to investigate possible differential speech/nonspeech processing. The results first showed that cortical tracking was most prominent at 5 Hz (syllable rate) for all three groups, but the French listeners showed enhanced tracking at 5 Hz compared to the English and the Japanese groups. Second, across groups, there were no differences in responses for speech versus non-speech at 5 Hz (syllable rate), but there was better tracking for speech than for non-speech at 10 Hz (not the syllable rate). Together these results provide evidence for both language-general and language-specific influences on cortical tracking.
The development of allophonic variants of phonemes is poorly understood. Thus, this study aimed to examine when children typically begin to articulate a phoneme with the same allophonic variant typically used by adults. Japanese children aged 5-13 years and adults aged 18-24 years participated in an elicited production task. We analyzed developmental changes in allophonic variation of the phoneme /z/, which is realized variably either as an affricate or a fricative. The results revealed that children aged nine years or younger realized /z/ as affricate significantly more than 13-year-old and adult speakers. Once the children reached 11 years of age, the difference compared to adults was not statistically significant, which denotes a similar developmental pattern as that of speech motor control (e.g., lip and jaw) and cognitive-linguistic skill. Moreover, we examined whether the developmental changes of allophonic realization of /z/ are due to speech rate and the time to articulate /z/. The results showed that the allophonic realization of /z/ is not affected by these factors, which is not the case in adults. We also found that the effects of speech rate and the time to articulate /z/ on the allophonic realization become adult-like at around 11 years of age.
Development of speech rate is often reported as children exhibiting reduced speech rates until they reach adolescence. Previous studies have investigated the developmental process of speech rate using global measures (syllables per second, syllables per minute, or words per minute) and revealed that development continues up to around 13 years of age in several languages. However, the global measures fail to capture language-specific characteristics of phonological/prosodic structure within a word. The current study attempted to examine the developmental process of speech rate and language-specific rhythm in an elicited production task. We recorded the speech of Japanese-speaking monolingual participants (18 participants each in child [5-, 7-, 9-, 11-, and 13-year old] and adult groups), who pronounced three types of target words: two-mora, two-syllable words (CV.CV); three-mora, two syllable words (CVV.CV); and three-mora, three-syllable words (CV.CV.CV), where C is consonant and V is vowel. We analyzed total word duration and differences in two pairs of word types: a pair of three-mora words (to show the effect of syllables) and a pair of two-syllable words (to show the effect of moras). The results revealed that Japanese-speaking children have acquired adult like word duration before 11 years of age, whereas the development of rhythmical timing control continues until approximately 13 years of age. The results also suggest that the effect of syllables for Japanese-speaking children aged 9 years or under was stronger than that of moras, whereas the effect of moras was stronger after 9 years of age, indicating that the default unit for children in speech rhythm may be the syllable even when the language is mora-based.(c) 2022 Elsevier Inc. All rights reserved.
Over the past 50 years, scientists have made amazing discoveries about the origins of human language acquisition. Central to this field of study is the process by which infants' perceptual sensitivities gradually align with native language structure, known as perceptual narrowing . Perceptual narrowing offers a theoretical account of how infants draw on environmental experience to induce underlying linguistic structure, providing an important pathway to word learning. Researchers have advanced perceptual narrowing theory as a universal developmental theory that applies broadly across language learners. In this article, we examine diversity and representation of empirical evidence for perceptual narrowing of speech in infancy. As demonstrated, cumulative evidence draws from limited types of learners, languages, and locations, so current accounts of perceptual narrowing must be viewed in terms of sampling patterns. We suggest actions to diversify and broaden empirical investigations of perceptual narrowing to address core issues of validity, replicability, and generalizability.
Human infants acquire motor patterns for speech during the first several years of their lives. Sequential vocalizations such as human speech are complex behaviors, and the ability to learn new vocalizations is limited to only a few animal species. Vocalizations are generated through the coordination of three types of organs: namely, vocal, respiratory, and articulatory organs. Moreover, sophisticated temporal respiratory control might be necessary for sequential vocalization involving human speech. However, it remains unknown how coordination develops in human infants and if this developmental process is shared with other vocal learners. To answer these questions, we analyzed temporal parameters of sequential vocalizations during the first year in human infants and compared these developmental changes to song development in the Bengalese finch, another vocal learner. In human infants, early cry was also analyzed as an innate sequential vocalization. The following three temporal parameters of sequential vocalizations were measured: note duration (ND), inter-onset interval, and inter-note interval (INI). The results showed that both human infants and Bengalese finches had longer INIs than ND in the early phase. Gradually, the INI and ND converged to a similar range throughout development. While ND increased until 6 months of age in infants, the INI decreased up to 60 days posthatching in finches. Regarding infant cry, ND and INI were within similar ranges, but the INI was more stable in length than ND. In sequential vocalizations, temporal parameters developed early with subsequent articulatory stabilization in both vocal learners. However, this developmental change was accomplished in a species-specific manner. These findings could provide important insights into our understanding of the evolution of vocal learning.
Songs and speech play central roles in early caretaker-infant communicative interactions, which are crucial for infants' cognitive, social, and emotional development. Compared to speech development, however, much less is known about how infants process songs or how songs affect their development. Lyrics and melody are two key components of songs, and much of the research on song processing has examined how the two components of the songs are processed. The current study focused on the roles of lyrics and melody in song perception, by examining developmental patterns and the ways in which lyrics and melody are processed in the infants' brains using near-infrared spectroscopy (NIRS). The results revealed that developmental changes occur in infants' processing of lyrics and melody in a similar timeline as perceptual reorganization, that is, from 4.5 and 12 months of age. We found that 4.5-month-olds showed a right hemispheric advantage in the processing of songs that underwent a change in either lyrics or melodies. Conversely, 12-month-olds showed significantly higher activation bilaterally when lyrics and melody changed at the same time. These results suggest that 4.5-month-olds processed songs in the same manner as music without lyrics. Moreover, 12-month-olds processed lyrics and melody in an interactive manner, a sign of a more mature processing method. These findings highlight the importance of investigating the independent development of music and language, and also considering the relationship between speech and song, lyrics and melody in song, and speech and music more broadly.
A prominent hypothesis holds that by speaking to infants in infant-directed speech (IDS) as opposed to adult-directed speech (ADS), parents help them learn phonetic categories. Specifically, two characteristics of IDS have been claimed to facilitate learning: hyperarticulation, which makes the categories more separable, and variability, which makes the generalization more robust. Here, we test the separability and robustness of vowel category learning on acoustic representations of speech uttered by Japanese adults in ADS, IDS (addressed to 18- to 24-month olds), or read speech (RS). Separability is determined by means of a distance measure computed between the five short vowel categories of Japanese, while robustness is assessed by testing the ability of six different machine learning algorithms trained to classify vowels to generalize on stimuli spoken by a novel speaker in ADS. Using two different speech representations, we find that hyperarticulated speech, in the case of RS, can yield better separability, and that increased between-speaker variability in ADS can yield, for some algorithms, more robust categories. However, these conclusions do not apply to IDS, which turned out to yield neither more separable nor more robust categories compared to ADS inputs. We discuss the usefulness of machine learning algorithms run on real data to test hypotheses about the functional role of IDS.
Infants come to learn several hundreds of word forms by two years of age, and it is possible this involves carving these forms out from continuous speech. It has been proposed that the task is facilitated by the presence of prosodic boundaries. We revisit this claim by running computational models of word segmentation, with and without prosodic information, on a corpus of infant-directed speech. We use five cognitively-based algorithms, which vary in whether they employ a sub-lexical or a lexical segmentation strategy and whether they are simple heuristics or embody an ideal learner. Results show that providing expert-annotated prosodic breaks does not uniformly help all segmentation models. The sub-lexical algorithms, which perform more poorly, benefit most, while the lexical ones show a very small gain. Moreover, when prosodic information is derived automatically from the acoustic cues infants are known to be sensitive to, errors in the detection of the boundaries lead to smaller positive effects, and even negative ones for some algorithms. This shows that even though infants could potentially use prosodic breaks, it does not necessarily follow that they should incorporate prosody into their segmentation strategies, when confronted with realistic signals.
This chapter covers theoretical frameworks and experimental findings showing that young infants are already sensitive to language prosody prenatally and can use it to learn about the lexical and morphosyntactic features of their native language(s). Specifically, the chapter first summarizes how prosody relates to the lexicon and the grammar in different languages. It then reviews empirical evidence about prosodic perception in infants. Subsequently, it shows how this early sensitivity to prosody facilitates language learning. Three areas are discussed. First, evidence is presented showing that infants can use their knowledge of lexical stress to constrain word learning. Second, the chapter argues that infants use prosody to learn about basic word order. Third, the chapter shows that infants can use prosody to constrain syntactic analysis and thus word learning. The chapter concludes by discussing the theoretical implications of the reviewed findings, and by highlighting open questions.
Is infants’ word learning boosted by nonhuman social agents? An on-screen virtual agent taught infants word–object associations in a setup where the presence of contingent and referential cues could be manipulated using gaze contingency. In the study, 12-month-old Japanese-learning children (N = 36) looked significantly more to the correct object when it was labeled after exposure to a contingent and referential display versus a noncontingent and nonreferential display. These results show that communicative cues can augment learning even for a nonhuman agent, a finding highly relevant for our understanding of the mechanisms through which the social environment supports language acquisition and for research on the use of interactive screen media.
Infants learn about the sounds of their language and adults process the sounds they hear, even though sound categories often overlap in their acoustics. Researchers have suggested that listeners rely on context for these tasks, and have proposed two main ways that context could be helpful: top-down information accounts, which argue that listeners use context to predict which sound will be produced, and normalization accounts, which argue that listeners compensate for the fact that the same sound is produced differently in different contexts by factoring out this systematic context-dependent variability from the acoustics. These ideas have been somewhat conflated in past research, and have rarely been tested on naturalistic speech. We implement top-down and normalization accounts separately and evaluate their relative efficacy on spontaneous speech, using the test case of Japanese vowels. We find that top-down information strategies are effective even on spontaneous speech. Surprisingly, we find that at least one common implementation of normalization is ineffective on spontaneous speech, in contrast to what has been found on lab speech. We provide analyses showing that when there are systematic regularities in which contexts different sounds occur in-which are common in naturalistic speech, but generally controlled for in lab speech-normalization can actually increase category overlap rather than decrease it. This work calls into question the usefulness of normalization in naturalistic listening tasks, and highlights the importance of applying ideas from carefully controlled lab speech to naturalistic, spontaneous speech.