Existing studies have emphasized the important effect of native phonotactic experience and cross-language perceptual similarity on non-native speech perception, but have seldom considered this issue when the targeted non-native segments are undergoing a phonological merger. To supplement this discussion, this study investigated Mandarin Chinese (MC) speakers' perception of syllable-final segments (⁄-р⁄, ⁄-t⁄, ⁄-k⁄, ⁄-m⁄, ⁄-n⁄, ⁄-ŋ⁄, ⁄-Ø⁄) that are undergoing a coronal-dorsal merger (⁄-t⁄↔⁄-k⁄, ⁄-n⁄↔/-ŋ/) in Hong Kong Cantonese (HKC). The HKC speakers completed an Identification Test, which confirmed the documented merger of Cantonese syllable-final segments in previous research. A Perceptual Assimilation Test and an Identification Test were prepared for MC speakers to assess their perceptual mapping and perception performance of HKC syllable-final segments. As observed, the MC speakers failed to map the merging HKC ⁄-n⁄-⁄-ŋ⁄ onto the contrastive MC ⁄-n⁄-⁄-ŋ⁄. Moreover, they perceptually collapsed the merging HKC ⁄-t⁄ and ⁄-k⁄ into one category via an illicit alteration of the place of articulation. However, neither native Mandarin phonotactic experience nor cross-MC-HKC perceptual similarity can sufficiently explain the MC speakers' perceptual mapping and rendering of the merging HKC targets. Accordingly, the current study proposes that the non-native phonological merger has a supplemental mediating effect on non-native speech perception.
Children with autism spectrum disorder are known to exhibit both social and language difficulties. Speech prosody is known to be easily noticeable, which has been shown to have far-reaching influences in the academic and social life of autistic individuals. This study examined two training programs on the speech prosody of autistic children, who tend to avoid social speech signals. The first program is a lab perceptual training program without social interaction, while the second utilizes a social robot to provide training with controlled, simulated social interaction. Ninety-two children in total were recruited with sixty-nine participants formally diagnosed with ASD and twenty-three children were typically developing children without any language or speech disorder. Our results showed that both lab perceptual training and robot-assisted training with simulated social interactions led to improvement in the use of speech prosody by autistic children. Although social interaction is considered critical in language acquisition for typical population, autistic individuals tend not to prefer social speech signals, which is hypothesized to lead to their social and language deficits. This study hence proposes two successful alternative ways to facilitate their learning of language through lab perceptual training and simulated human-robot interaction.
Prompt compression shortens LLM input to reduce inference cost, yet existing methods score token importance through LM forward passes. It remains questionable whether such nuanced, costly token selection is necessary. Compression requires identifying informative content, a problem that linguistic research has long addressed through cues that can be operationalized as deterministic rules. We therefore ask: can linguistic rules alone serve as effective prompt compressors, without LM-based scoring at compression time? To address this, we conduct offline evolutionary search over lexical, syntactic, semantic, and discourse seeds to find competitive rule combinations. The resulting linguistic compressor requires no LM forward pass at deployment and uses only CPU-side processing for compression. We evaluate it with a dual-path protocol to balance compression quality and reconstruction fidelity. Across short passages, multi-document reasoning, and dialogue-memory QA datasets, evolved compressors achieve performance similar to that of recent advanced prompt-compression strategies. Performance is strongest under light-to-moderate compression and degrades as compression becomes more aggressive, while the Direct and Reconstruction paths exhibit distinct patterns. Evolutionary analysis reveals that effective compression fuses signals across linguistic levels and, as the compression ratio increases, rules shift from token pruning to sentence extraction.
Quality of life encompasses social, economic, and environmental dimensions, with environmental indicators posing challenges due to their diverse nature. This study proposes a flexible measure of the physical environment that considers individual preferences. The research focused on neighbourhoods in Hong Kong, representing various levels of urbanization and spatial distribution, with a fixed neighbourhood size to enable meaningful comparisons. Utilizing Geographic Information System and Remote Sensing techniques, the study analysed four domains of urban morphology at the neighbourhood level: education-health-recreation facilities, street patterns, land use diversity, and building density. Principal Component Analysis was employed to reduce the dimensionality of each domain and derive an Environmental Quality Sub-index (EQ-I). This index can be standardized and personalized based on individuals’ values and preferences. The EQ-I facilitates both quantitative and qualitative comparisons through visual representations. The study exemplifies a methodological approach to consolidating multiple variables into a single index, considering the differential weighting of preferences. The methodology employs direct and objective measures that can be adapted and replicated in other cities, providing standardized yet personalized scores for regional and international comparisons. The findings present valuable insights for policymakers and urban planners aiming to enhance the environmental quality of ultra-dense urban settings.
Generative AI tools, particularly those utilizing large language models (LLMs), are increasingly used in everyday contexts. While these tools enhance productivity and accessibility, little is known about how Deaf and Hard of Hearing (DHH) individuals engage with them or the challenges they face when using them. This paper presents a mixed-method study exploring how the DHH community uses Text AI tools like ChatGPT to reduce communication barriers and enhance information access. We surveyed 80 DHH participants and conducted interviews with 9 participants. Our findings reveal important benefits, such as eased communication and bridging Deaf and hearing cultures, alongside challenges like lack of American Sign Language (ASL) support and Deaf cultural understanding. We highlight unique usage patterns, propose inclusive design recommendations, and outline future research directions to improve Text AI accessibility for the DHH community.
Introduction: Music and speech prosody share notable parallels, and music-based interventions have shown promise in fostering language development and social responsiveness. Song-based training, leveraging acoustic similarities between song and speech, is especially effective. This study examined whether short-term song-based training could enhance prosodic focus-marking in nondominant languages for autistic children. Specifically, it explored improvements in focus-marking strategies, such as on-focus expansion (OFE) and post-focus compression (PFC), and the number of prosodic correlates used. Method: A short-term sung speech training intervention was designed, aligning melodic patterns with Mandarin's prosodic focus marking. Eighteen native Cantonese-speaking children with autism spectrum disorder underwent short-term sung speech training, and their pre- and posttraining performance was compared with two control groups: 18 Cantonese-speaking and 20 Mandarin-speaking typically developing children. Comparisons were made across participant groups as well as within the autistic group before and after the training. Results: Sung speech training improved OFE use, particularly in fundamental frequency range, for noncontrastive focus marking in autistic children. Effects on PFC were less evident, and the training primarily enhanced OFE rather than increasing the number of prosodic correlates used. Control Cantonese-speaking participants showed no comparable improvements. Conclusion: These findings highlight the potential of short-term, perception-based sung speech training as a supplementary intervention for improving prosodic focus marking in trilingual autistic children's nondominant languages, indicating positive cross-domain effects on speech-processing abilities. Supplemental Material: https://doi.org/10.23641/asha.30347731
Studies indicate that women generally express emotions more visibly than men. It is unclear whether such variance is prevalent across different languages. This work explores gender-related differences in acoustic features of emotional prosody across Mandarin, Cantonese, and English, examining both simple and complex emotions. Six speakers produced 40 neutral sentences in 12 emotional states in each language. Results revealed systematic gender and language interactions. Females exhibited higher baseline pitch and greater modulation for high-arousal emotions (e.g., happy, angry), whereas males showed expanded pitch ranges in English but not in Mandarin or Cantonese. English speakers used wider pitch spans for surprise, whereas Mandarin/Cantonese relied on tone-driven contours. Cantonese speakers showed tighter face-voice coupling for complex emotions, suggesting cultural specificity in multimodal expression. The findings highlight how biological tendencies (e.g., female expressivity) interact with linguistic structure (e.g., tonal constraints), suggesting sociophonetic variance constrained by both gender and language. Tonal languages limit pitch flexibility, leading to compensatory strategies like duration adjustments, whereas English allows freer pitch modulation. These observations underscore the need for language-specific models in affective computing. Future research could expand speaker samples and include perceptual validation to refine emotion recognition systems and avoid gender biases in AI applications.
Purpose: This study investigated the effect of musical training on phonetic accommodation in a second language (L2) after interacting with a social robot, exploring the motivations and reasons behind their accommodation strategies. Method: Fifteen L2 English speakers with long-term musical training experience (musician group) and 15 speakers without musical training experience (nonmusician group) were recruited to complete four conversational tasks with the social robot Furhat. Their production of a list of key words and carrier sentences was collected before and after conversations and used to quantify their phonetic accommodations. The spectral cues and prosodic cues of the production were extracted and analyzed. Results: Both groups showed similar convergence patterns but different divergence patterns. Specifically, the musician group showed divergence from the robot's production on more prosodic cues (mean fundamental frequency and duration) than the nonmusician group. Both groups converged their vowel formants toward the robot without group differences. Conclusions: The findings reflect individuals' assessment of the robot's speech characteristics and their efforts to enhance communication efficiency, which might indicate a special speech register used for addressing the robot. The finding is more noticeable in the musician group compared to the nonmusician group. We proposed two possible explanations of the effect of musical training on phonetic accommodations: one involves the training of auditory attention and working memory and the other relates to the refinement of phonetic talent in L2 acquisition, contributing to theories on the relationship between music and language. This study also has implications for applying musical training to speech communication training in clinical populations and for designing social robots to better serve as speech therapy partners.
Prosodic cues provide suprasegmental access to speech-act perception, yet they remain underexplored in tonal languages. This study analysed the production of six speech acts (Statement, Doubt, Suggestion, Command, Celebration, and Complaint) by twenty Cantonese-speaking adults. Mean F₀, F₀ range, intensity, and duration jointly achieved classification rates well above chance level, despite competition from lexical tone. Doubt and Suggestion exhibited the highest mean F₀ and widest pitch range, while Statement and Complaint occupied the lowest pitch region. Command showed the narrowest pitch range and greatest loudness, whereas Complaint displayed the longest duration. The findings confirm that global intonation, intensity, and timing reliably encode pragmatic force in Cantonese.
Purpose: Children with autism spectrum disorder (ASD) often show abnormal speech prosody. Tonal languages can pose more difficulties as speakers need to use acoustic cues to make lexical contrasts while encoding the focal function, but the acquisition of speech prosody of non-native languages, especially tonal languages has rarely been investigated. Methods: This study aims to fill in the aforementioned gap by studying prosodic focus-marking in Mandarin by native Cantonese-speaking children with ASD (n = 25), in comparison with their typically developing (TD) peers (n = 20) and native Mandarin-speaking children (n = 20). Natural prosodic marking of different types of focus was elicited by picture-based prompt questions, recorded and analyzed acoustically. Results: The autistic children made use of fewer acoustic cues and produced less evident on-focus expansion in these cues than TD, especially the native-Mandarin speaking peers. They also demonstrated a clear preference to on-focus expansion than to post-focus compression. These children, together with their native Cantonese-speaking peers, also hyper-performed in tone realization, prioritizing lexical prosody over focus marking. Such hyper-performance may further limit their use of prosodic cues in focus marking. However, the difficulties the autistic children faced in the acquisition of speech prosody in a non-native tone language, though found, are not more than those they face in their mother tongue. Conclusion: Multilingual exposure may help the autistic children master the use of some focus marking strategies though they still need interventions to help them to implement their focus-marking knowledge more sufficiently in both native and non-native languages.
Abnormal speech prosody has been widely reported in individuals with autism. Many studies on children and adults with autism spectrum disorder speaking a non-tonal language showed deficits in using prosodic cues to mark focus. However, focus marking by autistic children speaking a tonal language is rarely examined. Cantonese-speaking children may face additional difficulties because tonal languages require them to use prosodic cues to achieve multiple functions simultaneously such as lexical contrasting and focus marking. This study bridges this research gap by acoustically evaluating the use of Cantonese speech prosody to mark information structure by Cantonese-speaking children with and without autism spectrum disorder. We designed speech production tasks to elicit natural broad and narrow focus production among these children in sentences with different tone combinations. Acoustic correlates of prosodic focus marking like f0, duration and intensity of each syllable were analyzed to examine the effect of participant group, focus condition and lexical tones. Our results showed differences in focus marking patterns between Cantonese-speaking children with and without autism spectrum disorder. The autistic children not only showed insufficient on-focus expansion in terms of f0 range and duration when marking focus, but also produced less distinctive tone shapes in general. There was no evidence that the prosodic complexity (i.e. sentences with single tones or combinations of tones) significantly affected focus marking in these autistic children and their typically-developing (TD) peers.
PURPOSE:The current study investigated English prosodic focus marking by autistic and typically developing (TD) Cantonese trilingual children, and examined the potential differences in this regard compared to native English-speaking children.METHOD:Forty-eight participants were recruited with 16 speakers for each of the three groups (Cantonese-speaking autistic [CASD], Cantonese-speaking TD [CTD], and English-speaking TD [ETD] children), and prompt questions were designed to elicit desired focus type (i.e., broad, narrow, and contrastive focus). Mean duration, mean fundamental frequency (F0), F0 range, mean intensity, and F0 curves were used as the acoustic correlates for linear mixed-effects model fitting and functional data analyses in relation to groups and focus conditions (i.e., broad, narrow, and contrastive pre-, on-, and post-focus).RESULTS:The CTD group had post-focus compression (PFC) patterns via reducing mean duration, narrowing F0 range, and lowering mean F0, F0 curve, and mean intensity for words under both narrow and contrastive post-focus conditions, while the CASD group only had shortened mean duration and lowered F0 curves. However, neither the CTD group nor CASD group showed much of on-focus expansion (OFE) patterns. The ETD group marked OFE by increasing mean duration, mean F0, mean intensity, and higher F0 curve for words under on-focus conditions.CONCLUSIONS:The CTD group utilized more acoustic cues than the CASD group when it comes to PFC. The ETD group differed from the CASD and CTD groups in the use of OFE. Furthermore, both the CASD and CTD groups showed positive first language transfer in the use of duration and intensity and, potentially, successful acquisition in the use of F0 for prosodic focus marking. Meanwhile, the differences in the use of OFE between the Cantonese-speaking and English-speaking groups, not PFC, might indicate that Cantonese-speaking children acquire PFC prior to OFE.
Purpose: Nonword repetition (NWR) has been described as a clinical marker of developmental language disorder (DLD), as NWR tasks consistently discriminate between DLD and typical development (TD) cross-linguistically, with Cantonese as the only reported exception. This study reexamines whether NWR is able to generate TD/DLD group differences in Cantonese-speaking children by reporting on a novel set of NWR stimuli that take into account factors known to affect NWR performance and group differentiation, including lexicality, sublexicality, length, and syllable complexity. Method: Sixteen Cantonese-speaking children with DLD and 16 age-matched children with TD repeated two sets of high-lexicality nonwords, where all constituent syllables are morphemic in Cantonese but meaningless when combined, and one set of low-lexicality nonwords, where all constituent syllables are nonmorphemic. Low-lexicality nonwords were further classified on sublexicality in terms of consonant–vowel (CV) combination attestedness (whether or not CV combinations in nonword syllables occur in real Cantonese words). Results: Children with DLD scored significantly below their peers with TD. Effect sizes showed that high-lexicality nonwords and nonword syllables with attested CV combinations offered the greatest TD/DLD group differentiation. Nonword length and syllable complexity did not affect TD/DLD group differentiation. Conclusions: NWR can capture TD/DLD group differences in Cantonese-speaking children. Lexicality and sublexicality effects must be considered in designing NWR stimuli for TD/DLD group differentiation. Future studies should replicate the present study on a larger sample size and a younger population as well as examine the diagnostic accuracy of this NWR test. Supplemental Material: https://doi.org/10.23641/asha.25529371
Purpose: Literature on apraxia of speech (AOS) in Chinese speakers is sparse compared to the English literature. This study aims to examine the pitch variation skills of Cantonese adults with AOS poststroke in terms of perceptual tone accuracy, acoustic fundamental frequency ( f o ) changes, and repetition durations on items with different syllable structures, lexical status, and tone syllables in various positions in a sequencing context. Method: Six Cantonese adults with AOS poststroke (AOS group), six adults without AOS poststroke (nAOS group), and six healthy controls (HC group) performed the tone sequencing task (TST), which was adapted from oral diadochokinetic tasks, with three different tone syllables. Tone accuracy, f o values across 10 time points, and acoustic repetition durations were compared within and between the groups. Results: The AOS group produced significantly lower tone accuracy and different f o changes on the three Cantonese tone syllables compared with the control groups and significantly longer repetition durations than the HC group. The AOS group showed more difficulty with the tone syllables with the consonant–vowel structure, while a priming effect was observed on the T2 (high-rising) syllables with lexical meanings. A unique lowering of f o in the final syllable of the trisyllabic items was observed only in the AOS group. Conclusions: The AOS group showed degraded pitch variation skills. The effects of the three linguistic elements were discussed. Future investigations are called for to adapt the TST in other tonal languages to determine if degraded pitch variation skills are present in other tonal language speakers with AOS.
Cross-linguistically, nonword repetition (NWR) tasks have been found to differentiate between typically developing (TD) children and those with Developmental Language Disorder (DLD), even when second-language TD (L2-TD) children are considered. This study examined such group differences in Cantonese. Fifty-seven age-matched children (19 monolingual DLD (MonDLD); 19 monolingual TD (MonTD); and 19 L2-TD) repeated language-specific nonwords with varying lexicality levels and Cantonese-adapted quasi-universal nonwords. At whole-nonword level scoring, on the language-specific, High-Lexicality nonwords, MonDLD scored significantly below MonTD and L2-TD groups which did not differ significantly from each other. At syllable-level scoring, the same pattern of group differentiation was found on quasi-universal nonwords. These findings provide evidence from a typologically distinct and understudied language that NWR tasks can capture significant TD/DLD group differences, even for L2-Cantonese TD children with reduced language experience. Future studies should compare the performance of an L2-DLD group and evaluate the sensitivity and specificity of Cantonese NWR.
The current study investigated the production of English prosody (i.e., focus marking) of trilingual Cantonese children with autism spectrum disorder (ASD) and their typically developing (TD) peers (i.e., Cantonese and American English children without ASD) using declarative questions. Speech materials were segmented at word and syllable levels, and word duration, f0, f0 range and intensity were extracted. Acoustic data were fitted using linear mixed-effects models with different explanatory variables followed by a likelihood ratio test. Between group comparison showed that the ASD group had significantly more fluctuating f0 range in post-focus words than the TD groups, which is likely to be an indication of hypercorrection, i.e., over-application of perceived prosodic pattern in English declarative questions. Within groups, the Cantonese children showed different patterns to the English children in terms of the interaction between the acoustic measures and on-focus expansion and post-focus compression. The Cantonese ASD group showed some degree of post-focus compression in terms of duration and mean f0, while the Cantonese TD group only had such pattern in terms of mean f0. The English TD group had a tendency of on-focus expansion in terms of duration and f0 range, but post-focus words showed significantly higher mean f0 than the pre- and on-focus ones, probably due to the question intonation.
IntroductionSpeech communication is multi-sensory in nature. Seeing a speaker’s head and face movements may significantly influence the listeners’ speech processing, especially when the auditory information is not clear enough. However, research on the visual-auditory integration speech processing has left prosodic perception less well investigated than segmental perception. Furthermore, while native Japanese speakers tend to use less visual cues in segmental perception than in other western languages, to what extent the visual cues are used in Japanese focus perception by the native and non-native listeners remains unknown. To fill in these gaps, we test focus perception in Japanese among native Japanese speakers and Cantonese speakers who learn Japanese, using auditory-only and auditory-visual sentences as stimuli.MethodologyThirty native Tokyo Japanese speakers and thirty Cantonese-speaking Japanese learners who had passed the Japanese-Language Proficiency Test with level N2 or N3 were asked to judge the naturalness of 28 question-answer pairs made up of broad focus eliciting questions and three-word answers carrying broad focus, or contrastive or non-contrastive narrow focus on the middle object words. Question-answer pairs were presented in two sensory modalities, auditory-only and visual-auditory modalities in two separate experimental sessions.ResultsBoth the Japanese and Cantonese groups showed weak integration of visual cues in the judgement of naturalness. Visual-auditory modality only significantly influenced Japanese participants’ perception when the questions and answers were mismatched, but when the answers carried non-contrastive narrow focus, the visual cues impeded rather than facilitated their judgement. Also, the influences of specific visual cues like the displacement of eyebrows or head movements of both Japanese and Cantonese participants’ responses were only significant when the questions and answers were mismatched. While Japanese participants consistently relied on the left eyebrow for focus perception, the Cantonese participants referred to head movements more often.DiscussionThe lack of visual-auditory integration in Japanese speaking population found in segmental perception also exist in prosodic perception of focus. Not much foreign language effects has been found among the Cantonese-speaking learners either, suggesting a limited use of facial expressions in focus marking by native and non-native Japanese speakers. Overall, the present findings indicate that the integration of visual cues in perception of focus may be specific to languages rather than universal, adding to our understanding of multisensory speech perception.
Many studies showed that prosodic cues such as f0, duration and intensity are used in focus marking cross-linguistically. Usually, on-focus words exhibit expansions of acoustic cues such as f0 expansion, whereas post-focus words may show compression of acoustic cues. However, how features in a sub-syllabic level are employed in focus marking remain to be investigated. F0 perturbation refers to the phenomenon that vocal folds vibration is affected by the preceding non-sonorant consonant. The current study aims to examine how f0 perturbation is realized in focus marking in two languages Japanese and Korean. Tokyo Japanese is a pitch-accent language and Seoul Korean is considered to be at the stage of quasi-tonogenesis. Our results showed that f0 perturbation effects were enhanced in on-focus positions and compressed in pre- and post-focus positions for both narrow and contrastive focus in both languages. In addition, our results showed that pitch accent can also affect the realization of f0 perturbation in various focus conditions. Compared to Korean, our results in Japanese showed that f0 perturbation effects were less restricted. These results provide new insights into the current model of communicative functions that sub-syllabic level acoustic cues such as f0 perturbation can also be employed in focus marking.