
This study examined the association and agreement between microstructure calculations from four large language model (LLM) chatbots and those generated by Systematic Analysis of Language Transcripts (SALT) software. Eighty-four typically developing children (35 girls; ages 2–9) participated, with 30 Japanese-English bilinguals and 54 English monolinguals. Narrative language samples were collected and coded with SALT conventions. Transcripts were randomly ordered and analyzed using two prompts. Using a general prompt, transcripts were entered into each LLM chatbot (ChatGPT, Copilot, MagicSchool, and Gemini), microstructure outputs were recorded, then the process was repeated with the same order using the more specific prompt. Results showed that across LLM chatbots, measures requiring relatively simple word or morpheme counts (MLUw, MLUm, NTW, NDW) generally exhibited stronger associations with SALT than the subordination index (SI), which showed greater variability. While intraclass correlation coefficients within measures were more variable, association and agreement coefficients were generally consistent with each other. Children's age and language backgrounds were associated with differences in the magnitude of some LLM–SALT coefficients. Finally, prompt specificity did not strengthen microstructure calculations for any of the LLM chatbots. Findings provide preliminary methodological evidence that some SALT-calculated microstructure measures may be more amenable to LLM chatbot-assisted calculation than others, with substantial associations and generally strong agreement observed for several measures; however, due to the exploratory nature of the current study, their clinical utility, decision-level accuracy, and workflow implications require further investigation.
BackgroundSwearing in digital communication is conventionally viewed as a straightforward marker of offensiveness, but its pragmatic value is systematically reshaped by platform norms, interactional goals, and sociocultural inference, challenging the adequacy of binary offensive-language models.MethodsThis study adopts a mixed-method corpus-pragmatic design to examine Indonesian swearing in Google Play comments from League of Legends: Wild Rift, Netflix, and X/Twitter, using lexicon-based detection across 500,599 non-empty comments, keyword-in-context interpretation, and contextual coding of six pragmatic categories: cognitive devaluation, animalized reference, corporeal taboo, death/dirt/disgust, moral abuse, and expletive intensification.ResultsIndonesian digital swearing performs catharsis, complaint, humor, solidarity, evaluative intensification, and ideological stance-taking, showing it is not reducible to aggression; platform differences are systematic, with gaming reviews foregrounding competence-based blame and cathartic release, streaming reviews concentrating service dissatisfaction and complaint intensification, and social-media reviews mobilizing public stance and identity positioning, while a Cohen's kappa of 0.3858 between lexicon-based and contextual detection confirms that lexical detection alone cannot explain pragmatic force.ConclusionThe study contributes a culturally situated, data-grounded model for interpreting Indonesian digital swearing beyond binary offensive-language classification and offers practical implications for platform-sensitive annotation, sentiment analysis, and content moderation.
IntroductionEasy German (Leichte Sprache) is a rule-guided variety of German intended to increase the accessibility of written information for people with cognitive disabilities and adults with low literacy skills. Although existing guidelines often discourage the use of idioms, experimental evidence on idiom processing in these target groups remains scarce. We investigated how literacy, functional familiarity, transparency, and literal plausibility jointly shape idiom comprehension and preferences in adult low-literacy readers.MethodsIn a norming study with literate German speakers (Experiment 1; Prolific; 71 verb-phrase idioms), we obtained seven-point ratings of transparency and literal plausibility (reverse-coded so that higher values indicate a more plausible literal scene) and selected 24 idioms to instantiate a 2 × 2 design (high/low transparency × high/low plausibility). Literacy in the target group was assessed with otu.lea (Experiment 2) and transformed into a quasi-continuous index (0–100). In Experiment 3 (no context), 30 low-literacy adults explained each idiom, providing a measure of functional familiarity. In Experiment 4 (same cohort after ~6 months; n = 26), idioms were embedded in Easy German contexts and tested using an endorsement task (idiom vs. paraphrase vs. distractor) and a ranking task (gold/silver/bronze).ResultsMixed-effects models showed that, in Experiment 3, figurative accuracy increased with literacy (β = 0.028, p = 0.0019) and was related to literal plausibility (β = 0.569, p < 0.001), with a strong transparency × plausibility interaction (β = −1.457, p < 0.001): idioms that were both transparent and literally plausible yielded the lowest figurative accuracy. In Experiment 4, idiom endorsement was predicted by functional familiarity from Experiment 3 (β = 0.702, p = 0.025), while ranking responses showed robust distractor rejection and no reliable disadvantage for idioms relative to meaning-consistent paraphrases.DiscussionOverall, idiom accessibility in low-literacy adults appears to be driven primarily by literacy, functional familiarity, and literal-figurative competition indexed by plausibility, whereas transparency does not emerge as a reliable stand-alone facilitator. These findings challenge Easy German recommendations that prioritize transparency and instead support applied strategies centered on audience familiarity, low literal plausibility, and contextual or paraphrastic scaffolding.
Prior research on so-called “elderspeak” or patronizing talk directed at older adults—often interpreted by communication accommodation theory (CAT) as a classic example of “overaccommodation”—has shown that such communicative behaviors are viewed negatively by recipients and observers. However, we provide evidence suggesting that these evaluations are by no means uniform and may vary across contexts. To account for this variability, we draw on system justification theory to argue that attributions and evaluations of unfavorably delivered messages may, nonetheless, be accepted insofar as they align with internalized social norms and contribute to maintaining a coherent representation of the social order. We argue that this theoretical position, aligned with CAT, offers a more innovative and comprehensive way of understanding how intergenerational communication may reproduce age-based inequalities.
Memory for who said what can help native and non-native listeners identify critical information for use in conversations. In this study, source memory for object-speaker associations was tested for listeners differing in language background and nativeness. Native and non-native participants of German first heard the voice of a child and an adult speaker name photographs of objects with age-typical characteristics (e.g., a board book for children vs. a standard hardcover novel) before they had to indicate who had previously named each object. Speakers either consistently named objects typical for their age group, or they randomly named objects from both age groups. Both listener groups associated objects with speakers successfully and showed better source memory in the age-consistent condition over the random condition. Critically, non-native listeners performed comparably to native listeners and benefited equally from age-typical patterns, suggesting that stereotypical knowledge about age can facilitate source memory encoding and retrieval even when processing a second language. The results confirm the role of source memory representations in communication and suggest that memory for object-speaker associations based on salient social categories like age remains robust in non-native listening.
In recent decades, there have been considerable advances in the area of bilingual cognition, but research on cognition across dialects remains limited. Here, we consider how the intuitive link between these experiences of linguistic diversity could drive empirical and theoretical work on the cognitive basis of dialect production. We begin with descriptive evidence to motivate contrasting hypotheses: the Difference Threshold Hypothesis, which argues for a qualitative difference between bilingual and bidialectal cognition; and the Linguistic Continuum Hypothesis, which argues that bidialectalism and bilingualism represent the same phenomenon across different degrees of linguistic overlap. To differentiate these hypotheses at the level of individual cognition, we argue that (1) research must expand beyond the lexical level to consider phonological and morphosyntactic differences across codes and (2) extra consideration is warranted for cases of bilingualism with apparent dense code-switching, which more closely approximates “bidialectal” practices.
Attention-Deficit/Hyperactivity Disorder (ADHD) is one of the most common neurodevelopmental disorders, characterized by symptoms of (i) inattention, (ii) hyperactivity-impulsivity or (iii) both. Atypical patterns of attention are a common phenotype in this population. To date, research on bilingualism in ADHD has not reached a consensus regarding the impact of bilingualism on executive functions (EFs) and ADHD-related behavior. This review focuses primarily on bilingualism in ADHD and highlights how outcomes within the ADHD literature, considering child, adolescent and adult populations, may vary systematically according to contextual moderators. First, socioeconomic status (SES) should be carefully controlled for, as it independently affects EFs, receptive vocabulary and attention abilities. Second, detailed profiling of language background is required to better interpret reaction time and accuracy findings, allowing performance differences to be linked to variables such as context of language usage, frequency of use, or number of languages acquired. Furthermore, implementation of tools such as the Q-BEx and LEAP-Q questionnaires is encouraged to precisely profile one's language background and capture language-experience variables, thus clarifying the role of bilingualism in cognitive performance in ADHD. Third, linguistic distance (LD) potentially imposes dynamic effects on EFs for bilinguals with ADHD with the literature showing bilingual advantages in EFs given one's proficiency and distance between languages. Co-occurring Autism Spectrum Disorder (ASD) is acknowledged given the comorbidity's prevalence rates in ADHD groups, as well as the fact that it has been associated with additional challenges in language and/or social skills. For future research, hypotheses regarding effects of bilingualism in ADHD groups with comorbid ASD are presented. Additionally, the incorporation of longitudinal studies in bilingual ADHD research is emphasized given that cross-sectional study designs can overlook outcomes related to EFs or ADHD-related behavior and diagnosis at different developmental time-points. Together, these insights call for a shift from global comparisons toward context-sensitive approaches that capture the diversity of bilingual experience in ADHD.
Lexical network analysis provides a powerful framework for examining how words are organized in the mind, as well as the semantic, categorical, and associative links that structure the mental lexicon. This study investigates the organization, structural properties, and types of lexical relations present in individual lexical networks of typically developing children and children with Down syndrome. Lexical relations were examined using network-based structural analyses derived from co-occurrence patterns, and measures of density, centrality, modularity, hubs, semantic clustering, and word-type distributions were computed, together with the analysis of taxonomic and thematic relations. The results reveal significant differences between the two populations, despite comparable vocabulary sizes. These differences are discussed in terms of variation in the mechanisms underlying lexical acquisition and organization, as well as the role of experience in psycholinguistic development. Overall, this study provides empirical evidence on early lexical organization in typical development and Down syndrome, with implications for theories of language development and for educational and therapeutic approaches aimed at strengthening vocabulary.
IntroductionGerman modal particles pose a major challenge for translation because their meanings are primarily discourse-pragmatic, interactional, and context-dependent rather than lexical or propositional. This study examines how the German modal particle denn is rendered in Turkish learner translations of literary texts and why it is frequently omitted.MethodsThe empirical material was drawn from Thomas Mann's Der Zauberberg. A corpus search yielded 848 occurrences of denn. After excluding homonymous and non-modal uses, 200 modal-particle occurrences were identified, from which 29 examples were selected according to their pragmatic functions. Fifteen Turkish-speaking students of German Language and Literature at approximately B2 level translated these examples into Turkish, resulting in 435 translation units. In addition, semi-structured interviews were conducted with five participants. The translations were coded according to four strategies: omission/zero correspondence, functional substitution/transposition, paraphrase, and literal translation.ResultsOmission was the dominant strategy, occurring in 333 cases (76.55%), while functional substitution/transposition occurred in 102 cases (23.45%). No cases of paraphrase or literal translation were identified. Functional substitutions were realized through Turkish particles, adverbs, connectors, interjections, and pragmatic formulas, especially ki, peki, zaten, gerçekten/gerçekten de, and bu yüzden.DiscussionThe interview data indicate that omission was linked to insufficient pragmatic awareness, perceived structural differences between German and Turkish, the perception of denn as optional, simplification strategies, and a lack of explicit translation strategies. The study shows that translating denn requires pragmatic awareness and context-sensitive translation competence rather than direct lexical equivalence.
This study revisits English stop voicing in Arabic-speaking learners through a design that combines perception, production, phonological context, and a brief instructional manipulation within the same participants. Forty participants took part: 30 Arabic-speaking learners (10 Novice-High, 10 Intermediate-High, and 10 Advanced) and 10 native English controls. Learners completed perception and production pretests, a live standardized mini-lesson, and immediate posttests in the same session; controls completed a matched interval without instruction. Perception cue weighting was estimated with participant-by-time logistic coefficients for voice onset time (VOT), onset F0, and vowel duration. Production analyses were separated into singleton /p/-/b/ contrasts and /sp/ aspiration suppression so that laryngeal category and context were not conflated. Learners underweighted VOT at pretest relative to native controls, with the largest gap in the novice group, and the learner groups showed immediate posttest increases in VOT weighting. Secondary-cue reliance was also proficiency-sensitive: novice and intermediate learners entered with heavier duration weighting, which declined after instruction. In production, learners produced smaller singleton VOT contrasts than native controls and showed persistent difficulty suppressing aspiration in /sp/ onsets, especially at lower proficiency levels. Immediate posttest adjustment was strongest in the same learners who had shown the weakest pretest control. Participant-level correlations linked change in perceptual VOT weight to change in both /sp/ suppression and singleton voiceless VOT. The results are consistent with gradient, shared phonetic restructuring, but the single-session design warrants a cautious interpretation of instruction as short-term adjustment rather than durable learning.
Normed linguistic stimuli are fundamental in psycholinguistics because they capture lexical and semantic properties that influence comprehension. However, generating these norms at scale is challenging, often leading researchers to rely on ad hoc norms collected from small samples, which can introduce inconsistencies and limit cross-study comparisons. In the present study, we investigated how large language models (LLMs) can support psycholinguistic research by prompting eight current LLMs to norm 300 English two-word metaphor combinations, such as sharp mind. We selected the dimensions of familiarity, aptness, concreteness, metaphoricity, and constituency, as these tap distinct cognitive processes and may provide insight into which aspects LLMs capture accurately and which they do not. We varied stimulus presentation (in context vs. in isolation) and response format (categorical vs. numerical) to examine which manipulation yields norms most closely aligned with human ratings. We then assessed the reliability and validity of model responses and used them to replicate existing analyses of metaphor comprehension. Overall, LLM-generated norms aligned best with familiarity and metaphoricity, which rely on word co-occurrence. In contrast, aptness, concreteness, and constituency—which require reasoning about the relationship between the topic (e.g., mind) and the vehicle (e.g., sharp)—proved more challenging for LLMs.
Adult language learners frequently struggle with possessive pronouns, and classroom evidence confirms these difficulties for learners of German. This study examines the acquisition of possessive agreement in L3 German by L1 Polish-L2 English speakers, focusing on third-person singular pronouns. The primary aim is to determine whether difficulties with possessive agreement arise from structural differences between Polish and German, and to explore the role of developing proficiency in L2 English, a language that is invariably part of the Polish education system, in the acquisition process. Using a subtractive language group design, L3 learners were compared with L1 English-L2 German learners to isolate the effect of Polish. Participants completed an untimed acceptability judgement task comprising felicitous and infelicitous items, further manipulated for possessor gender and for gender match between the possessor and the possessee noun. The results show that both groups performed similarly, indicating no non-facilitative influence from L1 Polish. Moreover, higher general L2 English proficiency predicted better performance in L3 German, suggesting proficiency-dependent facilitation from structurally similar English. Finally, participants' accuracy for infelicitous items was below chance, confirming that possessive agreement constitutes an inherently challenging domain in adult language learning.
Glossing has been extensively employed and researched as a tool for vocabulary learning, with numerous studies investigating how it aids second language incidental vocabulary learning by directing learners' attention to the connection between word forms and their meanings. Nevertheless, findings across studies have been mixed, highlighting the need for a comprehensive and systematic synthesis. In response, the present study conducted a meta-analysis of 26 empirical studies on English as a second or foreign language to provide a quantitative synthesis of how glossing affects incidental English vocabulary learning through reading across varied learning conditions, and to examine variables that may moderate its effectiveness within the Involvement Load Hypothesis Plus (ILH+) framework. The results showed that (1) glossing was associated with a significant facilitative effect on L2 incidental vocabulary learning, although this pooled estimate should be interpreted as a broad average across highly heterogeneous study conditions rather than as a uniformly robust general effect; (2) different operationally defined glossing subgroups were associated with varying effect-size estimates, with interactional and experimental contrasts showing relatively larger estimates and positional and modality-based contrasts showing relatively smaller estimates; and (3) input frequency, test length, treatment duration, and vocabulary knowledge type emerged as significant moderators.Systematic review registrationhttps://www.crd.york.ac.uk/PROSPERO/view/CRD420251237727, identifier: CRD420251237727.
In public and academic debates, the gender star has been criticized for potentially hindering text comprehensibility and imposing an additional burden on learners of German by complicating an already complex language. We investigated in an experiment whether the use of the gender star (e.g., Student*innen) complicates text comprehensibility compared to masculine generics for learners of German. Participants were 80 Flemish students studying German as a foreign language. They were asked to read 13 short news articles, of which ten texts served as experimental items in either gender star or masculine form. Participants' accuracy in answering content questions, subjective ratings of text comprehensibility and sentence difficulty, and their full-text reading times were assessed. Additionally, we collected data on participants' proficiency in German, prior knowledge of, and attitudes toward gender-fair German. Our results suggest that the gender star does not pose a strong hindrance to text comprehensibility for Flemish students learning German: while proficiency in German had a significant effect on content question accuracy and sentence difficulty as well as subjective comprehensibility ratings, this effect was independent of gender form.
Human language likely emerged from pre-symbolic social cognition rather than from a sudden, language-specific innovation. Across social animals and human infants, gaze following, affective attunement, and intention reading point to early forms of mind reading that precede explicit theory of mind and symbolic communication. In development, language and theory of mind appear to support each other: prelinguistic social cognition scaffolds later belief reasoning, while growing linguistic competence enables more flexible representations of others' mental states. In evolution, this trajectory may be explained by a neural threshold hypothesis, according to which quantitative increases in hominin brain size and connectivity yielded qualitatively new computational capacities. On this view, Homo erectus may have marked a critical transition, with expanded working memory, hierarchical integration, and social inference supporting rudimentary symbolic thought and early language. Rather than attributing language to a single mutation or an isolated neural mechanism, this account proposes that it emerged from an interacting trait package, a synthesis shaped by ecology, energetics, development, and social behavior. Modern language is therefore best understood as a culturally elaborated expression of a deeper neurobiological capacity rooted in early hominin evolution.
IntroductionThis study investigated the influence of grammatical information-specifically, the use of masculine forms as generic—on sentence comprehension in Italian.MethodsA sentence evaluation paradigm was used. In each trial, a first sentence introduced a professional role name (e.g., Gli architetti uscivano...), followed by a second sentence providing gender information (e.g., una delle donne...). Participants were asked to judge whether the second sentence was a sensible continuation of the first.ResultsResults indicated that participants' judgements were biased by the masculine plural form used as generic. The analysis of reaction time data showed that female continuations led to longer reaction times than male continuations, consistent with the judgement data.DiscussionThese findings align with align with previous research showing that in gender-marked languages, such as German and French, masculine generics bias mental representations in an androcentric way.
This study investigated the role of L1 influence on how L2 speakers interpret aspectual semantics of English past simple accomplishments. English past simple is aspectually underspecified and is thus vulnerable to the transfer of L1 aspectual representations. Slavic languages (Russian, Polish) grammaticalize dual perfective-imperfective aspect. Norwegian does not grammaticalize aspect. L1 Polish (n = 57), Russian (n = 20), and Norwegian (n = 50) speakers participated in two web-based Visual World Paradigm eyetracking experiments. Offline judgments and online gaze preferences from both experiments showed that L1 Slavic speakers associated English past simple with completed event pictures. This categorical association strengthened in offline judgments with higher L2 proficiency. L1 Norwegian speakers associated English past simple with ongoing events both online and offline. This association weakened in offline judgments with higher L2 proficiency. These findings provide evidence of L1 transfer during L2 aspectual interpretation and elucidate the interaction between crosslinguistic influence and L2 proficiency. Implications for theoretical models of cross-linguistic influence are discussed.
In this paper we raise the question “why is agreement so common across natural languages?”. We will argue that the challenge of grammar inference in natural and artificial languages provides key insights into the ubiquity of agreement. By grammar inference, we mean the discovery of a procedure that (i) determines string well-formedness on the basis of exposure to unannotated expressions of a language; and (ii) allows for the construction of structural descriptions for well-formed strings. The idealized version of this problem results in the identification of the generator of a stringset, or -more realistically- a restriction of the class of possible generators. We argue that agreement plays a crucial role not only in flagging dependencies between expressions at the string level, but also, considering that agreement relations occur in restricted structural configurations, in restricting the class of structural descriptions compatible with a string. As such, agreement mediates between strings and structure, providing a parser with information to solve the grammar inference problem. We will furthermore argue that the mechanisms involved in grammar identification are not restricted to natural language acquisition and processing, but in fact extend to a class of problems that motivated much research in the theory of symbolic encoding of dynamical systems and machine learning.
Linguistic labels have been shown to facilitate visual recognition and categorization. Labels have also been shown to be better facilitators than sounds. However, less is known about how nouns in contrast with verbs and non-verbal sounds differentially activate action-related conceptual information. In the present study, we investigated whether words and non-verbal sounds vary in their ability to facilitate visual processing depending on their congruence with an implied action. Participants were presented to an image verification task in which visual stimuli were preceded by either nouns, non-verbal sounds or verbs that were either congruent or incongruent with the action associated with the target stimulus. Reaction times were analyzed using linear mixed-effects models. Results revealed a robust interaction between cue type and action congruence. Non-verbal sounds and verbs produced significantly faster responses when the action was congruent compared to incongruent trials, whereas nouns were optimal only when action was incongruent. These findings suggest that action information plays a role in the representations activated by words and non-verbal sounds. In addition, the evidence shows that verbs appear to activate action-related information similar to that evoked by non-verbal sounds. More broadly, the results contribute to ongoing debates about the mechanisms by which language and non-linguistic cues shape perceptual processing, highlighting the role of action-related cues in predictive conceptual activation.
This study examines voice onset time (VOT) in Arabic first-language speakers learning English and Spanish as a multilingual test of the speech learning model-revised (SLM-r). The model predicts continuous phonetic learning, interaction among categories in a shared phonetic space, and bidirectional influence across languages. VOT production was examined in Arabic monolinguals, an Arabic–English bilingual group, an Arabic–English–Spanish multilingual group, and monolingual speakers of English and Spanish using a controlled reading task. Mixed-effects models incorporated group, phonetic condition, and continuous measures of proficiency and relative dominance. English aspirated /p/ was robustly differentiated from English /p/ after /s/ across groups, with a 57.21 ms model-estimated contrast. Spanish /p/ remained in the short-lag range and did not differ reliably between the multilingual group (M = 15.10 ms) and Spanish monolinguals (M = 14.17 ms). For voiced stops, Arabic monolinguals showed substantially more negative Arabic /b/ VOT (M = −85.55 ms) than the Arabic-English (M = −44.08 ms) and Arabic–English–Spanish (M = −43.85 ms) groups, indicating reduced prevoicing in the Arabic L1 learner groups. English /b/ and Spanish /b/ also showed higher prevoicing among Arabic L1 learners than among the respective monolingual controls. Relative English–Spanish dominance did not reliably predict voiceless-stop VOT within the multilingual group. The findings support SLM-r claims about lifelong phonetic plasticity and selective cross-language interaction, while also suggesting that Spanish /p/ may reflect facilitative transfer from Arabic short-lag timing as well as stable category separation.