
Abstract Based on the frequency and functions of pragmatic features used in 24 conversations of different lengths between tourism workers and tourists recorded in two tourist information centers in Zagreb, Croatia, this paper investigates the pragmatics of advising. The paper unveils which specific pragmatic strategies are used in each of the examined subsets and which functions they fulfill. Findings suggest that time-sensitive conversations include slightly fewer but more complexly used pragmatic features in the following order of frequency: solidarity markers > discourse markers > repetitions > hesitation markers > rephrasings = multilingual elements = co-constructions. In conversations that are longer and not time-sensitive, pragmatic strategies are used less frequently than expected and in the following order of frequency: discourse markers > hesitation markers > multilingual elements > solidarity markers > repetitions = code-switches > wrap-ups. Furthermore, the findings suggest that solidarity markers, especially okay , function as solidarity and discourse markers, representing a common adaptation of conversational norms in interactions within transient groups. Building on these insights, the paper concludes that the setting of the interaction is more influential on the tourism workers’ use of pragmatic features than their shared L1 and sociocultural background.
Abstract Phonological alternations that arise due to morphological concatenation often serve to enforce static phonotactic generalizations within the lexicon. Due to this close link, knowledge about alternations and phonotactics are encoded using a single mechanism in constraint-based models like Optimality Theory. Evidence of this link in learning, however, is equivocal. In this study we examine whether learners readily extend a generalization from alternations to phonotactics. Learners were trained on consonant harmony alternations that involved harmony between a suffix and stem consonant, while stems were ambiguous in terms of phonotactic generalizations. We then examined whether learners extended the alternation generalization to phonotactic judgments about novel stems. While learners successfully learned the alternations in training, they failed to extend this generalization to phonotactic judgments about novel stems. Taken together with previous findings, our results suggest that learners are conservative in extending learned generalizations to unseen morphological contexts. Overall, these results suggest that the link between phonotactics and alternations might not be as tight as previously suggested.
Abstract Generative AI tools have implications for corpus linguistics from practical and theoretical angles, both of which the current paper addresses. The first part of the paper establishes whether and how generative AI tools – here ChatGPT o3 – can in practice support a corpus-linguistic workflow comprising data extraction, cleaning, and annotation for the dative alternation. While certain steps in this workflow such as data extraction or object length annotation can be completed reliably in ChatGPT o3, more taxing processes like cleaning extracted data for relevant instances, identifying clause boundaries and clause elements, or annotating objects for animacy are marked by errors and inconsistencies, which even the addition of steps dedicated to facilitating AI support in the corpus-linguistic workflow cannot remedy. The second part of the paper relates to corpus-linguistic theory by discussing whether and how AI-generated texts should be represented in linguistic corpora. It argues that the notions of authenticity and representativeness as key features of linguistic corpora are compatible with the integration of AI-generated texts in linguistic corpora. As said integration marks a departure from corpus-linguistic tradition, various arguments in favour of and against featuring AI-generated texts are provided, yielding a pragmatic practical suggestion for future corpus projects.
Abstract This paper revisits the question of mode as an influencing factor for variation in clausal complexity. In a comparative corpus study using the Lang*Reg corpus, clausal embedding is compared between the phonic and graphic mode produced by the same language users with the same linguistic content in German and Persian. The corpus study reveals that mode affects syntactic complexity in terms of clausal embedding in Persian as well as in German. Participants used a higher number and higher depth of clausal embedding in the written storytelling context compared to the oral storytelling context. With respect to types of embedded clauses, the study shows that language users mainly increase their use of adverbial clauses but not complement or nominal dependent clauses due to mode.
Abstract Hybrid nouns such as Mädchen (‘girl’) pose a challenge to agreement in German, as they combine neuter grammatical gender with the semantic feature [+female]. This study investigates how modality and structural factors shape the choice between grammatical and semantic agreement in pronominal reference. Building on earlier work, the analysis draws on a large corpus dataset from the German Web Corpus (deTenTen; n = 1,628), which contains conceptually oral online discourse. The results show that semantic agreement strongly dominates (88 %), with grammatical agreement persisting mainly in structurally bound positions (e.g., relative pronouns) and in contexts of immediate proximity between antecedent and target. On the one hand, these findings confirm the predictions of the agreement hierarchy; on the other hand, they underscore the relevance of modality: compared to norm-oriented written registers, conceptually oral registers favor semantic over grammatical agreement even more strongly. These results contribute to the understanding of how modality impacts linguistic choices and agreement resolution strategies.
Abstract While previous studies have examined various aspects of Ghanaian parliamentary discourse, little attention has been paid to evidential constructions. To address this gap, this corpus-based study examines four first-person evidential constructions ( I think , I believe , I know , I mean ) in Ghanaian parliamentary discourse and traces changes in their functions over time. A corpus of 1,729 parliamentary Hansards (41.7 million words) spanning 2005–2024 was processed using Python 3.11 on Google Colab. Findings show that subjective inference markers ( I think , I believe ) are dominant, while the direct evidence marker ( I know ) and the discourse repair marker ( I mean ) were less frequent. A diachronic functional analysis shows systematic changes in the functions of these markers over time, with chi-square tests confirming statistically significant differences in the frequencies of these markers across five electoral periods. The findings contribute to understanding how evidential practices evolve in parliamentary discourse, highlighting gradual shifts in communicative strategies.
Abstract This paper examines the syntactic properties of the taboo-based intensifier mayyit ‘dead’ in Urban Jordanian Arabic (UJA). We argue that mayyit developed its intensifying function through metaphorical extension rather than grammaticalization. Despite this shift, mayyit retains its adjectival category and shows agreement and other features of lexical items. Two forms are attested: a prepositional form and a non-prepositional, compound-like form. While semantically similar, they differ syntactically, with the latter displaying traits of synthetic compounds. We propose a unified analysis in which both forms are subject–predicate structures, with mayyit heading its own projection and selecting a nominal or prepositional complement. This parallels genitive constructions in UJA, including construct and free state. The study contributes new data on taboo-based intensifiers and offers broader insights into their structure and typological variation.
Using the examples of her, hers, herself, and her own, this paper examines the internal syntax and morphological properties of English pronouns through the lens of Distributed Morphology. The central argument is that her own acts as a complex reflexive possessive pronoun, serving as the functional equivalent of the nonexistent form herself's. Specifically, the possessive pronoun her is not an inherent reflexive; rather, its reflexive character in specific contexts is derived from the presence of the reflexive pronoun own in the syntax, which may be absent or present at PF. With own, her (own) functions as a reflexive pronoun governed by Binding Condition A, thereby avoiding a Condition B violation. As regards the derivation of her own, formal features such as phi-features [phi], possession [poss], and reflexivity [ref] are distributed across multiple nodes in the syntax, which ultimately cluster as single words her and own. The nonexistent form self's is blocked by own for reasons of economy.
Russian-speaking Koryo-saram descend from Koreans who migrated to Soviet Russia in the nineteenth and twentieth centuries and were later forcibly deported to Central Asia by the Soviet government. Following the Soviet Union's collapse, many Koryo-saram responded to South Korea's diaspora engagement and migrated there to seek job opportunities, later constructing enclaves featuring Russian-language signage and imagery and products from Russian-speaking countries. Employing Bourdieu's concept of the linguistic market which posits that language users occupy different social positions based on their linguistic capital, this article analyses interviews with Russian-speaking bi/multilingual Koryo-saram migrants who possess at least some Korean proficiency. In their metapragmatic evaluation of the enclaves' Russian-language signage, the interviewees reproduce dominant ideologies that necessitate migrants' linguistic assimilation into South Korean society, drawing on their linguistic capital (Korean proficiency) to invoke perceived differences in socio-economic integration between themselves and most Koryo-saram who are monolingual Russian speakers. They contrast Korean with Russian, framing the former as linguistic capital for accessing socio-economic opportunities in South Korea and the latter as linguistic deficit that engenders the socio-economic marginalization of Koryo-saram in the country. This study underscores how bi/multilingualism becomes a resource for negotiating social (im)mobility during ethnic return migration.
This work investigates why recurrent neural networks (RNNs) tend to learn phonological patterns that are unattested or dispreferred by humans. Specifically, we explore the hypothesis that their over-generation is caused by their excess expressive capacity - they are beyond the limited complexity class that contains the set of attested phonological patterns. We compared these over-expressive RNNs against the weaker convolutional neural networks (CNNs) on a battery of string recognition tasks. We find that the expressivity of a model's architecture does not predict the string classes that it excels at recognizing. Instead, we suggest that CNNs' position-invariant biases better explain their successes in our experiment.
In health communication, access to language is essential during emergencies and crises. The Guangdong-Hong Kong-Macao Greater Bay Area (GBA), a region of superdiverse multilingualism through increasing migrant flows, features the coexistence of languages (e.g., Chinese, English, and Portuguese) and dialects (e.g., Cantonese, Hakka, and Min). This research studies how multilingual resources are represented and how multilingual communication is practiced. It examines physical representations evidenced in the linguistic landscape of hospitals in the GBA, drawing on 460 real-world signs, while investigating 10 practitioners' experience via in-depth interviews. The multilingual landscape, although an essential marker of multilingual awareness and professional identity, exhibits inequivalence in information load and mistranslations, which overshadow the functional meanings of multilingual signs. These function more as cultural tokens and do not prioritize accuracy. Four factors shape multilingual communication: (1) accessibility to professional interpreters, (2) cultural and cognitive gaps, (3) technology-enhanced practice, and (4) ethical considerations. Comparative analysis of the physical and practical dimensions of multilingual communication reveals a disjunction between symbolic inclusion and communicative functionality, institutional and ethical paradoxes, and the stratification of languages and linguistic justice. We consider performative, unregulated, and hierarchical multilingualism to shed light on a more critical understanding of multilingual healthcare communication.
This paper presents a study of regional and ethnic variation in the use of morphosyntactic features associated with African American English (AAE) among 35 Black high school students in Philadelphia and Boston. Twenty-five morphosyntactic features are examined by three measures of use: range of features, dialect density measure, and copula deletion rates. Results demonstrate marked differences in the use of AAE morphosyntax between cities, with Philadelphia speakers showing wider ranges, higher dialect density measures, and higher rates of copula deletion. Within ethnically diverse Boston, speakers with Afro-Latinx ancestry have the most restricted use of AAE features. This work represents the first description of AAE morphosyntax in Boston, and contributes to the study of diversity both within and between Black communities.
Translation is critical in development work in Vietnam, especially when introducing largely Western and sometimes ill-defined development concepts into local knowledge systems. Our study analysed online, semi-structured interviews with 18 development stakeholders in Vietnam and a 1.1-million-word bilingual corpus of development texts to examine how participants designate and perceive the concept of empowerment in their contexts and practices of translation when aiming for empowerment in their work. Our study suggests participants use varied equivalents in Vietnamese for empowerment, with an interesting tension between designations that could imply either power or rights. This tension led to different perceptions among participants of the translation of empowerment into Vietnamese: as a potential source of misunderstanding and political sensitivity, or as a possible catalyst to promote rights and empower marginalized groups. Our study also reveals that the Vietnamese practice of n & ocirc;m na - a form of intralingual translation and simplification - is a core competence in these stakeholders' translation work. N & ocirc;m na is used by them to empower stakeholders by generating accessible language, eliciting feedback, and permitting negotiation of power relations in development encounters. Such empowerment through contextually and culturally sensitive translation supports stakeholders' rights to participate in, contribute to, and enjoy development. Dich thuat & dstrok;& oacute;ng vai tr & ograve; thiet yeu trong c & ocirc;ng t & aacute;c ph & aacute;t trien tai Viet Nam, & dstrok;ac biet khi & dstrok;ua c & aacute;c kh & aacute;i niem tu phuong T & acirc;y v & agrave;o thuc tien & dstrok;ia phuong. Nghi & ecirc;n cuu n & agrave;y ph & acirc;n t & iacute;ch du lieu phong van voi 18 nguoi l & agrave;m viec trong l & itilde;nh vuc ph & aacute;t trien v & agrave; mot kho ngu lieu song ngu gom 1,1 trieu tu, nham t & igrave;m hieu c & aacute;ch ho tiep nhan v & agrave; dich kh & aacute;i niem empowerment. Ket qua cho thay nhung nguoi tham gia & dstrok;& atilde; su dung nhieu c & aacute;ch dich kh & aacute;c nhau; & dstrok;ieu n & agrave;y phan & aacute;nh nhung c & aacute;ch hieu v & agrave; dien giai kh & ocirc;ng thong nhat ve empowerment. Su kh & aacute;c biet n & agrave;y dan toi c & aacute;c quan & dstrok;iem tr & aacute;i nguoc: c & oacute; nguoi xem & dstrok;& acirc;y l & agrave; van & dstrok;e nhay cam, trong khi nguoi kh & aacute;c lai coi & dstrok;& oacute; l & agrave; co hoi mo rong c & aacute;c kh & iacute;a canh ve quyen cho c & aacute;c nh & oacute;m yeu the. Nghi & ecirc;n cuu c & utilde;ng chi ra rang n & ocirc;m na (c & aacute;ch dien giai & dstrok;on gian trong tieng Viet) l & agrave; c & ocirc;ng cu quan trong v & agrave; huu & iacute;ch, v & igrave; c & aacute;ch tiep can n & agrave;y & dstrok;am bao noi dung & dstrok;uoc tr & igrave;nh b & agrave;y r & otilde; r & agrave;ng, de hieu, & dstrok;ong thoi tao & dstrok;ieu kien cho su tham gia v & agrave; th & uacute;c & dstrok;ay quyen con nguoi trong c & aacute;c hoat & dstrok;ong ph & aacute;t trien.
This case study article demonstrates the linguicism in French emergency planning which became evident as a result of Cyclone Chido hitting Mayotte, a French overseas territory, on 14 December 2024. It investigates the communicative practices and legislative framework adopted to inform the residents and evaluates them in relation to literature on multilingual crisis communication to assess how French emergency planning and crisis management response accommodated for the diverse linguistic and cultural needs of the people of Mayotte. Through an examination of emergency planning policies and various levels of national and local communication, this case study presents how Metropolitan France's approach to Mayotte's multilingualism appears to be disconnected from the needs of the local population, and the importance of including local stakeholders in crisis communication strategies. The disconnection is here investigated in terms of disaster linguicism. Potential areas for future research emerge from the evidence of failures in supporting Mayotte's populations in accessing critical information. By looking at the case of Cyclone Chido through a linguistic and cultural lens, this article advocates for additional research into multilingual support in Mayotte, the practices of supporting non-Francophone populations, and the notion of co-design and co-creation of materials to enhance emergency communication planning. Cette & eacute;tude de cas d & eacute;montre la discrimination linguistique qui est devenue & eacute;vidente dans la planification d'urgence fran & ccedil;aise & agrave; cause du Cyclone Chido qui, le 14 d & eacute;cembre 2024, a frapp & eacute; Mayotte, un territoire d'outre-mer fran & ccedil;ais. Nous enqu & ecirc;tons sur les pratiques de communication et sur le cadre l & eacute;gislatif adopt & eacute;s pour informer les r & eacute;sidents et nous les & eacute;valuons en les comparant avec de la litt & eacute;rature sur la communication de crise multilingue afin d'& eacute;valuer comment la planification d'urgence et la r & eacute;ponse de gestion de crise fran & ccedil;aises ont r & eacute;pondu aux diverses besoins linguistiques et culturelles des mahorais. En examinant les politiques du plan d'urgence aussi bien que de plusieurs niveaux de communication nationale et locale, cet article pr & eacute;sente la mani & egrave;re dont l'approche de la France m & eacute;tropolitaine envers le multilinguisme de Mayotte para & icirc;t & ecirc;tre d & eacute;connect & eacute;e des besoins de la population locale. Cette d & eacute;connexion s'enqu & ecirc;te ici en termes de discrimination linguistique dans des situations de crise. Cet article milite pour plus de recherche dans le soutien multilingue & agrave; Mayotte, dans les pratiques de soutenir les populations qui ne sont pas francophones, et dans les notions de co-design et de cocr & eacute;ation des mat & eacute;riaux en incluant des parties prenantes locales pour am & eacute;liorer la planification d'urgence en termes de communication.
Language production includes multiple subsystems such as the lexical, syntactic, and semantic domains. These subsystems do not operate in isolation but instead interact during production. Previous research has documented complexity trade-offs, cases in which an increase in complexity in one linguistic domain tends to be accompanied by a reduction in others. However, existing studies have primarily focused on trade-offs between syntax and lexicon using small datasets. Our study is the first to extend research on trade-offs into the semantic domain. Two innovative semantic complexity indices (semantic specificity and semantic richness) are incorporated to offer insights into the interaction with semantic complexity. Additionally, we evaluate the effect of sentence length. We conduct Spearman correlation tests and fit generalized additive mixed models to assess the trade-offs. Our results provide converging evidence of complexity trade-offs in natural language production across lexical, syntactic, and semantic domains. We then confirm that longer sentences make complexity trade-offs more prominent. Complexity trade-offs may be attributed to the constraints imposed by the limitations of cognitive resources and the need for efficient communication. Our findings may represent a universal feature of languages, offering insights into how efficiency and human cognitive resources shape language production.
This study explores how the climbing-centred Instagram account @cruxingincolor showcases the specialized register of mountaineers as part of its "Term Tuesdays". The transmission of mountaineering vocabulary via Instagram represents a kind of "citizen lexicography", with the affordances of social media platforms being utilized in creative - but also ultimately transient - ways to give attention to a highly specialized (and, on a global scale, relatively small) community and its register. Furthermore, by investigating the definitions of showcased mountaineering expressions and how they are presented on @cruxingincolor, the study underscores that mobility is a concept inherent to the sport to the extent of being an important part of many of the sport's key expressions. The results reveal that, while the investigated entries typically focus on physical mobility, some also pertain to aspects such as the impact of mountaineering on the environment or its role in upcoming events. Thus, overall, the contribution shows that the citizen lexicography approach evident on @cruxingincolor is affected by and represents multiple layers of transient mobility.
This paper investigates the mechanics of phi-feature valuation by comparing agreement patterns in French and Arabic, showing how differences in [+/- human] specification regulate the outcomes of the Agree operation. French exhibits maximally transparent agreement: probes fully value phi-features regardless of the controller's [+/- human] specification. In Arabic, [+plural, -human] controllers, whether animate or inanimate, regularly yield deflected agreement, surfacing as third person feminine singular morphology. We argue that [-human] disrupts phi-feature valuation without eliminating Agree, thereby yielding a distinct failed-agreement default, which is separate from the no-agreement default (third person masculine singular). This dual-default system challenges the standard view of default agreement as morphologically uniform and reveals a three-way outcome for Agree: full valuation, failed valuation, and no valuation. The findings refine the typology of agreement mismatch and show that semantically interpretable features can be syntactically active , influencing valuation pathways and morphological realization in ways unattested in morphosyntactically inert systems like French.
Online travel reviews (OTRs) are an under-researched digital genre directly reflecting users' geographical and digital mobility. Based on a corpus of hotel and restaurant reviews from Airbnb and Tripadvisor as well as a control corpus of reviews from a Lonely Planet travel guide, this paper analyzes the linguistic means used to construe spatiality and mobility in OTRs. The analysis reveals that linguistic descriptions of spatiality and mobility are central to OTRs. In addition to the linguistic strategies described in the literature, noun phrases referring to salient elements of the spatial scene frequently occur. Deictic projection is the norm, with the origo typically shifted to the reviewed establishment to convey involvement or highlight salient features. Subcorpus differences mirror thematic priorities: Tripadvisor reviews focus on food, drink, and service, whereas Airbnb reviews emphasize location. These emphases are reflected in the relative frequencies of noun phrases, deictics, and (non-deictic) static or dynamic spatial descriptions. These patterns show that OTRs have developed (sub)genre-specific linguistic conventions. More broadly, these findings suggest that OTRs function not only as evaluative or informational texts but also as socially situated performances within digitally mobile, transient communities.
Dependency distance is a widely used measure of syntactic complexity, often considered to reflect cognitive abilities such as memory. However, empirical support for this claim remains limited. This study aims to examine whether dependency distance (DD) measures are associated with global cognitive function and single-domain memory performance, and identifies which variant best reflects these abilities. A sequential picture description task elicited clinical speech data from Mandarin-speaking older adults with varying cognitive levels. Multiple regression analyses are used to examine relationships between several DD variants and cognitive performance, as measured by the Montreal Cognitive Assessment-Basic (MoCA-B) and a single-domain Memory Index Score. Results show that all DD variants associated with sentence length are significantly correlated with global cognitive level, whereas removing the influence of sentence length substantially reduces their discriminative power. Among them, mean dependency distance per sentence is the best indicator of cognitive status. In contrast, no DD measure shows a significant association with the memory index. These findings suggest that DD measures can capture speakers' global cognitive status, with sentence length playing a key role. The study deepens our understanding of the relationship between language and cognition, providing empirical evidence for the Cognitive Commitment in usage-based linguistics.
In this article, we examine subordinate causal clauses headed by the declarative complementizer & zdot;e 'that' in Polish. By comparing them with canonical causal clauses headed by poniewa & zdot; 'because', we argue that causal & zdot;e-clauses function as speech act modifiers. The main evidence for this claim comes from: (i) movement to the left periphery of the matrix clause, (ii) variable binding, and (iii) sensitivity to material associated with the TP domain, CP domain, and the illocutionary force of the matrix clause. Additionally, we argue that they possess distinct discursive and semantic properties, i.e., they are non-at-issue and express epistemic causality. Diachronically, we provide evidence showing that & zdot;e 'that', as a complementizer, has undergone semantic narrowing, leading to its restricted behavior at the syntax-semantics interface in Present-Day Polish, in contrast to inherent causal complementizers.