
Abstract The observation that wh -expressions in a lot of languages can be interpreted either interrogatively or non-interrogatively has led to the widely accepted thesis that wh -expressions originate as variables, with their interpretations determined through being bound by a question operator or other non-interrogative operators ( Cheng 1991 , Li 1992 , Tsai 1994 , Cheng & Huang 1996 , among others). However, recent studies on a specific non-interrogative wh-construction in Mandarin, known as “ wh -conditionals,” challenge this variable-binding approach ( Liu 2016 , 2017 , Xiang 2016 , 2020 , Li 2021 ). In wh -conditionals, two wh -clauses are connected to express a conditional meaning. These studies advocate for a question-based approach, proposing that the meaning of wh -conditionals stems from interrogative meaning rather than variable binding. As a result, the question-based approach predicts that wh -conditionals should be linguistically closer to wh -questions than to other non-interrogative wh -constructions. This paper confirms the prediction using wh -constructions in Khorchin Mongolian. Despite their distinct theoretical commitments, the two approaches can be reconciled to address challenges pointed out by Cheng & Huang (2020) , Pan & Kuang (2023) , and Tsai (2023) to the question-based view.
Abstract In Eurasia, the hotbed of classifier languages as an areal feature is in East Asia, Southeast Asia, part of eastern India, and a less concentrated cluster in the Middle East and Central Asia. In Europe, only a few languages have been claimed to employ numeral classifiers. Russian, located in Eastern Europe, has a tripartite numeral construction [Num X N], where X’s are considered sortal classifiers by some researchers, e.g., Sussex (1976) and Goto (2012) , and is listed as a classifier language in the World Atlas of Classifier Languages (WACL). However, the status of Russian as a classifier language is controversial. In this paper, we first adopt the definition of classifier languages proposed by Her, Hammarström & Allassonnière-Tang (2022) . We then apply explicit syntactic criteria to evaluate Russian numeral constructions and examine the tripartite [Num X N] structure. Ultimately, our analysis demonstrates that the elements previously identified as sortal classifiers in Russian are in fact measure nouns, and we reject their classification after examining the reasons behind this misidentification. The findings suggest that it would be prudent to reexamine all putative classifier languages in Europe.
The Malay berprefix has not been discussed as extensively as other Malay verb prefixes. It has been accorded fuzzy meanings, such as "to mean something taking place" and "to be able to combine with any base form." Such vague definitions have hindered and prevented a thorough understanding of it. Our analysis took into consideration a middle voice marking system of situation types established by Kemmer (1993, 1994) and refined the categories based on instances collected from a large annotated corpus, the Malaysian Malay News Corpus (Chung & Shih 2019). From the corpus, 262,210 instances of berwere retrieved, analyzed and categorized into 14 distinct functions. The results revealed that 93.33% of the total word-tokens of berwere middle markers, whereas the remaining 6.67% were non-middle markers. We also refined and added new categories not mentioned in the past (e.g., "State of Having ROOT," and "Stative"). Some of the middle functions of ber-, especially the most cited and prototypical "Grooming/Body Care" function, were not the most frequently found functions. This indicated that berin Malay, as a middle marker in the form of reflexives in Malay, was not frequently found in the corpus.
In this paper, I argue that the auxiliary-like preverbal lai 'come,' which frequently appears before purposives in Mandarin Chinese, is actually a light verb located in the v position. The current light v analysis incorporates aspects of previous analyses of bare and gei purposives while providing a simpler explanation for relevant phenomena. Additionally, the proposal bolsters the argument that gei in the post-object gei phrase in various constructions is a verb (e.g., Lin & Huang 2015). The results of an empirical survey testing speaker judgments of sentences with lai and its antonym qu 'go' confirm the prediction that lai is a light verb that can appear in possible light v positions across sentences. Some differences in acceptance rates and preferences for lai or qu may be due to the bleached semantic content of lai and qu, frequency of usage, and processing difficulties arising from reference ambiguity.
In Eurasia, the hotbed of classifier languages as an areal feature is in East Asia, Southeast Asia, part of eastern India, and a less concentrated cluster in the Middle East and Central Asia. In Europe, only a few languages have been claimed to employ numeral classifiers. Russian, located in Eastern Europe, has a tripartite numeral construction [Num X N], where X's are considered sortal classifiers by some researchers, e.g., Sussex (1976) and Goto (2012), and is listed as a classifier language in the World Atlas of Classifier Languages (WACL). However, the status of Russian as a classifier language is controversial. In this paper, we first adopt the definition of classifier languages proposed by Her, Hammarstr & ouml;m & Allassonni & egrave;re-Tang (2022). We then apply explicit syntactic criteria to evaluate Russian numeral constructions and examine the tripartite [Num X N] structure. Ultimately, our analysis demonstrates that the elements previously identified as sortal classifiers in Russian are in fact measure nouns, and we reject their classification after examining the reasons behind this misidentification. The findings suggest that it would be prudent to reexamine all putative classifier languages in Europe.
This study explores methods to enhance the performance of offline Large Language Models (LLMs) using generative question-answer (QA) pairs. Existing research highlights the effectiveness of example-based prompts and QA pairs in improving LLM robustness and contextual understanding (Takahashi et al. 2023, Chowdhury & Chadha 2024). However, generating domain-specific QA pairs remains challenging due to the scarcity of datasets across diverse industrial sectors. To address this issue, we advance an innovative and adaptive approach that employs Generative Grammar (Chomsky 1957 et seq.) to convert industry-specific statements into questions, thereby facilitating QA pair creation. We compare the efficacy of this method with that of LLM-generated QA pairs. Our proposed approach not only reduces the labor-intensive process typically associated with prompt engineering but also provides a transparent and systematic framework for question generation through controlled wh-movement transformations. Initial findings indicate that QA pairs generated via these transformational rules substantially enhance LLM performance in industrial chatbot applications by enriching contextual information and highlighting promising directions for future LLM research and downstream applications.
This paper provides a comprehensive analysis of non-canonical wh -questions from a cross-linguistic perspective. This study claims that they are encoded through the interaction of various functional elements, prosodic constraints and pragmatic construals. Our investigation reveals that they combine a wh -expression with a modal element that can be either lexical or silent. Prosodically, they are characterized by distinct pitch and stress patterns. In terms of semantics, these constructions involve negation over modal quantification in conjunction with various not-at-issue contents such as expectations and presuppositions. Pragmatically, they change information-seeking into a denial/disapproval act, typically raising objections to the at-issue content within the scope of sentential wh -adverbs. Finally, we suggest that the origin of non-canonical wh -questions may well trace back to the hierarchical arrangement of causal and source questions: Namely, they are disrupted when the speaker is no longer interested in the cause-effect relationship, holding instead a negative attitude towards the interlocutor’s remarks or behavior. Our approach thus sheds new light on the complex nature of non-canonical wh -questions in relation to their interrogative counterparts.
Alongside superiority shi shenme 'what is superiority?,' which displays a canonical word order, the order shenme shi superiority is also possible in Chinese. In this paper I argue that this optionality results from a derivation that starts with merger of the two heads shenme 'what' and superiority, which generates a head-head structure that cannot be labeled under Chomsky's labeling algorithm. The system must therefore resort to movement of either shenme or superiority in order for the structure to be labeled. Evidence for this analysis comes from the fact that when the sentence contains phrasal elements like shenme yisi 'what meaning' or shenme ren 'what person,' the optionality disappears. I also examine examples involving other interrogative items in the language and claim that this optionality only obtains with shenme 'what' and shei 'who,' other superficially similar word order patterns being different in terms of their semantic properties. Finally, I discuss the implications of this analysis in the context of symmetry-breaking in the sense of Moro (2000) and of Chomsky's (2013) labeling algorithm.
The wh-expressions shei 'who' and shenme 'what' in Mandarin Chinese not only convey an interrogative meaning but also exhibit existential and universal readings in specific contexts (Huang 1982, Cheng 1991, 1995, Li 1992, Tsai 1994, Lin 1996, 1998). Focusing on the universal interpretation of shei, this paper has three objectives. First, we demonstrate that current state-of-the-art large language models (LLMs) such as ChatGPT, lack reliability in distinguishing these three distinct readings of shei. Second, we develop a specialized natural language processing and understanding (NLP/NLU) system capable of processing and interpreting shei across diverse contexts with greater accuracy, transparency, and consistency. Unlike current LLMs, our system is built upon Wang et al.'s (2019a, 2019b) generative linguistics-based NLP/NLU software tools, Articut and Loki, enabling it to require significantly less training data to interpret the universal reading of shei. Third, we compare our model's performance with that of ChatGPT, demonstrating its superior accuracy and robustness in interpreting the universal reading of shei.
In comparison to the extensive body of research on canonical expressions of homogeneous plurality, such as Mandarin bare nouns and the English plural morpheme -s, the study of similative plurality in natural language has received relatively little attention in the current literature (with notable exceptions such as Smith 2020a, 2020b). This study addresses this gap by contributing to a typological understanding of similative plurality through an analysis of the Mandarin expression shenme de, which conveys verbal similative plurality. Notably, in contrast to Japanese -tari, shenme de gives rise to two intriguing puzzles: it consistently yields an inclusive interpretation in both upward-entailing and downward-entailing contexts (the monotonicity puzzle) and resists overt contextual restriction (the domain restriction puzzle). Drawing on insights from Smith (2020a, 2020b), we propose a fully inclusive mixture analysis of shenme de and demonstrate how this framework accounts for both puzzles. Finally, based on empirical observations of shenme de, we propose three parameters to typologically characterize similative plurality in natural language: (a) the Category Parameter, (b) the Host Parameter, and (c) the Domain Argument Parameter.
Unlike earlier studies addressing rhetorical questions in which wh-arguments appear, this work investigates the restrictions imposed on the rhetorical use of Mandarin sentences containing the wh-adjunct, shenme-shihou 'what-time.' Sometimes, sentences that contain this wh-phrase can be used as rhetorical questions carrying the refutatory force, while at other times they cannot. I propose to account for this phenomenon by examining the interaction between shenme-shihou 'what-time,' modals and the sentence-final particle le, arguing that sentences in which shenme-shihou 'what-time' appears can be interpreted as rhetorical questions only when (a) shenme-shihou is a wh-adverb adjoining to a maximal projection that denotes a change of state or expresses an inchoative reading, and (b) the wh-phrase shenme-shihou is not deeply embedded within two bounding nodes. As long as these two conditions are satisfied, a rhetorical interpretation of a shenme-shihou 'what-time' sentence can surface. This study not only helps us better understand the mechanism underlying both the interrogative and rhetorical uses of shenme-shihou 'what-time' sentences, but also shows that syntax plays an important role in constructing rhetorical questions in Mandarin Chinese.
This paper observes two disjunction markers in Mandarin Chinese: haishi and huoshi. The aim of this study is to investigate the grammatical distinctions between them. Although both disjunction markers convey logical disjunction meaning or, haishi is primarily associated with interrogative contexts, such as alternative questions, and encodes exclusivity, whereas huoshi is primarily found in declarative sentences, allowing inclusive interpretations. Consequently, I propose that haishi functions as an exclusive disjunction marker, imposing mutual exclusivity on alternatives, while huoshi serves as an inclusive disjunction marker, allowing overlap among alternatives. This proposal thus provides a systematic analysis that accounts for their patterns of interchangeability and non-interchangeability by examining their distributions in the following linguistic environments: alternative and polar questions, embedded clauses of know-predicate, cleft constructions, and downward-entailing contexts.
This paper provides a comprehensive analysis of non-canonical wh-questions from a cross-linguistic perspective. This study claims that they are encoded through the interaction of various functional elements, prosodic constraints and pragmatic construals. Our investigation reveals that they combine a wh-expression with a modal element that can be either lexical or silent. Prosodically, they are characterized by distinct pitch and stress patterns. In terms of semantics, these constructions involve negation over modal quantification in conjunction with various not-at-issue contents such as expectations and presuppositions. Pragmatically, they change information-seeking into a denial/disapproval act, typically raising objections to the at-issue content within the scope of sentential wh-adverbs. Finally, we suggest that the origin of non-canonical wh-questions may well trace back to the hierarchical arrangement of causal and source questions: Namely, they are disrupted when the speaker is no longer interested in the cause-effect relationship, holding instead a negative attitude towards the interlocutor's remarks or behavior. Our approach thus sheds new light on the complex nature of non-canonical wh-questions in relation to their interrogative counterparts.
This study examines whether knowledge encoded in evidential categories is perceived differently by readers with varying levels of certainty. A questionnaire with a five-point scale was conducted to collect readers’ judgments about events introduced by evidential markers in Chinese. The evidential markers were manually categorized into categories, and inter-rater agreement was calculated. The findings reveal two hierarchies in terms of readers’ degree of certainty in judging the veracity of the knowledge: visual > hearsay and inference > quotative > hearsay. Additionally, three factors were observed to influence readers’ judgments about events: (a) whether a detailed event depiction is provided, (b) whether knowledge of an event is provided by more than one source, and (c) whether knowledge of an event has a close relationship with readers’ political standpoints or personal societal beliefs.
Extensive research has explored tonal contrasts, dialectal differences, and sandhi patterns of Taiwanese Southern Min tones. However, the duration reflexes of these tones, which hold theoretical significance, and their potential variations between the older generation, who use Taiwanese Southern Min as a first language, and the younger generation, who use it as a “heritage” or second language, have received comparatively less attention. In a sizable corpus study, we demonstrated that Taiwanese Southern Min syllables are best described as bimoraic, akin to other Chinese dialects, where syllables with fewer segments have comparable durations to those with more segments, and vowel durations in simpler syllable structures are longer than those in more complex structures. Furthermore, lexical tone durations partially follow patterns observed cross-linguistically, with rising tones produced longer than falling tones, and tones with higher F 0 produced shorter than those with lower F 0 . Finally, we observed a much narrower tonal space in young speakers compared to old speakers, with less separation in their F 0 trajectories in both level and contour tone productions. Our comprehensive study delves into underexplored Taiwanese Southern Min tonal duration reflexes, shedding light on potential generational variations and contributing to a better understanding of sound change trajectories.
This paper aims to investigate an array of morphosyntactic properties that constrain the formation of externally-headed relative clauses (EHRCs) in Siwkolan Amis, an Austronesian language in Taiwan. Additionally, it seeks to address two key issues related to the syntax of relative clauses: connectivity and modification. First, I adopt the head raising analysis, also known as the Ā-extraction analysis, in light of island effects and idiomatic discontinuity. This analysis is very much in line with the subject-only restriction ( Keenan & Comrie 1977 ) and the Austronesian Extraction Restriction hypothesis ( Erlewine, Levin & van Urk 2017 ). Second, I argue that the head noun is structurally integrated into the EHRC through complementation. This structural relation is formed by the linker a , which behaves similarly to a complementizer selecting TP as its complement, suggesting that the EHRC has a full-fledged CP structure. A structural analysis of Amis EHRCs is proposed to account for these associated properties and has implications for the syntax of modification in Amis. First, Amis modifier phrases consisting of ma -inflected verbs can be analyzed on a par with Amis EHRCs. Second, the subject-only restriction can be recast as saying that the head noun at [Spec, v P] is structurally privileged.
Lexical borrowings in Formosan languages from Japanese and Sinitic languages are frequently discussed in linguistic literature. Borrowings between Formosan languages themselves are more difficult to disentangle. This paper presents loanwords from Formosan languages in four Atayal dialects: Matu’uwal, Matu’aw, Plngawan, and Klesan . All four dialects were found to have been in contact with their immediate neighbors. The donor languages range from distantly related Pazih and Saisiyat (for Matu’uwal), to closely related Seediq (for Plngawan), to other Atayal dialects (for Matu’aw and Klesan). Knowledge of lexical borrowings is useful when reconstructing protolanguages. It also helps with understanding the cultural history of the Atayal people and their relationships with their neighbors.
This study analyzes cultural metaphors exhibited by a repository of plant-oriented proverbs in Hakka, explicitly demonstrating the intricate interplay between the universal framework of The Great Chain of Being theory and the parametric constraints of contextual factors. The salient selection of various types of plants, along with their biological traits together with factors including linguistic contexts, physical and social settings, and cultural aspects gives rise to their linguistic manifestations and conceptualization patterns. The reified virtues overlap with philosophical representations of Confucian ethics and are further consolidated into dealing with work and dealing with people. While most are universal values, some are more culturally specific, indicating a dynamic spectrum of the core virtues that may exhibit different significances in different cultures. The investigation makes empirical and theoretical contributions to metaphor and proverb studies by enhancing the explanatory power of the Great Chain of Being theory and by illustrating how various facets of contextual factors influence the conceptualization and interpretation of the cultural metaphors of plant proverbs.
Abstract Word stress is a structural property of increasing prominence. An established line of scholarship regarding word stress exists both in terms of theory and description in the Lhasa Tibetan (LT) language. Unlike LT, no such scholarly works are available that focus on Dharamshala Tibetan (DT), a dialectal variety spoken by Tibetan refugees living in the Dharamshala area in Himachal Pradesh, India. The current work aims to provide a systematic and concise theorisation of DT word stress based on the data collected from field in terms of parameters like culminativity, location of the head, direction, and quantity sensitivity. Optimality Theory is used to offer a theoretical judgment behind the analysis. A majority of DT words contain a trochaic, weight-insensitive, left-to-right stress pattern. The degenerate foot is accepted. Very few instances of words with an iambic stress pattern were found during the fieldwork. Similarly, few words containing heavy syllables are available in the word stress pattern inventory of DT.
The present study attempts to unveil patterns in interactive information structure of clause enhancement that conventionalize the actual representation of circumstantial elements which add information about time, place, manner, means, and reason/cause through circumstantial augmentation, tactic augmentation, or connectives in method sections of research articles (RAs). The dataset consisted of 120 method sections of randomly selected empirical RAs from ISI-indexed Q1-ranked applied linguistic journals published between 2020 and 2022. The results of the study point to a significant distinction in hypotactic augmentation to register circumstantial information in comparison to other choices available to authors. Moreover, the most striking observation to emerge from the data is that a logical information structure of circumstantial meaning is mostly facilitated through hypotactic non-finite enhancement rather than its finite counterpart. This viable preference lies in elliptical, evincing, and expressive functions of hypotactic clauses that make comprehension more attainable to the readers and assist authors in fulfilling the generic conventions and communicative purposes of academic writing and managing interactive information structures. A major theoretical implication of the current research entails the significance of the interactive functions of the options available to authors which can impose priorities on the system of choices within the context of information.