
This article examines data from a Piedmontese translation of the Gospels that throw new light on the development of presentational constructions involv ing an etymologically locative clitic. It builds on the evolution proposed in Parry (2013), namely that presentational or broad-focus VS constructions in Piedmontese, and some other northern Italo-Romance varieties, with the clitic i/-je ‘(to) there’ developed from locative sentences by means of a gradual reanal ysis of the anaphoric referential locative clitic as a non-argumental verbal mark er of agreement with an implicit locative Subject of Predication. Alongside its original locative meaning, i/-je in sentences with verb-subject inversion came to acquire a pragmatic-discourse function, namely the presentation of a new entity or event. Data from the 19th-century Testament Neuv dë Nossëgnour Gesu-Crist tradout in lingua piemonteisa (‘The New Testament of our Lord Jesus Christ trans lated into the Piedmontese language’) suggest a significant intermediary stage in the grammaticalization process from locative argument to presentational marker that may also apply to other Romance varieties with similar presentational con structions. It is proposed that the originally deictic clitic i/-je was used to attract the attention of the hearer/reader to a new appearance, hence new information, first in an objective, visual sense and then metaphorically. This is the essential function of the mirative.
This paper investigates the discourse-pragmatic expletive lu in Sovramontino, a conservative Gallo-Italic variety spoken in the northeastern part of Italy. Building on cross-linguistic research, the data suggest that lu functions as a dedicated marker of verum focus to emphasise the truth value of a proposition. Lu surfaces in clause-final position in contexts lacking a referential subject, such as weather verbs, impersonal and presentational constructions, and existential clauses. In contrast, in transitive and unergative clauses, verum focus is instead marked by a prosodically prominent pronominal copy of the subject in clause final position. The analysis seems to support the Lexical Operator Thesis on verum focus, positing that lu is a semantically vacuous lexical operator inserted to encode polarity emphasis. The paper proposes a syntactic account involving polarity fronting and TP-remnant movement, locating lu in the left periphery of the clause. The investigation contributes to our understanding of how null subject languages encode discourse-pragmatic information structurally; more specifically, it raises the broader questions about the interaction between sub jecthood and verum focus, offering insights into the typology of discourse-related expletives across Romance and beyond.
This study examines the morphosyntactic properties and the pragmatic func tions of talè in Sicilian. Its morphological invariability and its restricted syntactic distribution are taken as evidence for its status as discourse marker. Despite its uncertain origins, this marker is clearly related to the verb taliari ‘look’ and it shows a functional behaviour similar to the verb-based discourse markers equivalent to English look in other Romance languages, sharing in particular the function as attention-getter (cf. Italian guarda, Spanish mira). Unlike its Romance counterparts, however, talè has additionally developed a mirative value, which may complement its function as attention-getter or serve as its sole expressive and subjective meaning. This prominent and frequent mirative function, which only appears in limited or constrained contexts in other Romance languages, suggests a higher degree of grammaticalization and functional specialization for talè, and reflects more advanced stages in the evolution of pragmatic functions that characterize Romance verb-based discourse markers.
In this paper we discuss the function of topics and the emergence of expletives in thetic VS constructions from a corpus of early Romance texts of northern Italy. The sources date from the 13th to the early 16th centuries and are thus characterized by a ‘verb second’ (V2) syntax. We observe that aboutness and referential topics are distributed differently in the medieval vernaculars of northern Italy, as they are coded in pre- and post-verbal position, respectively. We also note that thetic VS structures without an overt aboutness topic exhibit an expletive pronominal form in preverbal position and lack V-S agreement. We propose that the expletive spells out the implicit, context-dependent topic that thetic VS sentences presuppose, signalling the contextually understood spatio temporal coordinates of the event or situation presented by the thetic construc tion. The evidence drawn from the written domain of the historical data suggests that there cannot be topicless sentences, as has been observed for the contempo rary Romance varieties spoken in northern Italy.
This paper investigates the interplay between information structure and ges ture from a generative perspective, focusing on a case study involving the ges tural topic marker [FINGER-BUNCH-OPEN-HAND] ([FBO]) in southern Italo Romance. This conventionalised gesture has previously been reported as mark ing the topic of utterances in non-formal linguistics (Kendon 1995). However, its precise grammatical contribution to the utterance remains unclear. Assuming the ‘Grammatical Integration Hypothesis’ (Colasanti 2023b) – which holds that any gesture assigned a semantic denotation and a phonological representation is the output of a syntactic representation – this paper shows that [FBO] behaves like spoken topic markers (e.g. Japanese wa), exhibiting a similar syntactic dis tribution but externalised in the visual-gestural modality instead. More gener ally, from a formal perspective, the relation between information structure and gesture is primarily a matter of syntax-phonology interface; it is by no means special, but rather parallels what is found in the spoken modality.
This paper investigates the syntax and semantics of the expression e anca anca – literally ‘and also also’ – attested in various Venetan dialects spoken in North Eastern Italy. Treated as a fixed phrase, it consists of the conjunction e (‘and’) and a reduplicated form of the additive adverb anca (‘also’). Depending on the context, it conveys either an emphatic increase (‘and even more’) or decrease (‘and even less’). We argue that e anca anca expresses the speaker’s epistemic attitude and functions as an instance of even-focus, a type of scalar, focus-sensitive expression. Our proposal is that e anca anca has undergone grammaticalization and behaves as a discourse marker. By ‘grammaticalized’ we mean, for instance, that the con junction e no longer functions as a canonical coordinator, and the reduplication of anca is not interpreted compositionally. Instead, the expression occupies a dedi cated position in the left periphery, specifically associated with corrective focus.
Sentences that describe existence and presence in a context and sentences that introduce a new event into discourse are known to share formal features in many languages and have been argued to be indistinguishable. Availing our selves of authentic corpus evidence and of the findings of interviews with native speakers, in this article we analyse such constructions in Kréol Rényoné, a French-based creole spoken on Reunion Island. We find that despite the shared word order and the marking with a form of the verb ‘have’ (nana), existentials, or descriptions of presence in a context, and cleft presentationals, which intro duce new events, differ in syntactic and semantic terms. We disentangle and analyse their meaning and syntax and we capture their intersection. The shared marking follows, in our analysis, from the way the two constructions interface with discourse: the propositions that they express are interpreted and evaluated in terms of the deictic coordinates present in the common ground when the sen tence is uttered.
For the morphological model proposed in this paper, we study Romance varieties spoken within the French department of Alpes-Maritimes. The analysis draws upon data which come from transcribed oral interviews. These interviews are in the form of questionnaires that each informant translates from French into his or her own local dialect of Occitan. Using data from nineteen different survey points, the linear orders of unstressed object pronouns are compared based on grammatical case. Three regions are identified within Alpes-Maritimes according to linear orders of pronominal sequences (ACC + DAT, DAT + ACC, and a variable linear order). Within the framework of Optimality Theory (Prince & Smolensky 1993), a hierarchical morphological model is used along with alignment constraints based on Case to account for these pronominal linear orders, some of which are otherwise unexpected using the morphological model alone.
This reply makes two main points. First, it lays out why very large language models [vLLMs], although useful as tools in linguistics, cannot be compared to linguistic theories. Linguistics as a science is necessary to understand the workings of human language, and to gain insights into the cognitive properties of human knowledge pertaining to language. Second, the reply clarifies the spectrum of generative grammar and points to major achievements in the field of theoretical linguistics.
Chesi's position on the relation between theoretical linguistics and large language models (LLMs) should be taken seriously by anybody who has worked on these two ways of approaching language. This commentary largely shares his general viewpoints, with some qualifications. One is that standardization (and the search for a broad core of compatible analyses) is more important than formalization. The second is that the size and opacity of the materials used for training the largest LLMs makes them essentially useless as models of human linguistic competence (though perhaps not entirely implausible as models of the evolution of a creature capable of using language). On the other hand, smaller models, trained on human-sized amounts of language (as in the the BabyLM challenge) could become useful tools to study human competence and understand the limits of exclusively data-driven approaches.
Chesi suggests that disagreements among leading researchers indicate that generative grammar is failing, and he proposes that large language models may provide a better kind of theory, one that rejects modularity in favor of using all aspects of surface linguistic context in every prediction. But the argument is not persuasive. Widely accepted and ongoing empirical advances suggest that alternative generative linguists agree on much more than recent programmatic disputes might suggest, and formal studies confirm this. While language models have significant instrumental, surface predictive power, as competent speakers do, these do not disconfirm claims of generative grammar or provide any alternative explanation of what human language is and why it has the properties it does.
We critically evaluate Piantadosi's claim that deep neural networks have obsoleted linguistic theory. The Generative Enterprise seeks to explain why human language exhibits discrete infinity, yet is not unrestricted. In fact, all languages appear to obey the same basic underlying properties. Through the lens of the Strong Minimalist Thesis (SMT), driven by evolutionary considerations, inquiry has been focused on maximally simple operations such as Merge for structure, and Minimal Search for establishing structural relations. By contrast, the computationally expensive setting of billions of parameters in current deep neural networks perform provides no biologically plausible explanation for human language. Moreover, we show through simple examples that the performance of current systems turn out to be highly overrated.
This paper responds to Chesi's paper Is it the end of (generative) linguistics as we know it? from the perspective of a computational linguist who works within one of several generative syntactic frameworks, namely Lexical Functional Grammar. Chesi's general conclusions and recommendations to Minimalism for the best way forward are found to hold, yet the diagnosis of the precise causes and therefore also the suggestion of remedies differs. Overall, I propose that the chances and opportunities offered by advances in machine learning in general and Large Language Models in particular should be studied carefully and integrated into a model of language which combines a rule-and-symbol driven approach with gradient and probabilistic information.
This commentary starts by discussing the future of theoretical linguistics in the context the rise of Large Language Models (LLMs). While comparative linguistics will likely persist due to universal interest in linguistic diversity, it is indeed questionable whether Chomskyan generative linguistics will continue as before. Its key contributions are mid-level generalizations, but its claims about innate linguistic knowledge have not been supported. One problem with Chesi's article is his focus on computational efficiency as this diverges from the Chomskyan goals. The commentary lists six core theoretical goals of linguistics and relates Chesi's article to these six goals. In particular, the conflation of theories (language-specific descriptions) and frameworks (general tools) has often created confusion. Ultimately, the text concludes that generative linguistics' decline does not doom the discipline.
This note intends to stress two complementary aspects of the issues raised by Cristiano Chesi's provocative paper: (i) Generative grammar and Large Language Models are two separate scientific endeavors, with different goals and methodologies: the first aims at the scientific description and explanation of a natural object, the human language faculty; the second is a technological program aiming at expressing linguistic knowledge in machines, in view of an efficient man-machine interaction. They should be kept carefully distinct. As far as I can tell, the second cannot determine the end of the first, much as the technological discovery of airplanes did not determine the end of the scientific study of flight in nature. (ii) The two endeavors both deal with the same object, natural language, and have common roots in the theory of computation (the common use of the adjective 'generative' in generative grammar and in generative artificial intelligence presumably is not a mere lexical accident). Rather than being considered in competition, they should be thought of as complementary in many ways. Various forms of collaborations should be envisaged in the future.
In this reply to Chesi's Is it the end of (generative) linguistics as we know it, I argue that the specifics of his vision for generative syntax in the 21st century remain hazy. Depending on how one interprets Chesi's methodological desiderata, they may well have a chilling effect on novel approaches instead of fostering them. As a concrete example of this dynamic, I discuss the problems with Chesi's focus on benchmarks and Minimum Description Length and how it would undermine recent efforts in subregular syntax that are in fact closely aligned with Chesi's goals. I conclude that Chesi's vision has merit, but only in moderation.
Chesi (this issue), responding to Piantadosi (2024), compares Large Language Models with work in Generative Linguistics. Chesi points out that syntactic tests exist for Large Language Models, but not for theories of Generative Linguistics, as well as that theories in Generative Linguistics lack proper formalization, which has led to Generative Linguistics becoming marginalized. In this paper, I take the position that Generative Linguistics develops theories of language, but Large Language Models are not theories. Also, while syntactic tests that examine the validity and scope of theories in Generative Linguistics could be useful, a number of large hurdles exist for their development.