
This study proposes a typology of referents designated by the second-person generic pronoun (DPG) based on a corpus of spoken French (OFROM). Far from being reducible to a unified usage, DPG emerges within diverse enunciative and interactional configurations, reflecting various discursive strategies of generalization, projection, and intersubjective negotiation. We identify three main configurations: (i) generalization of the speaker’s experience, which may take the form of shared knowledge co-construction, (ii) symmetrical projection into an unfamiliar situation, and (iii) hypothetical experimentation. The analysis shows that DPG isn’t a mere variant of on, but functions as a specific enunciative operator, modulating speaker-hearer relations in dialogic contexts in different ways. Semantically, we adopt Benveniste’s view of the second person as a “non-I”, whose under-specification fosters referential plasticity. By combining morphosyntactic, referential, and pragmatic approaches, this study sheds light on the linguistic, cognitive, and interactional mechanisms involved in the construction of genericity in spoken French. DPG thus emerges as a privileged linguistic device for discursively constructing situated norms and shared knowledge, as well as fostering intersubjective involvement.
The focus of this study is on pronouns with the [+hum] feature and a propositional base – qui tu sais, tu sais qui, qui vous savez, vous savez qui – which have been largely overlooked in research on French. Based on a corpus, this study analyzes the syntactic, semantic, and discursive functioning of these sequences, belonging to a large paradigm of forms referring to humans as well as original modes of co-constructing reference.
This article offers an in-depth analysis of the different behaviours of a pseudo-cleft construction such as exemplified by Ce qui est un problème, c’est X (What is a problem is X). This type of structure combines features from both (i) pseudo-clefts (What is problematic is X) and (ii) specificationals (The problem is X). Indeed this hybrid construction also contains a shell noun in its left segment and the specification of said shell noun in its right segment, thus exhibiting the same underlying structure. Given this interpretation based on the saturation of a role by a given value, one could argue that it summons similar cognitive schemas as specificational constructions. Our research is based on 247 occurrences from the frTenTen corpus, 2023 (23 billion words). It shows, among other things, that this construction is linked to textual rather than syntactic “constraints”. Furthermore, shell nouns in this construction seem to depend heavily on previous discourse, which highlights the many relations between that shell noun and its left context, be it extraction, reorientation or elaboration.
This paper investigates the discursive and pragmatic functions of it-clefts in contemporary English. The qualitative corpus-based study is based on the hypothesis that an it-cleft originates from the extraposition of a complex noun phrase in an underlying initial pseudo- or TH-cleft structure. From this perspective, it is shown that it-clefting pragmatically displays both specific properties and features shared with a related non-canonical construction, namely extraposition. The analysis demonstrates that it-clefts are not mere syntactic or stylistic devices, but that their use is motivated by pragmatic considerations and the discourse situation. Although it-cleft formation deviates from a canonical pattern and is therefore marked, it-clefts themselves can be either marked or unmarked. This study proposes a unified account of their discourse functions, arguing that their primary communicative function is to identify or specify an argument or circumstance, and that their uses parallel those of extraposed subject-subordinate clauses. It further provides evidence for a gradient ranging from the most prototypical to the least typical it-clefts, which supports the view that they form a unified discursive category, with differences residing only in co(n)textual factors and pragmatic effects.
This article aims to determine the nature of language processing units in written production, which can be defined as patterns used spontaneously during the writing process. Due to their internal solidarity, constructions (i.e., form-meaning pairings) are considered in this article as relevant processing units candidates. Our study focuses more specifically on the way French (semi-)copular constructions (e.g., Jean est devenu/reste tres autonome "John has become/remains very independent") are produced during the writing process: that is, based on a corpus of online recorded texts, we examine the occurrence of long pauses in relation to these constructions. The results of our research show that there is a high degree of internal solidarity within (semi-)copular constructions and that, despite the fact that there are some disparities in our data which are typical of real-time written production data, these outcomes support the hypothesis of a constructional blocs-based cognitive treatment during writing.
This article focuses on a category of statements in which a subject noun phrase conveys a viewpoint, a stance, or an ideology that prepares, supports, or is more important than the asserted content. We describe the argumentative and semantic relationships that emerge in discourse between the presupposed content and the asserted content. The examples we analyze are taken from Twitter/X. When the author of a tweet says, "l'Etat ne fait RIEN pour vous proteger" it is meaningful that they use the noun phrase l'Etat as a subject. Indeed, they mean that the protection the State fails to provide would fall under its responsibility - in other words, that the State is failing in fulflling its obligations as a State. This, we believe, is due to the meaning of the word Etat, whose use in a subject noun phrase triggers a presupposition that is argumentatively linked to the asserted content. We aim to account for this phenomenon within the framework of argumentative semantics developed by Anscombre and Ducrot (1983). More specifically, we draw on the "theory of semantic blocks" developed by Carel (2011,2023)
Our study investigates the way French speakers acquire additive expressions, especially the scope particles anche/pure ("aussi"), ancora ("encore") and the temporal adverbs sempre ("toujours"), ancora ("encore"), in Italian as a foreign language when building discourse cohesion in a complex narrative task. Three groups of learners with different levels have been compared to two reference groups (Italian and French native speakers). Our framework is functional and discursive (Klein & Stutterheim, 1987 et 1991). The data has been collected via an illustrated story with no written text created by Dimroth (2002). Our research aims at: (a) comparing the additive means (additive scope particles and possible alternative means) exploited by native speakers to the additive expedients employed by the learners at differentinterlanguage stages; (b) detecting the role of transfer; (c) identifying the "perspective" that the informants adopt in L1 and L2 (Thinking for Speaking Hypothesis-Slobin, 1996). The study points out some conceptual and formal differences between French and Italian with respect to the choice of additive means, and the learners' resort to syntactic transfer and to a type of transfer that we defined as "of frequency", as well as to some general interlanguage strategies (Klein & Perdue, 1992; Giacalone Ramat, 2003).
We examine the semantic evolution of Donald Trump's presidential speeches during his first term (2017-2021), using quantitative methods from corpus linguistics and natural language processing. Our corpus includes all transcribed official speeches, segmented into homogeneous speech units and then analyzed using word embeddings. Two scales of evolution are considered: an external dynamic, measured at the level of the presidency, and an internal dynamic, observed within each speech. The external analysis does not reveal a large semantic shift over time, suggesting relative stability in the presidential register. On the other hand, a clear internal evolution is evident: Trump's speeches begin in an institutional and impersonal register, then evolve towards a more subjective formulation, marked by the increasing use of the first person and a decrease in lexical density. This pattern indicates a specific structure of Trump's presidential discourse, characterized by a gradual transition from a representative role to personal expression.
The aim of our study is to contribute to the description and understanding of sentence organization in Spanish, taking into account the issue of dialectal variation and its formalization. More specifically, we aim to compare the probability of having a given structure, according to different dialectal varieties, and to quantify the weight of the factors involved in the choice of structure according to these varieties. To this end, we present a corpus study concerning subject position in Spanish yes-no questions: preverbal or postverbal subject (e.g., ¿Tomas tú café? vs. ¿Tú tomas café?). The data were extracted from the written part of the “Corpus del Español del Siglo XXI” (CORPES XXI), sampling 100 sentences per variety studied (Andes, West Indies, Continental Caribbean, Chile, Spain, Mexico/Central America, Río de la Plata), then manually annotated according to different linguistic categories (animacy, definiteness and length of the subject, transitivity of the verb, pragmatic value of the question). From a descriptive point of view, the position of the subject is correlated with all these variables. From a quantitative point of view, the use of conditional random forests shows that the most important variables for the choice of the subject position vary according to the seven dialectal varieties considered.
This article aims to determine the nature of language processing units in written production, which can be defined as patterns used spontaneously during the writing process. Due to their internal solidarity, constructions (i.e., form-meaning pairings) are considered in this article as relevant processing units candidates. Our study focuses more specifically on the way French (semi-)copular constructions (e.g., Jean est devenu/reste très autonome “John has become/remains very independent”) are produced during the writing process: that is, based on a corpus of online recorded texts, we examine the occurrence of long pauses in relation to these constructions. The results of our research show that there is a high degree of internal solidarity within (semi-)copular constructions and that, despite the fact that there are some disparities in our data which are typical of real-time written production data, these outcomes support the hypothesis of a constructional blocs-based cognitive treatment during writing.
This article focuses on a category of statements in which a subject noun phrase conveys a viewpoint, a stance, or an ideology that prepares, supports, or is more important than the asserted content. We describe the argumentative and semantic relationships that emerge in discourse between the presupposed content and the asserted content. The examples we analyze are taken from Twitter/X. When the author of a tweet says, “l’État ne fait RIEN pour vous protéger” it is meaningful that they use the noun phrase l’État as a subject. Indeed, they mean that the protection the State fails to provide would fall under its responsibility – in other words, that the State is failing in fulfilling its obligations as a State. This, we believe, is due to the meaning of the word État, whose use in a subject noun phrase triggers a presupposition that is argumentatively linked to the asserted content. We aim to account for this phenomenon within the framework of argumentative semantics developed by Anscombre and Ducrot (1983). More specifically, we draw on the “theory of semantic blocks” developed by Carel (2011, 2023).
This article offers a discussion and some elements on the characterization of the notion of restatement by focusing on the following question: in formal discourse approaches, should restatement be added as a discourse relation in its own or as a subtype of another relation? In order to provide some answers, we present the treatment of restatement in different formal approaches, and we examine the role of the discourse marker c’est-à-dire (“that is (to say)”) in recognizing it. We highlight a significant argument in favor of Segmented Discourse Representation Theory (SDRT): restatement arises from the relation of Elaboration, which is subordinating. Consequently, restatement acquires the status of a discourse strategy of subordination.
We examine the semantic evolution of Donald Trump’s presidential speeches during his first term (2017-2021), using quantitative methods from corpus linguistics and natural language processing. Our corpus includes all transcribed official speeches, segmented into homogeneous speech units and then analyzed using word embeddings. Two scales of evolution are considered: an external dynamic, measured at the level of the presidency, and an internal dynamic, observed within each speech. The external analysis does not reveal a large semantic shift over time, suggesting relative stability in the presidential register. On the other hand, a clear internal evolution is evident: Trump’s speeches begin in an institutional and impersonal register, then evolve towards a more subjective formulation, marked by the increasing use of the first person and a decrease in lexical density. This pattern indicates a specific structure of Trump’s presidential discourse, characterized by a gradual transition from a representative role to personal expression.
Our study investigates the way French speakers acquire additive expressions, especially the scope particles anche/pure (“aussi”), ancora (“encore”) and the temporal adverbs sempre (“toujours”), ancora (“encore”), in Italian as a foreign language when building discourse cohesion in a complex narrative task. Three groups of learners with different levels have been compared to two reference groups (Italian and French native speakers). Our framework is functional and discursive (Klein & Stutterheim, 1987 et 1991). The data has been collected via an illustrated story with no written text created by Dimroth (2002). Our research aims at: (a) comparing the additive means (additive scope particles and possible alternative means) exploited by native speakers to the additive expedients employed by the learners at different interlanguage stages; (b) detecting the role of transfer; (c) identifying the “perspective” that the informants adopt in L1 and L2 (Thinking for Speaking Hypothesis – Slobin, 1996). The study points out some conceptual and formal differences between French and Italian with respect to the choice of additive means, and the learners’ resort to syntactic transfer and to a type of transfer that we defined as “of frequency”, as well as to some general interlanguage strategies (Klein & Perdue, 1992; Giacalone Ramat, 2003).
This article offers a discussion and some elements on the characterization of the notion of restatement by focusing on the following question: in formal discourse approaches, should restatement be added as a discourse relation in its own or as a subtype of another relation? In order to provide some answers, we present the treatment of restatement in different formal approaches, and we examine the role of the discourse marker c'est-& agrave;-dire ("that is (to say)") in recognizing it. We highlight a significant argument in favor of Segmented Discourse Representation Theory (SDRT): restatement arises from the relation of Elaboration, which is subordinating. Consequently, restatement acquires the status ofa discourse strategy of subordination.
The aim of our study is to contribute to the description and understanding of sentence organization in Spanish, taking into account the issue of dialectal variation and its formalization. More specifically, we aim to compare the probability of having a given structure, according to different dialectal varieties, and to quantify the weight of the factors involved in the choice of structure according to these varieties. To this end, we present a corpus study concerning subject position in Spanish yes-no questions: preverbal or postverbal subject (e.g., & iquest;Tomas t & uacute; caf & eacute;? vs. & iquest;T & uacute;tomas caf & eacute;?). The data were extracted from the written part of the "Corpus del Espa & ntilde;ol del Siglo XXI" (CORPES XXI), sampling 100 sentences per variety studied (Andes, West Indies, Continental Caribbean, Chile, Spain, Mexico/Central America, Rio de la Plata), then manually annotated according to different linguistic categories (animacy, definiteness and length of the subject, transitivity of the verb, pragmatic value of the question). From a descriptive point of view, the position of the subject is correlated with all these variables. From a quantitative point of view, the use of conditional random forests shows that the most important variables for the choice of the subject position vary according to the seven dialectal varieties considered.
In literary dialogues, discourse markers are essential in reproducing the characteristics of spoken language, adding realism to interactions and the rhythm of the narrative. These elements reveal the characters'emotions, intentions, and social relationships. The aim of our study is to determine how these markers help replicate orality while maintaining the narrative integrity and readability of the written text. To explore this, we analysed the dialogues from the Norwegian novel Bienes historie by Maja Lunde and from its French translation Une histoire des abeilles. We focused on the markers da and jo, as they play a crucial role in structuring current oral interactions. Their frequent presence, particularly in informal dialogues, helps shape the rhythm and implicit agreement between speakers, making them key tools for our analysis. The analysis shows that da and jo structure discourse and enhance the fluidity of exchanges, but their French equivalents, such as donc for da and mais for jo, do not always fully capture their original functions. This complexity highlights the challenges of translating discourse markers, which are essential to faithfully represent human interactions and offer an immersive reading experience. Thus, the study of these markers sheds light on a new dimension of narration, where orality and writing combine to offer an authentic portrayal of human interactions.
This article seeks to explore the use of the word enfin in two oral corpora from a contrastive angle. First, it identifies and classifies the different values of enfin, then it highlights a significant distinction in the functions of enfin in each corpus, while carrying out quantitative and qualitative analyses. The aim is to understand how the degree of orality of oral productions influences the functions of enfin. The data come from two corpora: the former, from 49 minutes 6 seconds of recordings and transcriptions in the MPF corpus, consists of spontaneous speeches; the latter is made up of five speeches by Emmanuel Macron during his state visit abroad, totalling 22,369 words, and represents well-developed speeches. The results reveal that the diversity of functions of enfin as a discourse marker is much greater in the spontaneous oral productions, which present a high degree of orality. The degree of orality plays a crucial role in the variety of discourse marker functions: the more dynamic a corpus is, with frequent and unprepared interactions, the more diverse the functions it will present for discourse markers will be.
The French morpheme moi displays a variety of paraphrastic uses and functions, tying to what Blanche-Benveniste has referred to as discourse grammar. In this contribution, we offer to depict these uses as stemming from a unique discourse marker moi. To substantiate this claim, we have extracted all matching tokens from an oral corpus, CEFC-Gold, and we show that they indeed fulfill the criteria distinctive of this category. Moreover, they yield the properties that are expected from discourse markers, that is, the possibility to accumulate with other discourse markers, and an overall involvement in disfluencies and reformulations events. Next, we investigate the diachronic process of emergence of this marker, using historical corpus data from Frantext. Our main result is that the functional range of moi gradually emerges over time, starting from the 16th century up to the 20th. Additionally, this allows us to contribute the ongoing debate regarding the role of cooptation and pragmaticalization in the rise of discourse markers. While cooptation may trigger the first paraphrastic uses and feed a source context from which the marker may further develop, this development obeys the same patterns as those of grammaticalization, and more broadly of semantic expansion. Our study rather emphasizes the complex character of this diachronic process, which involve dynamically interrelating functions operating over different levels of the language structure.
In literary dialogues, discourse markers are essential in reproducing the characteristics of spoken language, adding realism to interactions and the rhythm of the narrative. These elements reveal the characters’ emotions, intentions, and social relationships. The aim of our study is to determine how these markers help replicate orality while maintaining the narrative integrity and readability of the written text. To explore this, we analysed the dialogues from the Norwegian novel Bienes historie by Maja Lunde and from its French translation Une histoire des abeilles. We focused on the markers da and jo, as they play a crucial role in structuring current oral interactions. Their frequent presence, particularly in informal dialogues, helps shape the rhythm and implicit agreement between speakers, making them key tools for our analysis.The analysis shows that da and jo structure discourse and enhance the fluidity of exchanges, but their French equivalents, such as donc for da and mais for jo, do not always fully capture their original functions. This complexity highlights the challenges of translating discourse markers, which are essential to faithfully represent human interactions and offer an immersive reading experience. Thus, the study of these markers sheds light on a new dimension of narration, where orality and writing combine to offer an authentic portrayal of human interactions.