
Abstract This study extends a paradigm-based framework for lexical conventionalization to multiword expressions (MWEs) in diachronic scientific English using the Royal Society Corpus (1665–1996). Building on prior work that quantified paradigmatic variability, the set of contextually available lexical alternatives, via distributional entropy, we model MWEs as single lexical units in distributional vector spaces to examine their role in long-term language standardization. Our aim is to verify whether MWEs undergo the same conventionalization process as the one previously observed for most word classes. Our results confirm that the overall decline in paradigmatic variability, indicating increasing conventionalization, continues through the 20th century. Different MWE classes follow distinct trajectories: compounds and phrasal verbs show substantial conventionalization in the 18th–19th and 19th–20th centuries, converging with or dropping below the variability of single nouns and verbs. In contrast, fixed expressions and academic formulas become conventionalized early and remain stably low thereafter. Flat expressions remain more variable, patterning with general nouns. These findings demonstrate that MWEs play an active role in the register-specific development of scientific English, with different types contributing differently to the reduction of paradigmatic choice.
Abstract The degree modifier practically , meaning ‘almost, nearly’ and functioning as a “approximator downtoner” (Quirk et al. 1985) or “approximator” in the “totality modifer class” (Paradis 2008), arises in Late Modern English, developing from an adverb of manner or respect meaning ‘in a practical manner, in practice, in reality’ in Early Modern English. This paper explores the historical development of practically. Modifying predicative and appositive adjectives and prepositional and adverbial phrases and occurring in noun and determiner phrases, practically comes to serve as a degree adverb. Modifying verbs and participles, it comes to function as a degree adjunct, often with negative implicature and emphasizer rather than degree function. Corpus findings point to the appearance of approximator uses of practically in all syntactic contexts in a confined period from 1830–1863 (cf. Núñez-Pertejo 2023). In ambiguous or “bridging” contexts, the manner meaning ‘in practice, in effect’ can be reinterpreted as ‘most often the case’ and hence as falling short of the expected level, thus giving rise to the approximator meaning ‘almost, nearly’. The change from a lexical adverb to a degree modifier is a process of grammaticalization, involving decategorialization (functional shift), host-class expansion, syntactic context expansion, desemanticization, semantic-pragmatic change, and subjectification. The later appearance of the degree modifier practically with bare verbs argues for a trajectory from degree adverb to degree adjunct rather than the reverse. With lexical adjectives, the approximator is grammaticalized first in predicative position and only later in attributive position, most likely because the predicative position is most similar to the “locus for reanalysis”, namely be/have + past participle (De Smet 2012). Finally, the negative implicature of the degree adjunct practically most likely develops because of the preponderance of verbs with negative semantic prosody, where ‘almost P’ is reinterpreted as ‘not (quite) P’, motivated by subjectification (Ziegeler 2015; 2016).
Abstract In the digital era, textbooks are increasingly required to be both adaptable to the digital environment and sensitive to genre conventions, posing new challenges for material selection. Although several corpora document authentic business communication, few are textbook-based and designed to serve as benchmarks for evaluating textbook content. To address this gap, the present study introduces the Corpus of Business English Textbooks (CBET), designed specifically to support corpus-informed compilation of digital Business English textbooks. The CBET is built from 69 widely adopted textbooks in Chinese tertiary education, comprising 7 genres and 35 sub-genres and amounting to 786,183 tokens. Beyond offering a comprehensive record of textbook discourse, the CBET enables systematic comparison between textbook and workplace texts, thereby providing empirical benchmarks for assessing the pedagogical appropriateness of instructional materials. To demonstrate its application, we employed Coh-Metrix to compare textbook letters from CBET with workplace letters, answering two questions: 1. How do business letters in the textbook scenario differ from those in the workplace? 2. What do these differences imply about CBET and its potential contribution to the selection of business letter materials for textbooks? Results reveal that textbook letters differ significantly from workplace letters in syntactic complexity, lexical concreteness, and cohesion, suggesting that they are more informationally compact and intentionally structured to develop learners’ command of abstract vocabulary and inference skills. These findings provide empirical insights for text selection in future digital Business English textbooks. Overall, this study highlights the CBET’s dual value as both a research resource and a practical tool for guiding the selection and digitalization of Business English teaching materials, thereby advancing corpus-informed pedagogy in the context of digital education.
Abstract Fiction set in the past is linguistically special, since representing the past in fiction typically requires making stylistic choices to convey a sense of ‘old-timey’-ness. Using the TV Corpus and the historical fiction dialogue it contains, which amounts to over 13 million words, this study investigates what makes the language of historical fiction distinct from other types of fiction on television. This is done by analyzing keywords, key parts of speech, and key semantic domains in the corpus. Results show that, compared to television generally, historical fiction greatly overuses modal verbs, terms of address, the perfect aspect and the passive voice, while it underuses features of orality, discourse markers, contracted forms, profanity and the progressive aspect. The paper argues that writers of historical fiction take advantage of the overlap of formal language and conservative (therefore older) speech to construct pseudo-historical dialogue. Importantly, the data implies that writers do pick up on real principles of language change and exploit this knowledge in their craft.
Abstract A recent survey of graph usage in corpus-based research has shown that the bar chart is the most widely used graph type for corpus data presentation. Motivated by this finding, the present paper offers a systematic review of bar chart usage in corpus-based research articles. It covers all papers ( n = 1,183) published in five corpus-linguistic journals up to and including the year 2024 ( International Journal of Corpus Linguistics , Corpus Linguistics and Linguistic Theory , Corpora , Research in Corpus Linguistics , and International Journal of Learner Corpus Research ). The aim of this survey is to arrive at a better understanding of the kinds of visualization tasks imposed on bar charts. We observe that they most commonly show percentages or absolute/normalized frequencies, and that they are often used for relatively complex visualization tasks involving multiple variables and subgroups. The survey is carried out against the backdrop of known limitations of this graph type and design recommendations found in data visualization guidebooks. Our critical examination of diagrams pays attention to issues that compromise the ability of the viewer to accurately perceive patterns in the data, and minor issues that affect the efficiency of a display. These observations are then distilled into a set of concrete recommendations, which are grounded in current usage and the advice given in the data visualization literature.
Abstract [ (If the) truth BE told ] is an idiomatic construction in English with a number of pragmatic purposes. It can suggest that a proposition is generally known but rarely admitted, or that a proposition is a previously unknown personal admission. It can also act as a pragmatically weaker discourse marker. This study initially looks for early uses in the Early English Books Online (EEBO) corpus, finding examples from the late 16th and early 17th centuries. Diachronic changes in the use and form of the construction in American English are then examined using the Corpus of Historical American (Davies 2010), which contains texts written between 1810 and 2020. In COHA the [ (if the) truth BE told ] construction becomes considerably more frequent after 1980. Coinciding with this increase in frequency the construction becomes more lexically fixed and reduced in length. It also becomes more likely to appear at the left periphery of a clause. In addition, it appears to be moving towards ‘extended intersubjectivity’ (Tantucci 2017). As such, it is increasingly losing its pragmatic purposes of marking a rarely admitted truth or a personal admission, and behaving more like a simple discourse marker, connecting clauses.
Abstract In this paper, I address the question of whether British and American web data pattern in the same way when it comes to attributive freak being used as an adjective or whether one variety is more tolerant towards the item crossing the nominal word class boundary than the other. In order to test this research question, I present a corpus-based analysis of 1,000 random tokens of attributive freak in the British and American sections of the Corpus of Global Web-based English (GloWbE). Contrary to my hypothesis that American English (AmE) favours noun to adjective transitions of freak over British English (BrE), the findings suggest that BrE is more open to freak crossing the nominal word class boundary than AmE. Further, the study reveals that nominal freak is almost restricted to the use of one to two frequent bigrams, both in BrE and AmE. In all other cases, freak is underdetermined for word-class status but leans towards an adjective interpretation.
Abstract Registers reflect the constraints of systematically recurring situational contexts and are therefore embedded in the lingua-culture in which these situations occur. Consequently, when a language – such as English – is used in widely differing cultural contexts, the question arises whether registers in different varieties of the language might not actually reflect cultural differences between similar types of situations. Previous studies have shown that varieties of English fall into different clusters and that informal spoken texts in particular reflect differences between the varieties. With a focus on register variation across varieties of English, Neumann & Evert (2021) suggest that register-related patterns of variation are much more pronounced than differences between varieties. However, they also observe divergence between texts in the same register from different varieties. The generality of both findings is limited, though, because their analysis was based on only three varieties of English. Our paper aims at exploring these questions more thoroughly by drawing on a larger set of nine components of the International Corpus of English (ICE) preprocessed for comparability (Lehmann & Schneider 2012) and by focusing the interpretation on registers that are expected to be more strongly affected by cultural differences. To this end, we extract the same set of 41 lexico-grammatical features from the ICE components as Neumann & Evert (2021), building on the corpus queries made available in their online supplement. In three steps, we first reproduce the geometric multivariate analysis (GMA) of Neumann & Evert (2021) and then replicate it in two increasingly different approaches. These methodological variations allow us to explore to what extent the results of Neumann & Evert (2021) depended on their specific choice of three ICE components and how stable the results of the exploratory analysis with the chosen multivariate approach are.
Abstract In the last few decades, much work in corpus linguistics has attempted to discover, and then interpret, differences in the frequencies of use of linguistic elements (words, patterns, constructions, discourse features, etc.). It is probably fair to say that such studies were particularly frequent in (i) learner corpus research, (ii) corpus-based varieties research, and (iii) sociolinguistically motivated studies. For instance, many studies have discussed the differences in how often certain elements are used (i) in corpus data from native speakers vs. corpus data from learner from different L1 backgrounds, (ii) in corpora representing different inner- and outer-circle varieties, or (iii) by speakers in corpora representing people of different gender or sexual identities. This paper will make the admittedly bold claim that any such study can in fact by definition unable to ‘prove’ what is often their main points, namely that the distributional differences found are in fact due to the one hypothesized explanatory variable(s) of L1, VARIETY, or, e.g., GENDER even when the distributional differences are significant and come with a decent effect size. To substantiate this claim, I will discuss some terminology from the family of methods known as multi-level modeling, namely the distinction between level-1, level-2, ... level-n variables and its relevance for many corpus studies. Second, I will then demonstrate how studies using only the above kinds of variables cannot distinguish the effect of their favored predictors from the effect of local/contextual level-1 variables. Third, in discussing this, I will exemplify how such effects need to be explored quantitatively instead.
Abstract Statistical approaches in linguistics seem to have gained in importance in recent times, especially in the field of Corpus Linguistics. In particular, the last ten years have seen an upsurge of linguists being dedicated to statistical methods and the improvement of statistical knowledge. This has repeatedly been described as ‘the quantitative turn’ in linguistics. In the present paper, we assess how real this quantitative turn actually is and whether statistics can be considered the ‘new normal’ in (corpus) linguistics. To this end, we have analyzed the contributions to six high-impact journals (Corpora, Corpus Linguistics and Linguistic Theory, ICAME Journal, English World-Wide, Journal of English Linguistics, and Language Variation and Change) for a period of eleven years (January 2011 until December 2021). Our results suggest that, indeed, statistical methods seem to be on the rise in linguistic studies. However, their frequency strongly varies between the journals, and, in general, we have identified some room for improvement in the use of advanced statistical methods, in particular the discussion of true prediction.
BackgroundCorpus linguistics, as a tool for research, is increasingly prevalent nowadays.It is relied on by researchers from diverse linguistics sub-disciplines, ranging from discourse analysis and sociolinguistics to computational linguistics, to name just a few.Designing and evaluating corpora that will be fit for the purpose they are being adopted for is imperative.This is particularly true since the need for robust and representative language corpora is well known.The authors of the book under review here, Jesse Egbert, Douglas Biber and Bethany Gray, have made significant contributions to the field of corpus linguistics, building on their research to advance the understanding and practice of corpus design and evaluation with a particular emphasis on representativeness, which dates back to Biber's work on corpus design in the 1990s (p.xii), honed over the years by means of their own extensive reading and research.In their introduction (p. 3) to Designing and evaluating language corpora: A practical framework for corpus representativeness, Egbert, Biber and Gray underline the fact that in definitions of the term 'corpus', 'representative' is, indeed, one of the most frequent key concepts.Corpus representativeness refers to the extent to which a corpus reflects the full range of variability in the language or language variety that it is intended to represent.It is important for a number of reasons.First, it allows researchers to generalize their findings from the corpus to the larger population of language users.Second, it ensures that corpora are useful for a wide range of research questions, from those that focus on specific linguistic features to those that focus on more general aspects of language use.Two key aspects of representativeness that are explored in this volume are domain and linguistic representativeness, which the authors see as being central to the design and evaluation of representative corpora.This volume, therefore, addresses the need, in particular, to focus on representativeness and does so with both academic rigour and practical insight.Corpus representativeness, in fact, is crucial for ensuring the reliability and generalizability of research outcomes.Achieving corpus representativeness, however, is not without its challenges.Several factors need to be considered, such as corpus size, composition, and balance.A larger corpus generally provides a more comprehensive and reliable representation of a language or language variety.However, the size of the corpus should be balanced with the available resources for data collection and analysis, together with the specific purpose for which that corpus is to be used.Careful sampling and consideration of
Turn-taking is a central topic in the theoretical framework of Conversation Analysis (CA).However, the analysis has mainly been based on British and American English, and differences can be expected if turn-taking is studied in varieties of English as suggested by anecdotal evidence.The present book, Conversation in World Englishes: Turn-taking and cultural variation in Southeast Asian and Caribbean English by Theresa Neumaier, sets out to fill this research gap by conducting an empirical analysis with the aim to describe, analyse, and compare turn-taking patterns in Caribbean and Southeast-Asian English face-to-face interactions.The major research questions are whether the turn-taking conventions in these two varieties correspond to those that have been established in previous work on turn-taking and whether there are differences between the varieties in terms of culture.The book contains eight chapters.Chapter 1 is the introduction where the author describes the aim of the book and states the research questions.In Chapter 2 the author provides an overview of Conversational Analysis (CA) and of the field of World Englishes catering for the needs of scholars in both traditions.Following Sacks et al.'s seminal article (1974), turn-taking is described as a two-part mechanism consisting of a turn-constructional and a turn-allocation component.The former deals with the construction of turn-constructional units (TUs) defined on the basis of syntactic completion and a variety of prosodic and lexico-semantic devices.The allocational component is split up into three hierarchically ordered rules describing how turn-taking is locally managed through techniques such as next-speaker selection, self-selection, and current speaker selection.Although there is some evidence that turn-taking strategies differ across languages, research on turn-holding and turn-claiming has only been carried out on few languages and "the question whether the turn-taking system might be culturally sensitive is still unanswered."(p.10).World Englishes are chosen as the ideal candidates to display if "the turn-taking system is fine-tuned to local cultural preferences" (p.12).If there are differences in the patterns between the varieties, they will be explained as culturally conditioned, and the patterns will be compared with regard to the cluster of context-sensitive features associated with each variety.One reason why CA and World Englishes have not been merged before has to do with "differences in their respective epistemologies" (p.15).The CA approach calls for 'theoretical indifference' based on the observation that "anybody who looks for differences hard enough will eventually find them in the data" (p.16).Combining "CA's fine-grained bottom-up analysis of interaction" with the study of World Englishes means that "only those contextual factors that show to be locally relevant for the interactants will be considered for the interpretation" (p.17).Studies of World Englishes, on the other hand, are based
Abstract This article investigates semantic prosody in a diachronic perspective. Although prosodies have been shown to change over time, there is no consensus regarding the source of such changes. The present study explores this further through a corpus study of the development of the lemmas fabric, fabricate and fabrication from the late 15th century to the late 20th century, drawing on material from Early English Books Online, the Corpus of Late Modern English Texts and the British National Corpus. The results of the study show that prosodic changes coincide with the emergence of new senses and indicate that these processes are related to and possibly caused by semantic transfer induced by persistent prosodies over time.
Abstract The modal verbs of necessity and obligation, a testing ground of grammatical change, have been shown to exhibit change and variation in world Englishes. Previous studies have primarily concentrated on English as a native language (ENL) and English as a second language (ESL) varieties. The present study extends this line of research and explores variation in modal verbs of necessity and obligation in English use as a Lingua Franca (ELF). Descriptive statistics indicate that ELF resembles American English and also shares similarities with ESL varieties. In addition, ELF further exhibits divergence from both ENL and ESL varieties that arises in multilingual interactions. The multivariate analysis of this study employs mixed-effects logistic regression on the use of must and have to. Integrating social and linguistic factors, this analysis exploits metadata gathered from the VOICE corpus, which has thus far been underused. The results of the inferential statistics indicate that the same sociolinguistic factors that influence the variation in ENL and ESL varieties also shape ELF grammar. These findings not only bring ELF closer to other English varieties but also demonstrate the advantage of studying ELF from a variationist sociolinguistic perspective.