This paper reports a Lexical Multidimensional Analysis (Berber Sardinha, 2014, 2019, 2020a, 2020b; Berber Sardinha & Fitzsimmons-Doolan, 2025) of English Google Books trigram data from publications spanning 21 decades (1800–2008). The study aims to compare social representations of the peoples of the Americas in publications written in English, focusing on recurrent noun phrase trigrams containing the collocations North American, Latin American, Central American, and South American. The findings indicate that regions outside North America are predominantly represented through a tripartite pattern involving economic opportunity, internal conflict, and natural abundance, reflecting enduring socio-historical perspectives in English-language discourse.
Abstract This entry examines the historical development of artificial intelligence (AI) and its transformative potential across various domains in linguistics. It highlights the emergence of generative AI, particularly large language models (LLMs), and their role in advancing tasks such as syntactic parsing and pattern recognition. The discussion addresses the mechanisms underlying AI models, including transformer and self‐attention architectures, as well as their limitations in replicating human linguistic nuance. By integrating insights from corpus linguistics and AI, this entry underscores the collaborative potential of these technologies while advocating for critical human oversight and future advancements in interactive and explainable AI systems.
Abstract This entry reviews key topics and methods in corpus‐based analyses of academic language, addressing both written and spoken communication. It highlights the range of variation across disciplines, registers, student levels, and research article sections, as well as the diversity of methodological approaches, which include manual, automatic, and statistical analyses employing both univariate and multivariate methods.
Abstract This entry provides an introduction to the section on Corpus Linguistics in The Encyclopedia of Applied Linguistics . It begins by examining how the term corpus has been defined over time and highlighting its key characteristics. It then explores how Corpus Linguistics as a field has been conceptualized, identifying common ideas across various perspectives. Lastly, it provides an overview of the entries on Corpus Linguistics in the EAL to offer a comprehensive perspective of the field, including its development, foundational concepts, notable researchers, and diverse applications to research questions.
Abstract A growing body of corpus linguistic research investigates social media, applying diverse methodological approaches to uncover its linguistic and discursive features. This entry provides an overview of two primary study types: those examining specific lexicogrammatical choices and those employing Multi‐Dimensional Analysis. These approaches offer detailed insights into the recurring linguistic patterns that define social media texts, highlighting their role in constructing distinct communicative styles and conveying specific discourses. Communicatively, corpus‐based studies confirm the personal, involved nature of social media communication while also uncovering a less recognized, information‐based objective style. Ideologically, research demonstrates how social media leverages particular lexical and grammatical features to disseminate practices such as policy legitimation, science mistrust, climate change denial, trolling, misogyny, sexism, and racism.
Multi-dimensional Analysis: Research Methods and Current Issues provides a comprehensive guide both to the statistical methods in Multi-dimensional Analysis (MDA) and its key elements, such as corpus building, tagging, and tools. The major goal is to explain the steps involved in the method so that readers may better understand this complex research framework and conduct MD research on their own.Multi-dimensional Analysis is a method that allows the researcher to describe different registers (textual varieties defined by their social use) such as academic settings, regional discourse, social media, movies, and pop songs. Through multivariate statistical techniques, MDA identifies complementary correlation groupings of dozens of variables, including variables which belong both to the grammatical and semantic domains. Such groupings are then associated with situational variables of texts like information density, orality, and narrativity to determine linguistic constructs known as dimensions of variation, which provide a scale for the comparison of a large number of texts and registers.This book is a comprehensive research guide to MDA
Abstract Representativeness is a fundamental consideration in corpus linguistics, as corpora are intended to accurately reflect a specific language, domain, or variety. Despite its significance, the concept is often overlooked in practice. Researchers frequently describe corpora as “representative” without providing statistical evidence to support this claim, raising questions about the clarity of the concept and highlighting the need for more careful attention to the principles of representativeness. This entry outlines the framework for corpus representativeness proposed by Douglas Biber, Jesse Egbert, and Bethany Gray. The framework draws attention to surveying the target domain and making informed decisions about how to capture it within a corpus. Additionally, it incorporates statistical analysis during the corpus compilation phase to assess which linguistic features are adequately represented and for which precise measurements can be reliably taken.
Abstract This article argues that a register-based Multi-Dimensional (MD) description is a suitable route for characterizing AI-generated language in corpus linguistics. The argument is illustrated with two sample studies: a grammar-oriented investigation that applies traditional MD analysis to English-as-a-foreign-language textbook texts and a discourse-oriented analysis that relies on lexical MD analysis to explore AI-generated pop music lyrics. In both cases, the results reveal sharp differences between AI-generated and human language. In the EFL texts, AI-written texts are more informational, abstract, and impersonal whereas human texts display interpersonal awareness, stance, and engagement. In the pop lyrics, AI generates moralized empowerment discourses that recast historical conflicts depicted in rap music as generalized virtue narratives. In both cases, AI demonstrates signs of register deficit (a limited awareness of register variation due to shallow knowledge of the linguistic constituency of human registers) and register metamorphosis (generation of texts that resemble one register on the surface but are realized linguistically as another).
Resumo: Este estudo visa descrever as características linguísticas de tratamentos médicos promovidos durante a pandemia de COVID-19. Um corpus contendo dois subcorpora foi coletado: o primeiro subcorpus consiste em artigos acadêmicos que recomendam tratamentos não endossados pelas agências reguladoras de saúde; o segundo subcorpus contém artigos que focam em várias questões relacionadas à COVID-19 sem endossar tais tratamentos não recomendados. A metodologia empregou um tipo de Análise Multidimensional Lexical (Berber Sardinha; Fitzsimmons-Doolan, 2025), que consistiu na detecção de deslocamentos de colocação. Foram identificadas cinco dimensões: intervenções médicas vs. impacto psicológico; ética em pesquisa vs. análise comparativa de tratamentos; análise estatística na pseudociência vs. compartilhamento de dados na ciência real; promoção de drogas reaproveitadas vs. avaliação crítica e práticas de ciência aberta; impacto de tratamentos reaproveitados vs. aprovação ética e conformidade reguladora. Essas dimensões capturam os principais recursos comunicativos do fazer científico em torno de tratamentos aprovados e contestados durante a pandemia. O trabalho mostra que embora se confundam, o fazer científico genuíno e o pseudocientífico utilizam linguagem distinta e se apoiam sobre discursos e formações diferentes. Tanto a linguagem quanto os discursos em questão são detalhados no artigo.
Multi‐dimensional (MD) Analysis originated in the 1980s as a method for describing register variation in corpora, providing comprehensive analyses of registers and domains. It employs a large set of linguistic features and statistical methods to determine the correlated sets of linguistic features corresponding to the underlying dimensions of variation. These dimensions reflect the major communicative functions realized by the shared linguistic features in the texts. MD Analysis has since evolved into a family of approaches, deriving from the original approach, which includes lexical, visual, and multimodal MD Analysis. The entry outlines the foundational principles of the approach and introduces its extensions.
Lexical Multidimensional Analysis (LMDA), an extension of Biber's (1988) Multidimensional Analysis, seeks to identify dimensions (correlated lexical features across texts in a corpus) unveiling underlying patterns of lexical co-occurrence and variation within texts that are operationalized as a variety of latent, macro-level discursive constructs. Initially developed in the 2010s, LMDA has been applied to diverse domains, including education policy, national representations, applied linguistics, music, the infodemic, religion, sustainability, and literary style. This Element introduces LMDA for the identification and analysis of discourses and ideologies, offering insights into how lexis marks discourse formations and ideological alignments. Two case studies demonstrate the application of LMDA: uncovering discourses on climate change within conservative social media and analyzing ideological discourses in migrant education.
Most – if not all – language use is multimodal, meaning it involves more than one semiotic mode. However, with rare exceptions, most language research in linguistics, including corpus linguistics, has been monomodal, focusing exclusively on the verbal/textual component. In this chapter, I present the results of the analysis of a corpus of social media posts and propose ways in which the findings can be explored in the classroom. The analysis was carried out using methodological extensions of multidimensional analysis – specifically, lexical, visual, and multimodal multidimensional analysis. The chapter further includes information on tools for collecting and visually annotating a corpus of texts and images.
Since the United Nations Agenda 2030 was set up, many countries have worked to find solutions to interconnected global issues such as hunger, poverty, education and sustainability just to mention a few. After the strikingofCOVID-19 pandemic, the world witnessed researchers gathering around in an international effort to find solutions to the problems previously pointed out. In order to describe the lexis used in research papers discussing the Sustainable Development Goals (SDGs), we compiled a 2 million-word corpus from the PLOs platform with the AntCorGen program using the query term "Sustainable Development Goal". After that, we part of speech-tagged the corpus using the Tree-tagger of Sketch Engine in order to carry out a lexical multidimensional analysis (LMDA). Our aim was to (i) identify the major dimensions of variation based on the lexico-grammatical characteristics in research papers published in English by international researchers and (ii) observe how the SDG themes stand out in the study corpus. Results showed six dimensions that were named according to the words concentrated in each one: Government Actions, Presenting Results, Data Interpretation, Data Presentation, Data Quality, Research Procedure. The first dimension is the one that best illustrates a theme related to the SDGs whereas the other five dimensions clearly show the lexis that illustrates how researchers describe their own articles.
The goal of this study is to assess the degree of resemblance between texts generated by artificial intelligence (GPT) and (written and spoken) texts produced by human individuals in real-world settings. A comparative analysis was conducted along the five main dimensions of variation that Biber (1988) identified. The findings revealed significant disparities between AI-generated and human-authored texts, with the AI-generated texts generally failing to exhibit resemblance to their human counterparts. Furthermore, a linear discriminant analysis, performed to measure the predictive potential of dimension scores for identifying the authorship of texts, demonstrated that AI-generated texts could be identified with relative ease based on their multidimensional profile. Collectively, the results underscore the current limitations of AI text generation in emulating natural human communication. This finding counters popular fears that AI will replace humans in textual communication. Rather, our findings suggest that, at present, AI's ability to capture the intricate patterns of natural language remains limited.
Embora a música popular tenha sido objeto de estudo linguístico, tem havido pouco interesse na descrição da canção do ponto de vista linguístico e musical simultaneamente. A maioria dos estudos descritivos foca no modo verbal da produção musical, isto é, o texto da composição (a letra da música). Neste estudo, buscamos realizar uma descrição tanto do texto escrito da canção popular em inglês quanto de sua manifestação musical, de tal modo a proporcionar uma visão holística desse produto cultural. Para tanto, empregamos uma abordagem baseada na Linguística de Corpus, por meio da qual foi possível coletar e analisar uma grande amostra de canções em inglês, incluindo mais de 200 mil letras de música e mais de 97 mil canções indexadas acusticamente. A metodologia baseou-se na Análise Multidimensional, uma abordagem baseada em corpus que permite identificar as dimensões de variação subjacentes a uma determinada variedade linguística. Em relação ao texto da composição musical, quatro dimensões foram identificadas, as quais refletem a predominância de determinados discursos. Em relação ao componente acústico, foram identificadas três dimensões, cada uma correspondendo a uma determinada musicalidade. A pesquisa ilustra a possibilidade de descrição linguística e acústica em larga escala por meio de metodologias baseadas em corpus. De modo geral, os resultados mostram que a música popular em inglês, embora extremamente variada, pode ser resumida em torno de quatro padrões textuais e três padrões musicais.
The present study explores the development of grammatical complexity in L2 English writing at the beginner, lower intermediate, and upper intermediate levels to see (i) to what extent the developmental stages proposed in Biber et al. (2011) are evident in low-proficiency L2 writing, and if so, what the patterns of progression are, and (ii) whether students gradually move away from speech-like production toward more advanced written production. We use data from COBRA, a corpus of L1 Brazilian Portuguese learner production, along with BR-ICLE and BR-LINDSEI. All the data were tagged using the Biber tagger (Biber, 1988) and the Developmental Complexity tagger (Gray et al., 2019), and subsequently analyzed using a technique developed in Staples et al. (2022) to quantify developmental profiles across levels. The technique considers not only overall change in frequency across levels, but also the incremental variation across each adjacent level (based on % frequency changes). The results show that the features were infrequent overall, with a majority of both clausal and phrasal features exhibiting an increase in frequency across the levels, albeit to varying degrees. This general pattern is contrary to predictions based on findings from previous studies, which found phrasal features increasing in use and clausal features decreasing in use. Nonetheless, for the features associated with each developmental stage, the frequencies generally increased, becoming more similar to advanced written production and more dissimilar to spoken production, as hypothesized in Biber et al. (2011).
Since the United Nations Agenda 2030 was set up, many countries have worked to find solutions to interconnected global issues such as hunger, poverty, education and sustainability just to mention a few. After the striking of COVID-19 pandemic, the world witnessed researchers gathering around an international effort to find solutions to the problems previously pointed out. In order to observe how research papers have communicated their studies discussing the Sustainable Development Goals (SDGs), we compiled a corpus of 2,000 million words from the PLOs platform using the AntCorGen program using the query term “Sustainable Development Goal”. After that, we part of speech-tagged the corpus using the Tree-tagger of Sketch Engine in order to carry out a lexical multidimensional analysis (MD). Our aim was to (i) identify the major dimensions of variation based on the lexico-grammatical characteristics in research papers published in English by international researchers and (ii) carry out a lexical multidimensional analysis to observe how the SDG themes stand out in the study corpus. Results showed six dimensions that were named according to the words concentrated in each one: Government Actions, Presenting Results, Data Interpretation, Data Presentation, Data Quality, Research Procedure. The first dimension is the one that best illustrates a theme related to the SDGs whereas the other five dimensions clearly show the lexis that illustrates how researchers describe their own articles.
This study identified and tracked the major discourses present in the first 50 years of TESOL Quarterly. A corpus of articles published in the journal was collected, tagged, and analyzed for lexical dimensions of variation (the lexical parameters underlyingvariation across texts in the journal). A factor analysis detected the sets of lexical words cooccurring in the texts. The factors were interpreted into five dimensions: (1) critical, social, cultural, discourse or identity versus language assessment and testing; (2) applications of linguistic theory versus language policy, education and planning; (3) quantitative research methods versus positivist teaching materials and techniques; (4) language teaching and learning versus word-based investigations; and (5) reading and writing versus listening and speaking. The dimension scores were entered in a cluster analysis that identified the two principal eras of the journal: the first from 1967 to the early 1990s, and the second from the early 1990s to 2016.
Although popular music has received attention in linguistics, there has been little interest in describing music from both a linguistic and a musical standpoint simultaneously. Most linguistic studies focus on the verbal mode of the musical production, that is, the text of the composition (the lyrics). In this study, we seek to conduct a description of both the written text and the acoustic indices of popular music, in such a way as to provide a holistic view of this cultural product. To do so, we employed a corpus approach, through which it was possible to collect and analyze a large sample of English songs, including more than 200,000 lyrics and more than 97,000 acoustically indexed songs. The methodology was based on Multidimensional Analysis, a corpus-based approach that enables the identification of the dimensions of variation underlying a given linguistic variety. Regarding the lyrics, four dimensions were identified, each reflecting the predominance of particular discourses. Regarding the acoustic component, three dimensions were identified, each corresponding to a particular musical pattern. This paper illustrates the possibility of large-scale linguistic and acoustic description using corpus-based methodologies. Overall, the results show that popular music in English, although extremely varied, can be summarized around four textual patterns and three musical patterns.
This paper introduces an initial text typology of social media posts from a multi-dimensional (MD) perspective. Text types are "[g]roupings of text that are similar in their linguistic form" (Biber 1989: 13). This text typology is based on a new MD analysis of social media messages presented in the paper. The corpus consists of 60,000 social media messages in English compiled from Facebook, Twitter, Instagram, Reddit, Telegram, and YouTube. After the texts were cleaned up, the corpus was tagged with the Biber Tagger and post-processed with the Biber Tag Count. Three dimensions of variation were determined, each representing an underlying parameter of variation. Once the texts were scored on each of the dimensions, a k-means cluster analysis was carried out, and the optimal number of clusters was determined using the Cubic Clustering Criterion statistic. A two-way typology was developed based on the dimensional characteristics of each cluster and on careful qualitative analysis of text samples.
Isabel Trancoso合作论文数Instituto Superior Tecnico, University of Lisbon2