
This study extends the analysis of cognitive verb use across academic disciplines by focusing on the Corpus of English Life Sciences Texts (CELiST), a subcorpus of the Coruña Corpus of English Scientific Writing (CC). Previous research has concentrated on the linguistic variation in English cognitive verbs over time, highlighting significant differences in the presence and distribution of these verbs across various fields, including Linguistics, Engineering, and Medicine (Carrió-Pastor 2017). This reveals distinct patterns of epistemic stance that reflect the research practices of each discipline. Based on this line of research, this study aims to determine whether similar patterns emerge in Life Sciences texts, where the communication of empirical knowledge holds significant importance. By analysing the frequency, distribution, and phraseological patterns of cognitive verbs within Life Sciences texts, this research will explore how eighteenth and nineteenth-century researchers in this field express certainty, attitude, and belief, and how these expressions match or differ from those observed in other disciplines. The findings will contribute to a deeper understanding of academic writing in the Life Sciences during that time. Further research may compare the results obtained with texts from the specific field of present-day life sciences.
This study investigates how culturally grounded conceptualizations shape English collocational usage in learner writing. Drawing on a 4.6-million-word corpus of Thai English as a foreign language (EFL) academic texts, collocations were extracted using corpus-driven association measures, including Mutual Information (MI) and t-score, and were manually validated against the British National Corpus (BNC). Of the 2,239 collocation types identified, 316 types (20,173 tokens) were classified through manual coding as cultural-transfer or culturally interpretable transfer-related forms when Thai linguistic patterning was accompanied by culturally salient conceptual motivation. Rather than viewing non-standard collocations as deficiencies, this study interprets this culturally classified subset as evidence of cultural meaning-making and lexical nativization. Quantitative and qualitative analyses revealed five domains: Kinship Hierarchy, Religious Practice, Food/Domestic Culture, Localized Service Economy, and Institutional/Media Discourse. Together, these domains accounted for 41% of all non-standard collocations, indicating that cultural transfer represents a systematic rather than incidental phenomenon in Thai learner English. The study contributes to applied corpus linguistics and World Englishes by showing how local cultural schemas are encoded in collocational choices. Pedagogically, the findings highlight the need to foster learners’ intercultural lexical awareness and to develop corpus-informed instructional practices that help learners negotiate cultural identity while maintaining communicative clarity.
Situated at the intersection of corpus stylistics, translation studies, and multifactorial statistics, this paper focuses on the identification of the predictors of preservation or avoidance of repetition in English-to-Polish translation of repeatedly used reporting verbs signalling direct speech. Using a sample of 20 literary texts, we fit multiple negative binomial regression with mixed effects models to assess the effect that seven predictor variables (i.e. frequency of a source-text verb, number of its translation equivalents in lexical databases, its number of senses, its semantic type, its length in characters, date of translation of a novel, and individual translators) have on the response variable: the number of Polish target-text reporting verbs (types) an English source-text reporting verb is translated into. The overall model fit per the lowest AIC (Akaike Information Criterion) and BIC (Bayesian Information Criterion) values obtained through backward elimination reveals that the frequency of a source-text reporting verb, its semantic type as well as the individual translators have the largest incremental contribution to the model’s fit. More precisely, the proportion of variance in the outcome variable explained by both fixed and random effects (76%) is higher than the proportion of variance explained by fixed effects alone (72%). Without the translators treated as a random intercept, some variation would have been unaccounted for. The findings attempt to explain the translator’s decisions, whether to avoid or preserve patterns of repetition in source texts, with respect to rendering reporting verbs, which play an important stylistic effect in literary prose.
The present study investigates clarifying when-constructions (e.g., when I say more, I don’t mean that we should consume more, but really it’s about creating a rich life) in a sample of 971 examples from the Corpus of Contemporary American English. In this construction, a phrase or clause from a different or the same turn is repeated in a when-clause to provide clarification through the main clause, often accompanied by a refining clause. The present research centers around identifying configurations of variables that are correlated with whether a speaker repeats a phrase or clause from a different or the same turn in a when-clause. We do so with a predictive modeling approach that uses the following variables: refining clause, main clause verb, main clause polarity, and mode. We also discuss the prototypical configurations associated with each type of repetition. Our findings go beyond traditional research on self-repetition and clarification by showing how their conversational functions interact with other grammatical domains in language use: morphosyntax, lexicon, and mode. This is in line with recent Usage-Based Construction Grammar work that has shown that constructionhood is a complex intersection of internal and external properties.
English for Academic Purposes (EAP) recognises that academic language varies according to disciplinary context and that learners’ linguistic needs are shaped by the norms of specific fields. This study investigates whether grammatical patterns in research article abstracts (RAA) differ across levels of disciplinary specificity. A Multidimensional Analysis (MDA) was conducted on a large corpus of RAAs, which were classified according to the Web of Science hierarchy at three levels of disciplinary granularity. Independent t-tests were then used to compare Mean Dimension Scores across disciplinary groups for each of the extracted MDA dimensions. The results show significant differences at the broadest level (e.g. Social Sciences compared to Physical Sciences), driven mainly by variation in tense and authorial stance (past-oriented empirical reporting versus present-oriented author-focused framing), nominalisation and phrasal complexity versus quantitative and conditional features, and evaluative predicative constructions. Smaller but still significant differences were also found at more fine-grained levels of classification, including between related subdisciplines such as Mathematics and Statistics & Probability. Overall, the findings suggest that grammatical variation in academic writing is systematic across multiple levels of disciplinary organisation, indicating that disciplinary specificity in EAP extends beyond broad ‘hard’ and ‘soft’ science distinctions to more fine-grained subdisciplinary differences.
The Common European Framework of Reference for Languages (2001) and its Companion Volume (2020) emphasise the importance of linking expressions for pragmatic competence. Research on contrastive linking has long attracted scholarly interest; however (pseudo)longitudinal studies across different levels or whether gender may affect learners’ written production in this respect have been neglected. This study aims to address this gap by analyzing how Spanish English as a Foreign Language (EFL) learners at different levels express contrast and whether gender impacts their use of concessive expressions. Surprisingly, lower-level (B1) users show a wide range of expressions similar to higher-level users, while those at B2 levels tend to avoid ‘risky’ options. Interestingly, gender does not significantly influence learners’ use of connectors in this corpus, contradicting earlier findings that suggested female learners use more connectors than males.
This paper examines the transitivity potential of a group of English unergative verbs that denote physiological processes, a syntactico-semantic verbal class which has not received enough attention in the literature. Through a qualitative and quantitative corpus-based analysis of 24 verbs conducted on the Corpus of Contemporary American English (COCA), British National Corpus (BNC) and Corpus of Global Web-Based English (GloWbE), it will be shown that the syntactic flexibility of this verbal class is higher than stated in previous studies since, in addition to the cognate object construction (Burp the same garlic burps), the substance object construction (Breathe the smoke, Lucien), and the resultative construction (He yawned open his mouth), these verbs have been documented in seven other transitive patterns in which they increase their valency with the addition of a non-canonical direct object: x’s way constructions (I sweated my way through a painful run), reaction object constructions (Emma hiccups a yes), caused-motion constructions (She spits phlegm into a Kleenex), the preposition drop object alternation (He shit the rug), the understood body-part object alternation (The elk snuffled her face through the snow), away constructions (Everyone laughs the evening away), and causative constructions (Let’s burp this baby!).
Today, digital crowdfunding platforms allow researchers to increasingly use digital resources to reach and engage diversified audiences, making scientific content accessible to everyone. This paper explores how evaluation in text contributes information relevant to understanding how scientists use language to express their expert opinions of scientific research and their attitudes about the value of their projects. Starting from the compilation and analysis of a 50-science project corpus from Experiment.com, evaluative stance expressions in this work were classified according to Biber’s (2004) taxonomy into the following stance categories: verbs, adverbs, adjectives and nouns. Subsequently, genre analysis was applied to identify the discourse functions of these evaluative words in each rhetorical section of the project proposals. Results show that the analysed crowdfunding proposals are rich in stance verbs (52.65%) and, to a lesser extent, stance adjectives (23.52%), serving to express values of effort, improvement and diligence in the proposed projects, as well as judgement regarding experiments and ‘Lab Notes’ updates, respectively. This can be useful for both theoretical advancement and pedagogical purposes, that is, to apply scientists’ findings to digital communication teaching and learning.
Although L1-English fluency has been extensively studied from many angles, few contrastive studies examine whether fluency develops similarly or differently across L1-varieties while taking sociolinguistic variation into consideration. This paper aims to close this research gap and examines the use of three core strategies of fluency (or fluencemes), i.e. discourse markers, filled pauses and unfilled pauses, across Australian, British, Canadian, and New Zealand English. These fluencemes were extracted and manually disambiguated from the private conversation sections of the respective components of the International Corpus of English (ICE-AUS, ICE-GB, ICE-CAN, and ICE-NZ). The data were normalised per speaker and linked with the sociobiographic metadata of the speakers. Analysis using random forests revealed a consistent fluenceme distribution across the four varieties, with unfilled pauses being the most common, followed by discourse markers, and then filled pauses. This pattern suggests a ‘common fluenceme core’ among L1-English varieties. The influence of sociolinguistic variables —gender, age, education, and occupation— was modest across varieties and exhibited diverse trends. Male speakers tend to use filled pauses more frequently but fewer unfilled pauses compared to female speakers. Increasing age did not significantly affect the frequency of these strategies; however, older speakers tend to use discourse markers less frequently. Both education and occupation showed a slight positive correlation with overall fluency.
The current study aims to increase the accessibility of Nelson’s (2024) recently suggested construction-based complexity measure by providing a tool that can calculate the measure for single or multiple texts. To validate the tool, complexity scores for the International Corpus Network of Asian Learners of English corpus (ICNALE) were compared with Nelson’s (2024) results. In addition, complexity scores were calculated for a new dataset, the Common European Framework of Reference English Listening Corpus (CEFR), along with the MERLIN corpus, which includes learner writing samples from learners of Czech, German, and Italian. Complexity scores generally increased across CEFR levels in all of the datasets. However, the complexity scores in the current study tend to be higher than the original study due to differences in the sentence splitting approach. The sentence tokenisation method used is deemed to be more appropriate, and it may be concluded that the Construction Complexity Calculator (ConPlex) tool accurately calculates Nelson’s measure. It is hoped that the tool will allow researchers to calculate the complexity of constructions at the text level for a wide range of research purposes.
The Multi-Feature Tagger of English (MFTE) provides a transparent and easily adaptable open-source tool for multivariable analyses of English corpora. Designed to contribute to the greater reproducibility, transparency, and accessibility of multivariable corpus studies, it comes with a simple GUI and is available both as a richly annotated Python script and as an executable file. In this article, we detail its features and how they are operationalised. The default tagset comprises 74 lexico-grammatical features, ranging from attributive adjectives and progressives to tag questions and emoticons. An optional extended tagset covers more than 70 additional features, including many semantic features, such as human nouns and verbs of causation. We evaluate the accuracy of the MFTE on a sample of 60 texts from the BNC2014 and COCA, and report precision and recall metrics for all the features of the simple tagset. We outline how that the use of a well-documented, open-source tool can contribute to improving the reproducibility and replicability of multivariable studies of English.
TreeTagger is a multilingual tagger capable of performing headword and POS tagging. However, before the completion of this project, Indonesian had not been supported. Thus, corpus query systems employing TreeTagger as a subsystem, such as CQPweb v.3.3.10 and LancsBox v.5, were incapable of annotating Indonesian texts. This context leads to the following research: 1) develop Indonesian language support for TreeTagger, 2) evaluate its performance, and 3) integrate the support into two popular corpus query systems, namely CQPweb and LancsBox, and demonstrate its functionalities. The research procedure can be concisely summarised as follows: training, annotation and evaluation, and incorporation. A pre-annotated corpus and lexicon were used in the training process. Headwords for the lexicon and corpus were semi-automatically added using MorphInd, augmented with expert revisions. The training produced an Indonesian TreeTagger parameter file, whose accuracy for POS and headword annotation was 96 per cent and 91 percent respectively. The parameter file has been incorporated into LancsBox v.6 and CQPweb 3.3.11, enabling support for the Indonesian language.
Research on noun phrase use in EFL writing has mainly focused on linguistic complexity and accuracy, lexical richness, and phraseological competence. However, the relationship between noun lexical diversity of nouns and the syntactic complexity of the noun phrases in which these nouns appear remains underexplored. To address this gap, this paper examines the lexical diversity of head nouns in noun phrases within a sample of emails written by L1 Spanish EFL learners at B1 and C1 proficiency levels, taken from the FineDesc Learner Corpus. The analysis considers both the lexical diversity of nouns and the syntactic complexity of the noun phrases they head. The findings reveal: a) a narrower range of nouns at the B1 level compared to the C1 level; b) a low percentage of nouns from both levels, based on the English Vocabulary Profile; and c) differences in NP complexity between the two proficiency levels (B1 and C1), depending on whether the head nouns are concrete or abstract. The paper underscores the importance of combining different complexity measures ––namely, lexical diversity and NP complexity analyses–– to gain a more comprehensive understanding of learners’ use of noun phrases.
Emoji (e.g., 🤪✈🧁) are increasingly used on social media by people of all ages, but little is known about the concept ‘emoji literacy’. To investigate different age groups’ emoji preferences, an exploratory corpus analysis was conducted using an innovative corpus-gathering method: children and adults were instructed to add emoji magnets to pre-constructed printed social media messages. The corpus (with 1,012 emoji) was coded for the number of emoji used per message, the type of emoji, their position and function in the message, and the sentiment they conveyed. Intuitions about emoji use turned out to be similar for children and adults, with greater use of facial emoji, emoji at the end of messages, emoji to express emotions, and emotional emoji to convey positive sentiment. Children’s emoji preferences were studied in more detail. Results revealed that their age, gender, smartphone ownership, and social media use related to differences in the number, position, and function of the emoji used. The data showed that older children, girls, children with their own smartphone, and children using social media exhibited a more advanced and sophisticated use of emoji than younger children, boys, and children without smartphones or social media experience. This study constitutes an important first step in exploring children’s emoji literacy and use.
This study explores the usage of nonbinary pronouns on X (formerly known as Twitter), focusing on THEY and neopronouns like ZE or XE within the nonbinary community. Building on the increasing practice of sharing pronouns, especially in online spaces, the research collects 1,980 X accounts using Followerwonk. Despite ideological differences across U.S. regions, no substantial variations in pronoun usage are observed. Notably, a preference for rolling pronouns (e.g., they/she) emerges, with fewer instances of monopronoun usage (e.g., they). When a single pronoun is chosen, it is often accompanied by the respective accusative form, while rolling pronoun users tend to omit the accusative. Users with binary pronouns often prioritize it as their first chosen pronoun. THEY remains the predominant nonbinary pronoun, with neopronouns being rare. The study highlights X profiles as valuable sources for understanding linguistic patterns related to social trends, particularly in the context of gender equality and network relations.
Computer-Mediated Communication is part of the everyday lives of a great many people of all ages, cultures, social statuses, and geographical locations. In the present study, I explore non-categorical syntactic variability in internet language with data from the Corpus of Global Web-Based English (GloWbE), which includes material from blogs, forums, comments, and other types of websites. The focus is on how the geographical area of internet users affects the use of the clausal complementation patterns available for the verb regret. The analysis of more than 10,000 examples from Indian, Sri Lankan, Pakistani, Bangladeshi, Singaporean, Malaysian, Philippine, Hong Kong, British, and American Englishes shows that geographical origin does have a bearing on the complementation system of this verb, in terms of both the factors that determine variability and the preferences for particular patterns. The varieties displaying more similarities are those that are geographically close, making the distinction between three geographical areas possible: South Asia (India, Sri Lanka, Pakistan, and Bangladesh), South-East Asia (with Singapore, Malaysia, and the Philippines) and East Asia (Hong Kong).
Twitter for academic purposes has been analysed from multiple perspectives such as genre analysis, the use of multimodality and hypertextuality, or type of participants; yet interactivity between writers and readers remains under-researched. This study analyses academic-related conversations from the Twitter conference genre, particularly focusing on the discussion session. Its objective is to identify the main interactional patterns, communicative functions, and digital discourse features in tweets. Dialogic turns were classified into comments, questions, responses, follow-up conversations, and automatic comments. Findings reveal that the main reasons behind online interaction correspond with community building and knowledge construction purposes. The digital medium does shape the form of tweets, which shows a high level of evaluative language, conversational style features, hedging, and emojis. All in all, these discursive features help create a welcoming and engaging style needed to engage in online science communication practices on social media.