
Interpretive researchers often face challenges in presenting the whole dataset for qualitative interpretation within the constrained space of academic articles, leading them to select only a limited number of extracts (Mann, 2016). However, the process of selecting these extracts is frequently conducted without sufficient rigour due to the lack of a well-defined, principled approach to selection. Without such an approach, interpretive researchers risk introducing bias during selection, and this potentially undermines their ability to identify those with confirmed salience. To address this critical gap, we developed a rigorous methodology for identifying salient extracts. The dataset for this study comprises approximately 650,000 words of Facebook posts discussing a culturally sensitive news topic in Thai society - a trans Thai net idol dressing as the Lord Buddha for Hallowe'en celebrations. Using this large dataset, we employ a corpus-based technique to outline a four-stage methodology for identifying salient extracts: (1) keyword identification, (2) plot generation and keyword thematisation based on the plots' main distribution areas, (3) investigation of the main communicators of posts linked to these keywords, and (4) the actual extract selection. Our findings demonstrate that this four-stage approach not only captures extracts with strong salience but also enhances the trustworthiness, providing a more rigorous framework for producing meaningful analysis and interpretation of interpretivist research findings.
This paper analyses the opposition of Us and Them in the context of power relationships. Power relationships are theorised and operationalised with the help of McClelland's human motivation theory. Power takes the most manifest forms in times of war where parties to the military conflict perceive each other as enemies. The Russo-Ukrainian War is no exception. The unique quadrilingual 320 million-word corpus of political and media discourses about the Russo-Ukrainian War assisted the critical discourse analysis. The data were collected in five countries (Ukraine, Russia, the United States, the United Kingdom, and France) from January 2022 to January 2025. The category ‘Us’ predictably tended to collocate with the words operationalising the need for affiliation. Although the Western leaders did not include Ukraine in the scope of ‘Us’, they frequently associated this country with the need for power in their war-related speeches. The Ukrainians, as ‘Them’, were perceived as needing to be empowered to defend themselves against aggression. The study confirms the premise of critical discourse analysis, requiring all categories to be interpreted in context.
Tests in the null hypothesis significance testing (nhst) framework are designed to identify differences between groups. Thus, researchers wishing to assess similarities are out of luck if they use widespread techniques in the field, such as t-tests, chi-square tests and linear regression. This is unfortunate given that researchers with many common study designs would benefit from knowing whether there is no difference - or at least no practical difference - between groups. For example, we might want to test empirically claims about there being no differences (regarding use of some linguistic feature) between first-language and advanced second-language English students. Or we might want to decide which sub-groups in our corpus to merge into one. This paper starts by outlining challenges inherent to the nhst framework in this context, and then provides two alternatives: equivalence testing and model-based techniques. We provide a guided introduction to both approaches for researchers wishing to expand their toolbox to carry out tests of group similarities.
Within the field of English for Research Publication Purposes (erpp), corpora can contribute valuable insights towards a better understanding of the issue of accuracy and role of language errors. Multilingual scholars, such as the Japanese scientists targeted in this corpus project, often know that their written English is prone to errors but, nevertheless, may feel powerless to deal with the problem. This short paper presents an account of corpus building for research and pedagogical application in erpp. It outlines the design and development of a corpus of research article manuscripts written by Japanese materials scientists. The Specialised Corpus of English for Materials Science (scems) comprises manuscript versions both before and after copyediting and, importantly, a sub-corpus of sentence-level grammar errors and their grammatical reconstructions. For comparative purposes, it also includes a set of research articles published in high-impact journals from the same field of materials science. With 2,490 sentences containing an estimated total of 4,500 errors, the scems, thus, establishes an empirical basis for the identification of error patterns in the research writing of the target population of Japanese materials scientists and the subsequent development of specialised pedagogical resources to support that population.
For many purposes, downsampling and using downsampled corpora are popular ways of extracting research data from larger text collections. Whilst most benefits of downsampling are gained from using it as a tool for qualitative inspection, it is not uncommon that the use of concordance or downsampled corpora is extended to collocation extraction. This is done with the underlying assumption that differences between concordance corpus or downsampled corpus and full corpus are trivial in this regard. This paper reports on results from an analysis, where this assumption was specifically tested. The results show, that whilst there can often be a relatively high degree of agreement between the two methods, they cannot be relied upon to produce correlating rank-orders or overlapping top collocate lists in every case. Further, the difference was more marked in the case of content words in contrast to function words.
AVATAR therapy is an innovative form of relational therapy for the treatment of distressing auditory verbal hallucinations, or voice-hearing, targeted at reducing voice-related distress. AVATAR therapy involves the creation of a digital simulation of a single voice, termed an 'avatar', which is used in a series of three-way therapeutic dialogues. This paper presents the AVATAR Therapy Dialogues Corpus, a specialised corpus containing orthographic transcriptions of AVATAR therapy sessions. We offer an overview of the corpus contents, and a detailed discussion of the design and construction of the corpus. We describe the processes and specialised tools created, transcription conventions, and mark-up designed to capture para-linguistic and non-speech features which may have clinical relevance. Finally, we discuss the potential of the corpus to provide a genuine innovation in clinical care, offering clinicians a data stream that could augment their understanding of patient experiences.
This study introduces the English Teacher Corpus (ETC), delineating its development, technical parameters, and research prospects. This spoken learner corpus contains spontaneous and semi-spontaneous speech tasks performed by Czech teachers of English as a foreign language (EFL). The tasks include a monologue, dialogue, picture-based narrative, reading-aloud assignment, and an interview conducted in the teacher's L1. Complementing this corpus is a reference counterpart featuring native English teachers based in the Czech Republic, mirroring the ETC's task design. In its 12.5 hours of recorded and transcribed text, the ETC consists of 76,122 tokens for the L2 and 31,898 tokens for the L1 sub-corpus. The corpus has been partly transcribed by Whisper AI and subsequently aligned using ExMARaLDA. The ETC marks a pioneering effort as the first spoken learner corpus produced by EFL teachers. Its innovation extends beyond its content, as it gave rise to a developmental and pedagogical project within a university teacher-training programme.
This study explores English noun sequences such as climate change, with a common noun modifier, and Harvard students, with a proper noun modifier, contrasting German and Swedish. The material is provided by the Linnaeus 5-million word non-fiction corpus. The results show that the most common type of translation correspondence - regardless of translation direction - is the German and Swedish (solid) compound noun (world war > Weltkrieg/v & auml;rldskrig). When specifically focussing on English proper noun modifiers, it is, however, evident that these are less likely to produce compound nouns in translations, due to language-internal preferences in German and Swedish. Apart from the formal properties of correspondences, this study also takes semantics into account. We show that some types of semantic relations between the head and its modifying noun, such as Composition, which identifies the material of the head noun (silk cloth), are more likely to be rendered as compound nouns in German and Swedish. Amongst the non-compound correspondences in German and Swedish, post-modifying prepositional phrases are one of the more prominent alternatives (climate signal > signal fr & aring;n ['from'] klimatet [Swedish]). This result is in line with our previous findings (Str & ouml;m Herold and Levin, 2019; and Levin and Str & ouml;m Herold, 2024), suggesting that Swedish, more than German, favours post-modification. Amongst the notable translation effects, we observe how translators sometimes make the content more explicit through the addition of a noun, but also that the opposite applies.
This paper examines how civil resistance is constructed in parliamentary discourse, focussing on naming choices for the 2013-14 protests in Ukraine known as the Euromaidan, Revolution of Dignity, or Maidan, and their surrounding contexts in speeches by Ukrainian Members of Parliament and foreign guests in full-house sittings of the Ukrainian parliament from 2013 to 2023. Using metadata annotation in the newly created corpus of Ukrainian parliamentary proceedings under the ParlaMint project, the study explores the interplay between naming and framing the protests over time and across collective actors at different levels of data aggregation. The results indicate a decline in explicit references to the 2013-14 protests in Ukrainian parliamentary discourse, but each name in question follows its unique trajectory, showing variations in relative frequency and semantic preference. The study also discusses the limits of using these names interchangeably, considering their non-arbitrariness and word-building potential within the context of competing framings of the events by different political players.
Lexical bundles are considered to be building blocks of written academic genres, including abstracts. Whilst the focus has primarily been on lexical bundles in abstracts of higher-level writing, little attention has been given to abstracts of lower-level theses. This study aimed to uncover differences between types and frequencies of lexical bundles and their structural patterns and textual roles within rhetorical moves in abstracts of students' theses written in L1 and L2 English. In addition to the two novice writers' corpora, we built a third corpus containing abstracts from research papers by L1 expert writers. Results revealed parallels and differences between L1 and L2 novice writing, and between novice and expert writing. L2 novice writers showed nearly over-cautious dependence on formulaic language, repeating the same patterns and functions and frequently failing to achieve their communicative purpose. Despite the similarities between the two types of novice writing, L1 novice writers showed similarities with L1 experts. A pedagogical implication is encouraging more diverse grammatical patterns and textual functions, which would result in a more accurate portrayal of the research conducted.
This paper presents the methodological, theoretical and practical aspects of the CoLaGe corpus (Corpus for the Study of Language and Gender in Spanish), an oral bi-dialectal corpus of Spanish collected in Valencia (Spain) and Guadalajara (Mexico). The corpus consists of three sub-corpora, one CoLaGe-GD) each containing three types of linguistic data: sociolinguistic interviews, roleplays and phonetic data, elicited through picture description tasks to elicit phonetic data. The linguistic data is complemented with a social psychological database published separately. Whilst the corpus has been designed for a research project, studying inter-relations between speaker's gender, sexuality and language use in two societies sharing the same language (but arguably differing in terms of gender norms and roles) it can be used for many different research areas ranging from gender studies to discourse analysis. The structure of the data permits quantitative comparisons across dialects, age groups and genders.
In this paper, we investigate audience responses to a controversial sub-genre of reality television: poverty porn. In so doing, we explore a previously neglected perspective in audience response research (i.e., that of the working class). Using a sample of 1,966 comments posted to YouTube, we examine the kinds of attitudes evoked by the BBC Three online series Britain's Forgotten Men and the underlying ideation driving such attitudes. Specifically, we contrast comments posted by those who self-identify as working class with those who make no such identification. Making a virtue of a limitation that has long dogged corpus-assisted discourse studies (i.e., being confined to the sentence level), we cut the data into forty-eight mini corpora (consisting of the clauses and clause complexes expressing the respective attitudinal categories) which allowed for the identification of salient differences. For example, whilst self-identified working class commenters used significantly more +TENACITY, +CAPACITY and +PROPENSITY to express a self-focussed discourse of merit and individual achievement, this did not exclude the expression of sympathy towards those on screen, nor criticism of structural factors. Whilst non-identified commenters largely used attitudes to negatively evaluate those on screen, the expression of -SECURITY (with an ideational focus on material deprivation) drove more sympathetic responses. We conclude with a discussion of the implications arising out of this study for researchers and content producers.
The corpus-assisted discourse studies (CADS) researcher has several established methods at their disposal, including concordance analysis, keyword analysis, and collocation analysis. Some researchers have recommended the use of topic modelling as a complementary approach, whilst others have highlighted methodological concerns. In this paper, we evaluate the application of a core CADS method, keyword analysis, alongside a newly introduced topic model, called keyword-assisted topic modelling (keyATM) (Eshima et al., 2023). We first establish what a strong and weak keyword is for keyATM using traditional topic modelling. Then we test the use of keyword analysis as a method for principally selecting keywords for use in keyATM. Keywords identified through traditional topic modelling serve as strong keywords for keyATM, but we find uneven success in using keyword analysis for keyword identification. We conclude with recommendations for researchers, a discussion of the limitations of topic modelling in CADS, and areas for future research.
Many universities use social media to communicate and engage with stakeholders, including students and staff. In recent years, universities were also faced with navigating the challenges resulting from the COVID-19 global pandemic and related restrictive measures that disrupted routine operations. In this paper, we examine a case study of a UK University and its posts on Twitter (now X) prior to, during and following the period of restrictive measures. With a focus on features of the 'Conversational Human Voice' (Kelleher, 2009), we report keywords and key emoji in a corpus of Twitter posts between 2018 and 2022. We demonstrate that despite the disruption of the pandemic and restrictive measures, the University maintained a consistent strategy, capitalising on the timeliness and broadcast functions of the platform to celebrate activities of its personnel and promote local events. Furthermore, we demonstrate how emoji and other paralinguistic elements can be incorporated into a multi-modal corpus analysis.
Previous studies on English in the Philippines (EngPh) have often investigated relativiser choice (e.g., who, that, zero) in restrictive relative clauses (RRc) in popular written and spoken modalities (e.g., student writing and dialogues), highlighting the conditioning effects of a narrow set of modalities on relativiser variation. However, EngPh extends beyond these modalities, presenting a gap in our knowledge: will the distribution of relativiser variants observed in these modalities align with lesser-known ones in EngPh? The extent to which relativiser variation is contributed to by factors beyond those identified in previous research is likewise unknown. Against this backdrop, this paper examines relativiser variation in human-antecedent RRcs on Twitter/X - a digital platform that has hybrid written, spoken, and electronically mediated characteristics. It employs quantitative, computational methods to examine RRcs extracted from the Twitter Corpus of Philippine Englishes (TcoPE). The findings show notable disparities in the ranking and distribution of relativiser variants between EngPh in speech and on Twitter. Moreover, it was found that within EngPh on Twitter, intra-linguistic, stylistic and extra-linguistic factors jointly restrict or constrain restrictive relativiser choice, with intra-linguistic factors exerting a stronger influence than others. The results provide support for a probabilistic representation of relativiser variation that involves all these factors but assigns greater significance to structural ones.