Swearing involves the use of words interpreted as obscene, derogatory, or offensive, and so can have a significant societal impact in the digital world. While it is widely acknowledged that the functions of swearing vary according to who is doing it and in what context, other dimensions of variability in online swearing remain to be fully explored. This article argues that we need to move beyond purely distributional text analytics of inferred social variables in analyzing online swearing if we are to systematically examine the social impacts of swearing in born-digital data. Using computational text analysis and network analysis methods, three large datasets representing snapshots of English Twitter/X communication in Australia, the United Kingdom, and the United States are examined through four interrelated research questions: (1) Do rates of vulgarity differ significantly across these three English-speaking regions and across different times of day? (2) To what extent do users employ non-standard orthographic and typographic variations to obscure vulgar expressions? (3) How does vulgarity correlate with users' positions within social networks, specifically their network integration and follower counts? (4) What patterns emerge when these dimensions of variability are examined together? Findings demonstrate that while it remains important to examine who swears online, it is also critical to examine other dimensions of variability, including where, when, how, and with whom swear words are used. The implications of extending our understanding of these different dimensions of variability in born-digital data for studies of online swearing and the data-intensive humanities more generally are also discussed.
Opposing social movements are groups that have conflicting objectives on a shared social justice issue. To maximize the probability of their movement's success, groups can strategically portray their group in a favourable manner while discrediting their opposition. One such approach involves the construction of victimization discourses. In this research, we combined topic modelling and critical discursive psychology to explore how opposing groups within the feminist movement used victimization as a lens to understand their movements in relation to transgender women. We compiled a dataset of over 40,000 tweets from 14 UK-based feminist accounts that included transgender women as women (the pro-inclusion group) and 13 accounts, that excluded transgender women (the anti-inclusion group). Our results revealed differences in how victimization was employed by the opposing movements: pro-inclusion groups drew on repertoires that created a sense of shared victimhood between cisgender women and transgender women, while anti-inclusion groups invoked a competitive victimhood repertoire. Both groups also challenged and delegitimised their oppositions' constructions of feminism and victimhood. These findings add to our understanding of the communication strategies used by opposing movements to achieve their mobilization goals.
Parliamentary discourse is highly regulated, leading to an almost blanket avoidance of explicit vulgarity or overtly offensive language. Yet it is nevertheless replete with examples in which the language used by members is construed as 'unparliamentary'. This study examines the occurrence of 'unparliamentary' as a metapragmatic label across the entire corpus of the Australian Federal Hansard from 1901 to 2024, and probes how it can be used to implement specific metapragmatic acts (i.e. doing something through labelling talk as 'unparliamentary'), as well as how it can also become an object of metapragmatic discourse (i.e. a topic of debate in its own right). In so doing we explore how the boundaries of offensive, objectionable or otherwise disorderly language use in parliamentary discourse are established, maintained, contested, as well as change and evolve overtime. As the Australian Federal Hansard constitutes a relatively large corpus of more than 900 million tokens, the study draws in a dialogic and iterative manner from both computational and interpretive methods of analysis. This dialogic form of analysis indicates that what is encompassed by the notion of 'unparliamentary' is broader and more complex than what is prescribed in the Standing Orders and associated codes of practice of the Australian Federal Parliament. (c) 2025 The Author(s). Published by Elsevier B.V. This is an open access article under the CC BY license (http:// creativecommons.org/licenses/by/4.0/).
This paper introduces and evaluates the Coordination Network Toolkit, an open-source software package and methodological framework designed to detect and analyse coordinated behaviour on social media platforms. As the dynamics of online communication continue to evolve, coordination analysis has emerged as an important field of study with significant implications for understanding online influence, digital astroturfing, and online activism. Recognising the absence of a comprehensive, open-source tool for constructing coordination networks, our approach fills this gap, catering to multiple behaviors across diverse social media platforms. Our approach synthesises and significantly enhances various methods to provide a methodological framework for ‘multi-behaviour’ coordination detection, utilising weighted, directed multigraphs to capture intricate coordination dynamics. We evaluate our approach by revisiting a case study of the 2020 #ReopenAmerica Covid protest movement on Twitter. The paper concludes with a set of recommendations for future work, emphasising the need for a tailored statistical framework for coordination analysis and a deeper exploration into the motives behind online coordination.
Responding to the challenge for qualitative researchers to claim a central place in conversations about big data, analytics, datafication, data mining and the role of algorithms, this article describes a mixed-method research partnership focused on algorithmic ethnography. In the debates about the opacity of online algorithms, qualitative researchers typically advocate for access to code. This standard discourse centralises the technical aspects of big data and networked ethnographies. Instead, this article outlines a research methodology that analyses algorithmic discourses by working alongside the technical expertise of data scientists and utilizes the affordability of big data methods to do qualitative work. The potential for qualitative research skills to investigate the underlying technical processes that frame online social interactions is proposed as a way to place how people understand the world at the centre of big data research.
AbstractIntroductionMedicinal cannabis is now legal in 44 US jurisdictions. Between 2020 and 2021 alone, four US jurisdictions legalised medicinal cannabis. The aim of this study is to identify themes in medicinal cannabis tweets from US jurisdictions with different legal statuses of cannabis from January to June 2021.MethodsA total of 25,099 historical tweets from 51 US jurisdictions were collected using Python. Content analysis was performed on a random sample of tweets accounting for the population size of each US jurisdictions (n = 750). Results were presented separately by tweets posted from jurisdictions where all cannabis use (non‐medicinal and medicinal) is ‘fully legalised’, ‘illegal’ and legal for ‘medical‐only’ use.ResultsFour themes were identified: ‘Policy’, ‘Therapeutic value’, ‘Sales and industry opportunities’ and ‘Adverse effects’. Most of the tweets were posted by the public. The most common theme was related to ‘Policy’ (32.5%–61.5% of the tweets). Tweets on ‘Therapeutic value’ were prevalent in all jurisdictions and accounted for 23.8%–32.1% of the tweets. Sales and promotional activities were prominent even in illegal jurisdictions (12.1%–26.5% of the tweets). Fewer than 10% of tweets were about intoxication and withdrawal symptoms.Discussion and ConclusionThis study has explored if content themes of medicinal cannabis tweets differed by cannabis legal status. Most tweets were pro‐cannabis and they were related to policy, therapeutic value, and sales and industry opportunities. Tweets on unsubstantiated health claims, adverse effects and crime warrants continued surveillance as these conversations could allow us to estimate cannabis‐related harms to inform health surveillance.
In this article, we examine two interrelated hashtag campaigns that formed in response to the Victorian State Government’s handling of Australia’s most significant COVID-19 second wave of mid-to-late 2020. Through a mixed-methods approach that includes descriptive statistical analysis, qualitative content analysis, network analysis, computational sentiment analysis and social bot detection, we reveal how a small number of hyper-partisan pro- and anti-government campaigners were able to mobilise ad hoc communities on Twitter, and – in the case of the anti-government hashtag campaign – co-opt journalists and politicians through a multi-step flow process to amplify their message. Our comprehensive analysis of Twitter data from these campaigns offers insights into the evolution of political hashtag campaigns, how actors involved in these specific campaigns were able to exploit specific dynamics of Twitter and the broader media and political establishment to progress their hyper-partisan agendas, and the utility of mixed-method approaches in helping render the dynamics of such campaigns visible.
Public discourse about the COVID-19 that appears on Twitter and other social media platforms provides useful insights into public concerns and responses to the pandemic. However, acknowledging that public discourse around COVID-19 is multi-faceted and evolves over time poses both analytical and ontological challenges. Studies that use text-mining approaches to analyse responses to major events commonly treat public discourse on social media as an undifferentiated whole, without systematically examining the extent to which that discourse consists of distinct sub-discourses or which phases characterize its development. They also confound structured behavioural data (i.e., tagging) with unstructured user-generated data (i.e., content of tweets) in their sampling methods. The present study aims to demonstrate how one might go about addressing both of these sets of challenges by combining corpus linguistic methods with a data-driven text-mining approach to gain a better understanding of how the public discourse around COVID-19 developed over time and what topics combine to form this discourse in the Australian Twittersphere over a period of nearly four months. By combining text mining and corpus linguistics, this study exemplifies how both approaches can complement each other productively.
Trust is fragile. The 2018 Facebook and Cambridge Analytica debacles highlighted how data harvested from social media platforms can be used not only for commercial purposes but also for political manipulation. This incident and the widespread discussion around it further demonstrated the following issues: unethical data collection enabled by a platform; unethical use of data for corporate and political interest; and unethical data sharing by an academic. Research needs to be credible to maintain social license. Data is the lifeblood of research. For research to remain credible, research needs to remain fundamentally ethical and research methods comprising data collection and data analysis need to be robust, transparent, repeatable, and auditable. Such methods alone cannot create credibility, but research data infrastructure design and implementation can provide a foundation for credibility by addressing these fundamental processes. Social science research has traditionally relied on data collection methods such as surveys, interviews, and ethnographic observations. However, an increasing proportion of human life is being mediated by online platforms, with approximately 2.3 billion active users on Facebook and 326 million active users on Twitter (Statista 2019). Social media data collection and analysis have become imperative for researchers interested in various phenomena playing out in these new media. This paper discusses the current state and issues of social media data collection and describes the Digital Observatory’s approach to establishing a credible and trusted research data infrastructure.
This paper presents the application of Lamb waves to detect and locate laminar damages using a beam forming imaging methodology. Beam forming is using a network of transducers that are used to sequentially scan the structure before and after the presence of damage by transmitting and receiving guided wave pulses. An image of the damage is reconstructed by analysing the cross correlation of the scatter signal with the excitation pulse and enables the detection and location of potential damage areas. The results of simulation and experimental studies show that the method enables the reliable detection of structural damages with locating inaccuracies in the order of a few millimeters within inspection areas of 300 x 300 mm2 using a transducer network of only four transducer elements.