The criminal case of Jeffrey Epstein has generated a complex, long-running global discourse on digital platforms, characterized by punctuated attention shocks, conspiracy theories, and blame attribution. To facilitate the computational study of these dynamics, we introduce the FUBU-EPSTEIN dataset, a large-scale, multi-dimensional research corpus of 54.38 million Twitter statuses authored by 7.11 million users and collected between August 2019 and April 2023. The source corpus was captured continuously in near-real time and enriched with a directed social contact graph of 37.03 million edge rows and annotations from Qwen2.5-7B-Instruct covering sentiment, conspiracy and misinformation stance, toxicity, moral emotion, and related dimensions. For public distribution, we created a textless, de-identified derivative that retains one row for every deduplicated status, categorical and numeric annotations, coarse temporal information, and 46.15 million internal status relationships. It excludes tweet text, original post and user identifiers, handles, profile fields, exact timestamps, and reverse mappings. The resulting release supports longitudinal content and diffusion analyses while reducing disclosure and platform-content redistribution risks. To request access to raw data for collaborative research under ethical and legal safeguards, contact fubu.dataset@gmail.com.
We introduce IO Factory, an AI-driven framework for simulating information and influence campaigns as fully integrated, traceable processes. The threat of digital manipulation now extends beyond persuasive text from individual language models to AI swarms, i.e., persistent groups of coordinated agents that adapt to platform feedback and disguise organized campaigns as ordinary social interaction. Because such campaigns cannot be identified from isolated messages alone, they must be analyzed across a continuous spectrum of planning, platform action, exposure, interpretation, measurement, and adaptation. IO Factory represents this process inside a controlled simulated platform, linking actor roles, platform actions, exposure records, structured model-based evaluations, and configured changes in the simulated population. We implement the architecture and evaluate it across configurations of up to 100,000 agents. The results show that IO Factory executes campaign timelines at scale and produces inspectable evidence of exposure and measured movement in configured belief variables. By recording the actors, objectives, action constraints, exposure paths, and measurement rules used in each run, IO Factory supports reproducible research and red-team analysis of coordinated influence.
Shortly after the first COVID-19 cases became apparent in December 2019, rumors spread on social media suggesting a connection between the virus and the 5G radiation emanating from the recently deployed telecommunications network. In the course of the following weeks, this idea gained increasing popularity, and various alleged explanations for how such a connection manifests emerged. Ultimately, after being amplified by prominent conspiracy theorists, a series of arson attacks on telecommunication equipment followed, concluding with the kidnapping of telecommunication technicians in Peru. In this paper, we study the spread of content related to a conspiracy theory with harmful consequences, a so-called Digital Wildfire (DW). In particular, we investigate the 5G and COVID-19 DW on Twitter before, during, and after its peak in April and May 2020. For this purpose, we examine the community dynamics in complex temporal interaction networks underlying Twitter user activity. We assess the evolution of this particular DW by appropriately defining the temporal dynamics of communication in communities within social networks. We observe that, for this specific DW, the number of interactions of the users participating in the DW, as well as the size of the engaged communities, both exhibit visual patterns suggestive of power-law distributions, consistent with other social networks, though based on exploratory inspection rather than formal statistical tests. Moreover, our exploratory analysis elucidates the possibility of conceptualizing the phases of a DW, as per established literature. We identify one such phase as a potential critical shift, marked by a shift from sporadic tweets to a global spreading event, highlighting patterns suggestive of dramatic scaling in misinformation propagation, used heuristically without quantitative physical parameters. Additionally, we argue that patterns suggest the observed shift is associated with influential users, who appear to amplify the spread of misinformation, though causality is exploratory. Lastly, our data suggest that the characteristics of such events could contribute to prediction models, at least in some instances. From this data, we hypothesize that monitoring minor peaks in user interactions, which precede the critical phase culminating in real-world consequences, could serve as an early warning system, aiding in the anticipation and potentially the mitigation of DWs.
This article examines the role of social media, specifically Twitter, in the amplification of violent divisive discourse during Ethiopia's Tigray conflict (2020-2022), introducing the concept of Hatetags to analyze the transforming of hashtags into vehicles of digital conflict. Through a mixed-method approach combining explorative digital methods, qualitative content analysis and qualitative interviews, the study investigates how hashtags evolved into emotionally charged communicative tools that mobilized polarized publics, reinforced ethnic boundaries, and contributed to conflict escalation. Situating Hatetags within broader frameworks of contemporary conflict studies, the article highlights how digital infrastructures do not merely reflect but actively shape conflict dynamics. It demonstrates how platform algorithms and state-diaspora interactions fuelled the normalization of hate speech and dehumanizing rhetoric. Case studies of hashtags like #junta and #cockroach reveal how repeated use and the interaction with visual memes reinforced hostility and legitimated violence. The study argues for a reconceptualization of digital platforms as active agents in conflict communication, underscoring the need for context-sensitive moderation in fragile media ecosystems. The concept of Hatetags offers a critical lens for understanding the evolving nexus between digital media and political violence.
The study combines domain expertise and computational community detection to uncover what role citizen journalists and social media platforms play in mediating the dynamics of conflict in Mali. Under conditions of the growing conflict in Mali, citizen journalists are opening Twitter (rebranded as X) accounts to stay updated and tweet about the ongoing socio-political tensions, chronicling life in a conflict-ravaged context. This article conceptualizes the rapid reliance on Twitter among citizen journalists consisting of bloggers, activists, government officials and NGO’s as a form of networked conflict and networked journalism. Networked journalism emerges as professional journalists adopt tools and techniques used by nonprofessionals (and vice versa) to gather and disseminate information while networked conflict involves the consequential and intricate relationship between social media and conflict in the Sahel region of Africa. Our findings show that Twitter is a source of action that promotes and mediates conflict, which exposes users to conflict-related content. The findings also show that what accounts for citizen journalism in a conflict setting is vague as those with access to Twitter and as such, the presumed ability to influence the narrative, unequivocally consider themselves citizen journalists.
While research on disinformation in Africa is starting to gain attention, only a few of these studies have sought to uncover its connection to conflict contexts, an important area we seek to draw attention to with this study. Furthermore, barely any studies have uncovered journalists and other media workers' perceptions of disinformation, particularly in countries dominated by conflict. This article uses participatory action research, specifically focus-group interviews with over 30 Malian and Ethiopian journalists to uncover rarely discussed perceptions and contestations emerging from authoritarian contexts dominated by both disinformation and conflict. Our aim is to make a significant contribution in the field of journalism studies by studying the rapid rise of disinformation within a context marked by both authoritarianism and conflict. We believe our findings, which reveal previously unknown evidential pointers to journalists, journalism educators and other news media workers' perspectives on disinformation in a non-democratic contexts potentially have wider applicability given disinformation is a problem that even democratic countries are also currently grappling with. Our findings show that there are varying actors in the disinformation landscape. Also, conflicting notions on disinformation exist while there is a noticeable lack of established truth-telling institutions. Besides, ethnicity influences people's perceptions of truth.
String alignment algorithms are an essential tool for understanding DNA and protein sequences. They demand substantial computation in real-world applications, and are thus a prime target for hardware acceleration. However, GPUs struggle to provide sufficient acceleration. Meanwhile, the recent MIMD-capable AI accelerators such as the Graphcore Intelligence Processing Unit (IPU) have become technologically viable. In this paper we present iPuma, a new implementation of Smith-Waterman sequence alignment for the IPU, which offers generalized short and medium length, one-to-one, and many-to-many high-throughput alignments for both DNA and protein sequences. iPuma is integrated into two bioinformatics pipelines, MetaHipMer2 and PASTIS. On protein datasets, iPuma shows speedups of 2.7 × and 1.6 × over state-of-the-art GPU and CPU implementations, respectively. We test the scalability on up to 64 IPUs, attaining a peak scoring performance of 1763 GCUPS for protein and 1168 GCUPS for DNA sequences.
This paper presents GraphMa, a framework aimed at enhancing pipeline-oriented computation for graph processing. GraphMa integrates the principles of pipeline computation with graph processing methodologies to provide a structured approach for analyzing and processing graph data. The framework defines a series of computational abstractions, including computation as type, higher-order traversal, and directed data-transfer, which collectively facilitate the decomposition of graph operations into modular functions. These functions can be composed into pipelines, supporting the systematic development of graph algorithms. For this paper, our focus lies in particular on the capability to implement the well- established computational models for graph processing within the proposed framework. In addition, the paper discusses the design of GraphMa, its computational models, and the implementation de- tails that illustrate the framework's application to graph processing tasks.
Shortly after the first COVID-19 cases became apparent in December 2020, rumors spread on social media suggesting a connection between the virus and the 5G radiation emanating from the recently deployed telecommunications network. In the course of the following weeks, this idea gained increasing popularity, and various alleged explanations for how such a connection manifests emerged. Ultimately, after being amplified by prominent conspiracy theorists, a series of arson attacks on telecommunication equipment follows, concluding with the kidnapping of telecommunication technicians in Peru. In this paper, we study the spread of content related to a conspiracy theory with harmful consequences, a so-called digital wildfire. In particular, we investigate the 5G and COVID-19 misinformation event on Twitter before, during, and after its peak in April and May 2020. For this purpose, we examine the community dynamics in complex temporal interaction networks underlying Twitter user activity. We assess the evolution of such digital wildfires by appropriately defining the temporal dynamics of communication in communities within social networks. We show that, for this specific misinformation event, the number of interactions of the users participating in a digital wildfire, as well as the size of the engaged communities, both follow a power-law distribution. Moreover, our research elucidates the possibility of quantifying the phases of a digital wildfire, as per established literature. We identify one such phase as a critical transition, marked by a shift from sporadic tweets to a global spread event, highlighting the dramatic scaling of misinformation propagation.
With the expansion of mobile communications infrastructure, social media usage in the Global South is surging. Compared to the Global North, populations of the Global South have had less prior experience with social media from stationary computers and wired Internet. Many countries are experiencing violent conflicts that have a profound effect on their societies. As a result, social networks develop under different conditions than elsewhere, and our goal is to provide data for studying this phenomenon. In this dataset paper, we present a data collection of a national Twittersphere in a West African country of conflict. While not the largest social network in terms of users, Twitter is an important platform where people engage in public discussion. The focus is on Mali, a country beset by conflict since 2012 that has recently had a relatively precarious media ecology. The dataset consists of tweets and Twitter users in Mali and was collected in June 2022, when the Malian conflict became more violent internally both towards external and international actors. In a preliminary analysis, we assume that the conflictual context influences how people access social media and, therefore, the shape of the Twittersphere and its characteristics. The aim of this paper is to primarily invite researchers from various disciplines including complex networks and social sciences scholars to explore the data at hand further. We collected the dataset using a scraping strategy of the follower network and the identification of characteristics of a Malian Twitter user. The given snapshot of the Malian Twitter follower network contains around seven million accounts, of which 56,000 are clearly identifiable as Malian. In addition, we present the tweets. The dataset is available at: https://osf.io/mj2qt/
The COVID-19 pandemic has been accompanied by a surge of misinformation on social media which covered a wide range of different topics and contained many competing narratives, including conspiracy theories. To study such conspiracy theories, we created a dataset of 3495 tweets with manual labeling of the stance of each tweet w.r.t. 12 different conspiracy topics. The dataset thus contains almost 42,000 labels, each of which determined by majority among three expert annotators. The dataset was selected from COVID-19 related Twitter data spanning from January 2020 to June 2021 using a list of 54 keywords. The dataset can be used to train machine learning based classifiers for both stance and topic detection, either individually or simultaneously. BERT was used successfully for the combined task. The dataset can also be used to further study the prevalence of different conspiracy narratives. To this end we qualitatively analyze the tweets, discussing the structure of conspiracy narratives that are frequently found in the dataset. Furthermore, we illustrate the interconnection between the conspiracy categories as well as the keywords.
The study of online social networks has become a major topic of research within the last decade, and many aspects of the networks and the behavior of their users have been investigated. The majority of research efforts has been directed at Twitter, which grants limited data accesses to researchers and provides detailed information. However, recently, important social and economic phenomena such as WallStreetBets or Antiwork have originated on Reddit, which has thus become an important field of investigation in its own right, and, due to its open nature, all Reddit data is available to study. As a consequence, in contrast to Twitter, where it is difficult to obtain large amounts of data, the main challenge of researching Reddit is to handle the vast amounts of data that are freely available. Here, we present the Reddit Dataset Stream Pipeline (RDSP), a simple and efficient parallel system based on Akka Streams that is capable of processing the entire Reddit dataset. We demonstrate how to build massive temporal graphs between subreddits from a parallel streamed dataset. We investigate the generated graphs and present experimental results. Moreover, we publish both the datasets as well as the codebase in order to invite researchers from different fields to contribute and profit from this work.
Minimum weighted vertex cover is the NP-hard graph problem of choosing a subset of vertices incident to all edges such that the sum of the weights of the chosen vertices is minimum. Previous efforts for solving this in practice have typically been based on search-based iterative heuristics or exact algorithms that rely on reduction rules and branching techniques. Although exact methods have shown success in solving instances with up to millions of vertices efficiently, they are limited in practice due to the NP-hardness of the problem. We present a new hybrid method that combines elements from exact methods, iterative search, and graph neural networks (GNNs). More specifically, we first compute a greedy solution using reduction rules whenever possible. If no such rule applies, we consult a GNN model that selects a vertex that is likely to be in or out of the solution, potentially opening up for further reductions. Finally, we use an improved local search strategy to enhance the solution further. Extensive experiments on graphs of up to a billion edges show that the proposed GNN-based approach finds better solutions than existing heuristics. Compared to exact solvers, the method produced solutions that are, on average, 0.04% away from the optimum while taking less time than all state-of-the-art alternatives.
Online social networks are ubiquitous, have billions of users, and produce large amounts of data. While platforms like Reddit are based on a forum-like organization where users gather around topics, Facebook and Twitter implement a concept in which individuals represent the primary entity of interest. This makes them natural testbeds for exploring individual behavior in large social networks. Underlying these individual-based platforms is a network whose “friend” or “follower” edges are of binary nature only and therefore do not necessarily reflect the level of acquaintance between pairs of users. In this paper,we present the network of acquaintance “strengths” underlying the German Twittersphere. To that end, we make use of the full non-verbal information contained in tweet–retweet actions to uncover the graph of social acquaintances among users, beyond pure binary edges. The social connectivity between pairs of users is weighted by keeping track of the frequency of shared content and the time elapsed between publication and sharing. Moreover, we also present a preliminary topological analysis of the German Twitter network. Finally, making the data describing the weighted German Twitter network of acquaintances, we discuss how to apply this framework as a ground basis for investigating spreading phenomena of particular contents.
Artificial neural networks have been used for a multitude of regression tasks, and their descendants have expanded the domain to many applications such as image and speech recognition, filtering of social networks, and machine translation. While conventional and recurrent neural networks work well on data represented in Euclidean space, they struggle with data in non-Euclidean space. Graph Neural Networks (GNN) expand recurrent neural networks to directly process sparse representations of graphs, but they are computationally expensive, which invites the use of powerful hardware accelerators. In this paper, we investigate the viability of the Graphcore Intelligence Processing Unit (IPU) for efficient implementation of Spatio-Temporal Graph Convolutional Networks. The results show that IPUs are well suited for this task.
The COVID-19 pandemic has severely affected the lives of people worldwide, and consequently, it has dominated world news since March 2020. Thus, it is no surprise that it has also been the topic of a massive amount of misinformation, which was most likely amplified by the fact that many details about the virus were not known at the start of the pandemic. While a large amount of this misinformation was harmless, some narratives spread quickly and had a dramatic real-world effect. Such events are called digital wildfires. In this paper we study a specific digital wildfire: the idea that the COVID-19 outbreak is somehow connected to the introduction of 5G wireless technology, which caused real-world harm in April 2020 and beyond. By analyzing early social media contents we investigate the origin of this digital wildfire and the developments that lead to its wide spread. We show how the initial idea was derived from existing opposition to wireless networks, how videos rather than tweets played a crucial role in its propagation, and how commercial interests can partially explain the wide distribution of this particular piece of misinformation. We then illustrate how the initial events in the UK were echoed several months later in different countries around the world.
The FakeNews Detection task at MediaEval 2022, running for the third time as part of the challenge, focuses on the detection of misinformation tweets and their spreaders. Like in the 2021 task, conspiracy theories related to COVID-19 in nine different categories have to be detected, along with the authors stance towards them. For the 2022 challenge, the size of the dataset has approximately doubled. Furthermore, we also provide a large interaction graph along with vertex features derived from the same Twitter dataset in which misinformation spreaders should be classified. As a final subtask, participants are asked to combine text and graph information to refine their classifications. This paper describes the tasks, including use case and motivation, challenges, the dataset with ground truth, the required participant runs, and the evaluation metrics.
The FakeNews: Corona Virus and Conspiracies Multimedia Analysis task, running for the second time as part of MediaEval 2021, focuses on the classification of tweet texts aiming detection of fast-spreading misinformation. Task of this year extends the number of target conspiracy theories and introduces new challenges in terms of analysis complexity of the imbalanced dataset. This paper de-scribes the task, including use case and motivation, challenges, the dataset with ground truth, the required participant runs, and the evaluation metrics.
The COVID-19 pandemic has been accompanied by a flood of misinformation on social media, which has been labeled an "infodemic". While a large part of such fake news is ultimately inconsequential, some of it has the potential to real-world harm, but due to the massive amount of social media contents, it is impossible to find this misinformation manually. Thus, conventional fact-checking can typically only counteract misinformation narratives after they have gained significant traction. Only automated systems can provide warnings in advance. However, the automatic detection of misinformation narratives is very challenging since the texts that spread misinformation may be short messages on Twitter. They may also transmit misinformation by implication rather than by stating counterfactual information outright, and satirical messages complicate the issue further. Thus, there is a need for highly sophisticated detection systems. In order to support their development, we created substantial ground truth data by human annotation. In this paper, we present a dataset that deals with a specific piece of misinformation: the idea that the COVID-19 pandemic is causally connected to the 5G wireless network. We selected more than 10,000 tweets that deal with COVID-19 and 5G and labeled them manually, distinguishing between tweets that propagate the specific 5G misinformation, those that spread other conspiracy theories, and tweets that do neither. We provide the human-annotated dataset along with an additional large-scale automatically (by using the human-annotated dataset as the training set) labelled dataset consist of more than 100,000 tweets.