Large-scale classification of social media content is a crucial technique for finding, studying, and analyzing misinformation in online social networks. Based on a manually labeled dataset of COVID-19 related conspiracy tweets, we train an NLP classifier and test methods for performing inference at scale using both GPUs as well as the Graphcore IPU AI accelerator.We apply our methods on a large dataset of about 2.5 billion tweets, demonstrating that using our methods, large scale inference is possible using affordable research infrastructures. Furthermore, we find that the IPU, due to its tile-centric design, is especially suited for such inference tasks.As a result, we obtain the AICO dataset of around 18 million tweets related to COVID-19 conspiracy theories that were posted between January 2020 and November 2021, which we make available for other researchers interested in studying the topic further under https://huggingface.co/datasets/Jlangguth/AICO.
Exposure-based therapies have shown promise in treating post-traumatic stress disorder (PTSD), but challenges exist in maintaining patient engagement and finding appropriate stimuli for graded exposure. Virtual reality (VR) technology has been used to enhance exposure therapy, but current software lacks customization and some patients remain treatment-resistant. A novel approach called multimodular motion-assisted memory desensitization and reconsolidation (3MDR) has the potential to solve some of the current limitations of VR-assisted exposure therapy. This study examines the efficacy of 3MDR treatment for individuals with treatment-resistant PTSD through a systematic review of relevant literature and clinical studies. Preliminary findings indicate promise for 3MDR in reducing PTSD symptoms, including emotional regulation and moral injury. However, further research with larger samples and controlled studies is needed to understand underlying mechanisms and validate these results. Moreover, this study highlights the importance of health-economic evaluations to assess costs and resource utilization associated with implementing 3MDR treatment in clinical services.
The COVID-19 pandemic has been accompanied by a surge of misinformation on social media which covered a wide range of different topics and contained many competing narratives, including conspiracy theories. To study such conspiracy theories, we created a dataset of 3495 tweets with manual labeling of the stance of each tweet w.r.t. 12 different conspiracy topics. The dataset thus contains almost 42,000 labels, each of which determined by majority among three expert annotators. The dataset was selected from COVID-19 related Twitter data spanning from January 2020 to June 2021 using a list of 54 keywords. The dataset can be used to train machine learning based classifiers for both stance and topic detection, either individually or simultaneously. BERT was used successfully for the combined task. The dataset can also be used to further study the prevalence of different conspiracy narratives. To this end we qualitatively analyze the tweets, discussing the structure of conspiracy narratives that are frequently found in the dataset. Furthermore, we illustrate the interconnection between the conspiracy categories as well as the keywords.
The COVID-19 pandemic has severely affected the lives of people worldwide, and consequently, it has dominated world news since March 2020. Thus, it is no surprise that it has also been the topic of a massive amount of misinformation, which was most likely amplified by the fact that many details about the virus were not known at the start of the pandemic. While a large amount of this misinformation was harmless, some narratives spread quickly and had a dramatic real-world effect. Such events are called digital wildfires. In this paper we study a specific digital wildfire: the idea that the COVID-19 outbreak is somehow connected to the introduction of 5G wireless technology, which caused real-world harm in April 2020 and beyond. By analyzing early social media contents we investigate the origin of this digital wildfire and the developments that lead to its wide spread. We show how the initial idea was derived from existing opposition to wireless networks, how videos rather than tweets played a crucial role in its propagation, and how commercial interests can partially explain the wide distribution of this particular piece of misinformation. We then illustrate how the initial events in the UK were echoed several months later in different countries around the world.
The COVID-19 pandemic has been accompanied by a flood of misinformation on social media, which has been labeled an "infodemic". While a large part of such fake news is ultimately inconsequential, some of it has the potential to real-world harm, but due to the massive amount of social media contents, it is impossible to find this misinformation manually. Thus, conventional fact-checking can typically only counteract misinformation narratives after they have gained significant traction. Only automated systems can provide warnings in advance. However, the automatic detection of misinformation narratives is very challenging since the texts that spread misinformation may be short messages on Twitter. They may also transmit misinformation by implication rather than by stating counterfactual information outright, and satirical messages complicate the issue further. Thus, there is a need for highly sophisticated detection systems. In order to support their development, we created substantial ground truth data by human annotation. In this paper, we present a dataset that deals with a specific piece of misinformation: the idea that the COVID-19 pandemic is causally connected to the 5G wireless network. We selected more than 10,000 tweets that deal with COVID-19 and 5G and labeled them manually, distinguishing between tweets that propagate the specific 5G misinformation, those that spread other conspiracy theories, and tweets that do neither. We provide the human-annotated dataset along with an additional large-scale automatically (by using the human-annotated dataset as the training set) labelled dataset consist of more than 100,000 tweets.
In the wake of the COVID-19 pandemic, a surge of misinformation has flooded social media and other internet channels, and some of it has the potential to cause real-world harm. To counteract this misinformation, reliably identifying it is a principal problem to be solved. However, the identification of misinformation poses a formidable challenge for language processing systems since the texts containing misinformation are short, work with insinuation rather than explicitly stating a false claim, or resemble other postings that deal with the same topic ironically. Accordingly, for the development of better detection systems, it is not only essential to use hand-labeled ground truth data and extend the analysis with methods beyond Natural Language Processing to consider the characteristics of the participant's relationships and the diffusion of misinformation. This paper presents a novel dataset that deals with a specific piece of misinformation: the idea that the 5G wireless network is causally connected to the COVID-19 pandemic. We have extracted the subgraphs of 3,000 manually classified Tweets from Twitter's follower network and distinguished them into three categories. First, subgraphs of Tweets that propagate the specific 5G misinformation, those that spread other conspiracy theories, and Tweets that do neither. We created the WICO (Wireless Networks and Coronavirus Conspiracy) dataset to support experts in machine learning experts, graph processing, and related fields in studying the spread of misinformation. Furthermore, we provide a series of baseline experiments using both Graph Neural Networks and other established classifiers that use simple graph metrics as features.
OPINION article Front. Psychol., 23 August 2021 | https://doi.org/10.3389/fpsyg.2021.662222
The COVID-19 pandemic constitutes a novel threat and traditional and new media provide people with an abundance of information and misinformation on the topic. In the current study, we investigated who tends to trust what type of mis/information. The data were collected in Norway from a sample of 405 participants during the first wave of COVID-19 in April 2020. We focused on three kinds of belief: the belief that the threat is overrated (COVID-threat skepticism), the belief that the threat is underrated (COVID-threat belief) and belief in misinformation about COVID-19. We studied sociodemographic factors associated with these beliefs and the interplay between attitudes to COVID-19, media consumption and prevention behavior. All three types of belief were associated with distrust in information about COVID-19 provided by traditional media and distrust in the authorities' approach to the pandemic. COVID-threat skepticism was associated with male gender, reduced news consumption since the start of the pandemic and lower levels of precautionary measures. Belief that the COVID-19 threat is underrated was associated with younger age, left-wing political orientation, increased news consumption during the pandemic and increased precautionary behavior. Consistent with the assumptions of the theory of planned behavior, individual beliefs about the seriousness of the COVID-19 threat predicted the extent to which individual participants adopted precautionary health measures. Both COVID-threat skepticism and COVID-threat belief were associated with endorsement of misinformation on COVID-19. Participants who endorsed misinformation tended to: have lower levels of education; be male; show decreased news consumption; have high Internet use and high trust in information provided by social media. Additionally, they tended to endorse multiple misinformation stories simultaneously, even when they were mutually contradictory. The strongest predictor for low compliance with precautionary measures was endorsement of a belief that the COVID-19 threat is overrated which at the time of the data collection was held also by some experts and featured in traditional media. The findings stress the importance of consistency of communication in situations of a public health threat.
We investigated the relation between emotional reactivity measured by Perth Emotional Reactivity Scale - Short Form (PERS-S) and trust in fictitious news stories on crime. In Study 1 we found on a sample of 508 older adults (M = 70.6 years) that their general positive and negative emotional reactivity was associated with trust in the presented misinformation, experienced negative emotions elicited by the news stories and willingness to share the news. For young adults in Study 2 (N = 186; M = 21.7) there was a weaker association between emotional reactivity and trust in misinformation, which involved only negative emotional reactivity. For both samples, trust in fictitious news stories was associated with trust in traditional and new media. There was no association between trust in fictitious news stories and the amount of news consumption and Internet use. Based on our findings, the focus on emotion control and critical reading seems to be important in the fight against misinformation.
The COVID-19 pandemic has been accompanied by a flood of misinformation on social media, which has been labeled an "infodemic". While a large part of such fake news is ultimately inconsequential, some of it has the potential to real-world harm, but due to the massive amount of social media contents, it is impossible to find this misinformation manually. Thus, conventional fact-checking can typically only counteract misinformation narratives after they have gained significant traction. Only automated systems can provide warnings in advance. However, the automatic detection of misinformation narratives is very challenging since the texts that spread misinformation may be short messages on Twitter. They may also transmit misinformation by implication rather than by stating counterfactual information outright, and satirical messages complicate the issue further. Thus, there is a need for highly sophisticated detection systems. In order to support their development, we created substantial ground truth data by human annotation. In this paper, we present a dataset that deals with a specific piece of misinformation: the idea that the COVID-19 pandemic is causally connected to the 5G wireless network. We selected more than 10,000 tweets that deal with COVID-19 and 5G and labeled them manually, distinguishing between tweets that propagate the specific 5G misinformation, those that spread other conspiracy theories, and tweets that do neither. We provide the human-annotated dataset along with an additional large-scale automatically (by using the human-annotated dataset as the training set) labelled dataset consist of more than 100,000 tweets.
The attempts to mitigate the unprecedented health, economic, and social disruptions caused by the COVID-19 pandemic are largely dependent on establishing compliance to behavioral guidelines and rules that reduce the risk of infection. Here, by conducting an online survey that tested participants’ knowledge about the disease and measured demographic, attitudinal, and cognitive variables, we identify predictors of self-reported social distancing and hygiene behavior. To investigate the cognitive processes underlying health-prevention behavior in the pandemic, we co-opted the dual-process model of thinking to measure participants’ propensities for automatic and intuitive thinking vs. controlled and reflective thinking. Self-reports of 17 precautionary behaviors, including regular hand washing, social distancing, and wearing a face mask, served as a dependent measure. The results of hierarchical regressions showed that age, risk-taking propensity, and concern about the pandemic predicted adoption of precautionary behavior. Variance in cognitive processes also predicted precautionary behavior: participants with higher scores for controlled thinking (measured with the Cognitive Reflection Test) reported less adherence to specific guidelines, as did respondents with a poor understanding of the infection and transmission mechanism of the COVID-19 virus. The predictive power of this model was comparable to an approach (Theory of Planned Behavior) based on attitudes to health behavior. Given these results, we propose the inclusion of measures of cognitive reflection and mental model variables in predictive models of compliance, and future studies of precautionary behavior to establish how cognitive variables are linked with people’s information processing and social norms.
We review the phenomenon of deepfakes, a novel technology enabling inexpensive manipulation of video material through the use of artificial intelligence, in the context of today’s wider discussion on fake news. We discuss the foundation as well as recent developments of the technology, as well as the differences from earlier manipulation techniques and investigate technical countermeasures. While the threat of deepfake videos with substantial political impact has been widely discussed in recent years, so far, the political impact of the technology has been limited. We investigate reasons for this and extrapolate the types of deepfake videos we are likely to see in the future.
We design a system for efficient in-memory analysis of data from the GDELT database of news events. The specialization of the system allows us to avoid the inefficiencies of existing alternatives, and make full use of modern parallel high-performance computing hardware. We then present a series of experiments showcasing the system's ability to analyze correlations in the entire GDELT 2.0 database containing more than a billion news items. The results reveal large scale trends in the world of today's online news.
ObjectiveProspective employers can nowadays easily access applicants' photos via Internet, for instance on professional and social networks or previous employers' websites. In our study, we investigated whether a facial expression in a picture affects evaluation of one's competence for a position where facial qualities are not crucial, namely a position of a software developer.MethodIn Study 1, both “models” and participants were employed in IT companies. The experiment followed a 3 x 3 x 2 design, with facial expression (smile, neutral, and thinking) and evaluator's experience in hiring as between‐subjects factors and gender of the model as a within‐subjects factor.Study 2 was a survey among software specialists where we investigated their awareness of the impact of applicant's face on the evaluation of his/her competence.ResultsWhen the models smiled, they were perceived as more competent than when they had a neutral expression. When models adopted a thinking pose, they were evaluated as the least competent. Fifty‐five percent of the sample was previously involved in hiring employees; the amount of hiring experience had no impact on this effect. Women were perceived as less competent than men and an interaction analysis revealed that this effect was driven by participants without prior experience in hiring. In Study 2, software specialists assigned a significant role in hiring decisions to the applicant's competent physical appearance, only 10% of participants thought that employers were hardly ever affected by the applicant's face.ConclusionFacial expression in a photo affects perceived competence of applicants for a position of a software developer regardless of evaluators' prior hiring experience for this type of job.
The FakeNews: Corona Virus and 5G Conspiracy task, running for the first time as part of MediaEval 2020, focuses on the classification of tweet texts and retweet cascades for the detection of fast-spreading misinformation, and therefore provides a lowthreshold introduction to natural language processing and graph analysis. This paper describes the task, including use case and motivation, challenges, the dataset with ground truth, the required participant runs, and the evaluation metrics.
Online social networks such as Facebook and Twitter are part of the everyday life of millions of people. They are not only used for interaction but play an essential role when it comes to information acquisition and knowledge gain. The abundance and detail of the accumulated data in these online social networks open up new possibilities for social researchers and psychologists, allowing them to study behavior in a large test population. However, complex application programming interfaces (API) and data scraping restrictions are, in many cases, a limiting factor when accessing this data. Furthermore, research projects are typically granted restricted access based on quotas. Thus, research tools such as scrapers that access social network data through an API must manage these quotas. While this is generally feasible, it becomes a problem when more than one tool, or multiple instances of the same tool, is being used in the same research group. Since different tools typically cannot balance access to a shared quota on their own, additional software is needed to prevent the individual tools from overusing the shared quota. In this paper, we present a proxy server that manages several researchers' data contingents in a cooperative research environment and thus enables a transparent view of a subset of Twitter's API. Our proxy scales linearly with the number of clients in use and incurs almost no performance penalties or implementation overhead to further layer or applications that need to work with the Twitter API. Thus, it allows seamless integration of multiple API accessing programs within the same research group.