
The emergence of social media has led memes to become a powerful mode of communication, blending text, images, and emojis. However, this surge in meme usage has also seen a rise in offensive material. With manual content moderation proving impractical due to the sheer volume of data, there's a pressing need for automated methods to identify harmful memes. Yet, existing research predominantly targets high-resource languages such as English, neglecting low-resource ones like Nepali. To bridge this gap, we introduce the first Nepali meme dataset annotated for hate speech and sentiment. Our contributions are threefold: (1) We create and release NeMeme, a unique dataset featuring Nepali and code-mixed Nepali memes (combining Nepali and English). (2) We evaluate NeMeme using cutting-edge unimodal and multimodal models to establish initial performance benchmarks. (3) We introduce MemeNePAL, a novel multimodal framework employing prompt-assisted learning to effectively categorize Nepali memes. MemeNePAL overcomes the shortcomings of prior state-of-the-art (SOTA) techniques, which were designed for high-resource languages and struggle with Nepali's linguistic differences and cultural subtleties. This work not only promotes inclusivity in content moderation research but also aligns with UN Sustainable Development Goals such as promoting well-being, reducing inequalities, and fostering peace. We adhere to FAIR principles by making the dataset publicly available.
This study examines how individual music listening behaviors evolved during the COVID-19 lockdowns in France, focusing on both listening volumes and rhythms. We combine passively collected individual listening history data, provided by a music streaming service and covering the 2019-2023 period, with survey data collected from the same users (n ≈ 10000). Using the Dynamic Time Warping method, we develop a typology of listening trajectories during the first lockdown. The results reveal significant and heterogeneous changes in listening behavior, with approximately one-third of respondents experiencing a significant decrease in listening volume, while a quarter experienced an increase. We then analyze the evolution of the intervals between consecutive music listening sessions — so called inter-session times — to assess disruptions in individual listening rhythms. We uncover an unprecedented shift in the listening rhythms at the onset of first lockdown, reflecting varying degrees of disruption in daily life rhythms. For half of the individuals this disruption lasted more than four weeks. Finally we show that age, educational attainment and household structure unevenly influence the reorganization of music listening activity during this period, shedding light on the social differentiations at work in the reorganization of an ordinary activity during this crisis period.
Ensuring the online safety of youth has motivated research towards the development of machine learning (ML) methods capable of accurately detecting social media risks after-the-fact. However, for these detection models to be effective, they must proactively identify high-risk scenarios (e.g., sexual solicitations, cyberbullying) to mitigate harm. This `real-time' responsiveness is a recognized challenge within the risk detection literature. Therefore, this paper presents a novel two-level framework that first uses reinforcement learning to identify conversation stop points to prioritize messages for evaluation. Then, we optimize state-of-the-art deep learning models to accurately categorize risk priority (low, high). We apply this framework to a time-based simulation using a rich dataset of 23K private conversations with over 7 million messages donated by 194 youth (ages 13-21). We conducted an experiment comparing our new approach to a traditional conversation-level baseline. We found that the timeliness of conversations significantly improved from over 2 hours to approximately 16 minutes with only a slight reduction in accuracy (0.88 to 0.84). This study advances real-time detection approaches for social media data and provides a benchmark for future training reinforcement learning that prioritizes the timeliness of classifying high-risk conversations.
This study delves into the mechanisms that spark user curiosity driving active engagement within public Telegram groups. By analyzing approximately 6 million messages from 29,196 users across 409 groups, we identify and quantify the key factors that stimulate users to actively participate (i.e., send messages) in group discussions. These factors include social influence, novelty, complexity, uncertainty, and conflict, all measured through metrics derived from message sequences and user participation over time. After clustering the messages, we apply explainability techniques to assign meaningful labels to the clusters. This approach uncovers macro categories representing distinct curiosity stimulation profiles, each characterized by a unique combination of various stimuli. Social influence from peers and influencers drives engagement for some users, while for others, rare media types or a diverse range of senders and media sparks curiosity. Analyzing patterns, we found that user curiosity stimuli are mostly stable, but, as the time between the initial message increases, curiosity occasionally shifts. A graph-based analysis of influence networks reveals that users motivated by direct social influence tend to occupy more peripheral positions, while those who are not stimulated by any specific factors are often more central, potentially acting as initiators and conversation catalysts. These findings contribute to understanding information dissemination and spread processes on social media networks, potentially contributing to more effective communication strategies.
Manually annotating data for computational social science tasks can be costly, time-consuming, and emotionally draining. While recent work suggests that LLMs can perform such annotation tasks in zero-shot settings, little is known about how prompt design impacts LLMs' compliance and accuracy. We conduct a large-scale multi-prompt experiment to test how model selection (GPT-4o, GPT-3.5, PaLM2, and Falcon7b) and prompt design features (definition inclusion, output type, explanation, and prompt length) impact the compliance and accuracy of LLM-generated annotations on four highly relevant and diverse CSS tasks (toxicity, sentiment, rumor stance, and news frames). Our results show that LLM compliance and accuracy are prompt-dependent. For instance, prompting for numerical scores instead of labels reduces all LLMs' compliance and accuracy. Concise prompts can significantly reduce prompting costs but also lead to lower accuracy on tasks like toxicity. Furthermore, minor prompt changes like asking for an explanation can cause large changes in the distribution of LLM-generated labels. By assessing the impact of prompt design on the quality and distribution of LLM-generated annotations, this work serves as both a practical guide and a warning for using LLMs in CSS research.
Cannabis use is on the rise, driven by relaxing legal regulations and declining perceptions of harm. This trend, coupled with the increasing reliance on social media for health-related information, has sparked interest in cannabis use during pregnancy (CanPreg). This study examines online discourse about CanPreg on Twitter, analyzing 53,183 unique tweets from 32,744 users in the USA and Canada between 2012 and 2021. We investigate the spatio-temporal distribution of CanPreg discussions, key topical contexts within these conversations, and their correlations with socioeconomic and health indicators. The analysis reveals regional differences, with a relatively higher interest in CanPreg discussions in Canada compared to the USA. The online discourse is primarily focused on research, alongside criticism, personal experiences, queries, news sharing, and advertisements. Additionally, correlations between CanPreg tweet activity, poverty rates, and mental health metrics suggest a connection between online discussions and real-world behaviors. This study highlights the role of social media in health communication and provides insights to inform targeted intervention strategies.
People are increasingly avoiding the news in a phenomenon termed news avoidance. Prior work has found that many people are not following political news, with the median American consuming zero articles per year from traditional news outlets. This behavior has troubling implications, as a well-informed public is typically considered essential to a functioning democracy. However, it is possible that people are accidentally picking up political information through entertainment or "soft news" outlets. The subject of this paper is People Magazine, a soft news outlet selected for its popularity (with an average of 187.1 million monthly visits). To understand how political news is covered by entertainment and soft news outlets, we make the following contributions: (1) we propose two potential and complementary frameworks for differentiating hard news content from soft news content; (2) we collect a large dataset of articles published on People Magazine's website; (3) we apply our frameworks to this dataset and contribute an analysis of the political content offered by a soft news outlet; (4) we evaluate our framework on two additional platforms. Together, these contributions additively result in the first exploration of the political content of individual news stories - all prior work has labeled entire outlets as hard or soft. As people increasingly avoid outlets with hard news, our work shows that they may still encounter political information through soft news outlets, a finding that has critical implications for political communication.
TikTok has emerged as a leading social media platform with increasing relevance for the consumption and distribution of news, especially for younger age groups. Despite its growing relevance, analyses of how traditional news outlets produce content for the platform and maintain journalistic news values are limited. Moreover, there are few large-scale datasets that are suitable for tracing larger journalistic trends and developing approaches for the automated annotation of multimodal content. This paper addresses these gaps and introduces “News on TikTok,” an annotated dataset of 8,623 TikTok videos published by 18 major German-speaking news outlets in 2023. Combining metadata with human-annotated data of 25 variables (incl. the presence of visual, auditory, and interactive elements, and journalistic news values), our dataset makes three significant contributions: First, it enables extensive analyses of news characteristics on TikTok. Second, it provides ground truth data to develop and validate automated tools for multimodal content analyses. Third, it offers a comprehensive guide for generating datasets with similar research interests.
We are witnessing a significant shift in social media platforms; we are transitioning from chronological social media feeds to feeds that are driven by AI recommendation systems. While the main goal of AI recommendation systems is to suggest engaging content to users, there are also some associated risks: AI recommendation systems can promote extreme content, causing negative consequences like online polarization and user radicalization. Overall, there is a pressing need to design powerful techniques that allow us to audit AI recommendation systems. Motivated by this, our work introduces ClipMind, a scalable and generalizable framework using advanced AI models to audit these recommendation algorithms on short-format video platforms like TikTok and YouTube Shorts. We demonstrate the merits of our framework by collecting social media feeds from TikTok. Our analysis shows that TikTok’s recommendation algorithm increasingly recommends similar videos when a user expresses interest in mainstream topics like Food and Beauty Care. On the other hand, by investigating niche interests (War and Mental Health), we find no evidence of informational rabbit holes of extreme content on TikTok. Our work contributes to efforts that leverage AI for social good, as our framework can be used by several interested stakeholders, including users, social media platforms, regulators, and researchers, to understand and audit video-based algorithmic recommendations.
This study investigates how viewers and creators on Japanese YouTube channels progress towards conspiracy theories. By categorizing channels based on ideological or financial motives and analyzing engagement metrics such as views, likes, and comments, we find that channels driven by monetization, particularly Monetized Conspiracists, promote conspiracy theories more vigorously. This indicates that financial incentives are a crucial factor in the proliferation of such content. Channels that package conspiracy theories in formats like entertainment or spirituality serve as gateways, facilitating viewers' progression towards more extreme conspiracy-laden content. Understanding these pathways is vital for crafting strategies to counteract the spread of conspiracy theories on social media.
Online communities on Reddit are a popular choice among people with opioid use disorder (OUD) to seek information on drug use, withdrawal symptoms, and recovery. LLM-powered chatbots (e.g., ChatGPT) are widely being adopted as question-answer systems for health-related queries. However, such online health information seeking could potentially be hindered by myths and misinformation on OUD, misleading or causing genuine harm to people with OUD. In this work, we examine the prevalence of 5 OUD-related myths, on treatment models and patient characteristics, within human- (taken from Reddit) and LLM-generated responses to queries on OUD. We further explore the framing strategies used within responses (both human- and LLM-generated) promoting and countering the myths. We found that all 5 myths were more widespread within human-generated responses. In addition, myth-promoting responses adopted trustworthy and authoritative framings, compared to knowledge-imparting linguistic cues within those countering the myths. Our work offers recommendations to reduce online OUD misinformation.
Authentic information is vital for a society's ability to make rational decisions. Fabricated and manipulative information can be harmful to society as seen in cases of threatening events that were consequences of foreign propaganda and radical ideologies. While past research has studied dis- and misinformation on social media platforms, the study of propaganda has received much less attention. This study explores the sharing intentions of propaganda on social media platforms and develops an intervention to help detect it. In a randomized controlled trial setting, we added indicators to social media posts that used propaganda techniques to advance an agenda, including techniques that rely on fallacious reasoning, emotional rather than logical reasoning, etc. We then asked our participants (n=1,187) about their intention to engage with these posts. We found that participants were significantly (2.4 times) less likely to share these posts with indicators. We also found that participants’ political affiliation moderated their sharing intentions. We believe our findings provide valuable insights for the study of propaganda on social media platforms.
This study presents an analysis of digital polarization on the topic of the Hijab by examining YouTube comments in Arabic. Employing a novel dataset of around 10K annotated comments, this research investigates the digital discourse using seven labels: Stance, Use of Sarcasm, Argumentation, Cordiality, Offensiveness, Hopefulness, and Apparent Gender of Commenters. The findings reveal significant insights into gender dynamics and the prevalence of specific rhetorical strategies within the debate. This study contributes to the broader field of polarization and argument mining, offering a unique lens on the intersection of digital culture and societal issues in the Arab context.
This paper studies a critical problem of explainable public health misinformation detection on social media, where clear explanations are essential for enhancing user understanding and trust, surpassing the limitations of black-box misinformation detection results. To tackle this problem, there is a growing trend of leveraging collective intelligence from diverse intelligence sources, such as deep neural networks (DNNs), human intelligence, and large language models (LLMs). However, integrating hybrid intelligence from different sources remains a challenge: DNNs excel in accurate and efficient classification, crowd workers provide contextual understanding and readable explanations, and LLMs offer extensive domain knowledge and advanced language generation. Moreover, current crowdsourcing and human-AI collaboration methods mainly focus on aggregating misinformation detection labels using traditional measures like consistency, often overlooking more complex and challenging inputs like textual explanations. We propose SynthX, a collective intelligence framework that incorporates a holistic prompting design to harness the language and reasoning capabilities of LLMs for synthesizing diverse detection and explanation results. It also integrates a novel estimation theory-LLM hybrid approach to assess the varying reliability of detection results from different intelligence sources. Our evaluation on a real-world social media misinformation dataset demonstrates that SynthX consistently outperforms a rich set of state-of-the-art baselines in both detection accuracy and explanation quality.
This paper proposes the development of a WCAG-compliant chatbot capable of generating multimodal content to enhance usability for all users. While LLM-based chatbots excel in generating varied responses, they often struggle with ambiguous or incomplete queries, leading to misaligned outputs. We introduce a framework that formulates domain-specific, persona-driven follow-up questions to clarify ambiguities, utilizing knowledge graphs and human feedback. The system refines queries before generating responses by employing a domain-Specific Multilayer Hierarchical Relational Graph (MHRG) to model user intent. Our preliminary evaluations indicate that the Accessibility Bot improves response relevance and quality as compared to existing techniques.
In the current social media landscape, the study of influence propagation and consensus formation has gained prominence. While user interactions like retweeting are apparent, the underlying pathways of influence often remain hidden and complex. This study proposes a novel network called Latent Influence Network (LIN), which advances the analysis of influence on social media. LIN's architecture and the process of parameter selection are meticulously discussed within the comprehensive Latent Influence Detection Framework (LIDET). Based on the user's behavior label, LIN identifies the optimal network configuration, revealing more accurate influence patterns. We applied the LIDET framework to four diverse datasets, each demonstrating substantial improvements in influence pattern recognition over traditional network models. Specifically, in a case study on a COVID-19 dataset, LIN achieved a classification accuracy of 99%, significantly outperforming conventional methods. These findings underscore the utility of LIN in capturing the dynamics of influence and enhancing our understanding of opinion formation on social media.
The increasing frequency of mass shootings in the United States has become alarmingly common, prompting discussions about gun control. While gun control in the US involves complex legal issues, cultural factors---particularly ``gun culture''---play a significant but often overlooked role. Although the role of social media in shaping culture is well-documented, the intersection of gun culture and fringe online communities, like 4chan, remains unclear. This gap is particularly concerning given the rise in mass shootings and the online radicalization of some shooters. To address this gap, we explore gun culture on /k/, 4chan's weapons board. More specifically, we employ various NLP techniques to analyze over 4M posts on /k/ and contextualize the discussion within the broader body of theoretical framework of gun culture. Our findings reveal that discussions on /k/ cover a wide array of topics, with a significant focus on law-related discussions---over 17% of gun-related conversations on /k/ revolve around legal matters. Additionally, our analysis uncovers the presence of extreme viewpoints surrounding firearms, often manifesting as gun fetishism. These insights can be valuable for a range of stakeholders including social media platform, in efforts to address content moderation and de-radicalization
Warning: This paper may contain triggering language for some readers, especially survivors of sexual violence. Survivors of sexual violence sometimes share their experiences on social media, revealing their feelings and emotions and seeking advice. On platforms such as Reddit, some stories can be long---up to 40,000 characters. We posit that such long stories are demanding for helpers to read and respond to. Prior research has indicated that parts of these stories describing the incident, the effects on the poster, and advice requested by the poster are important. Highlighting those parts can draw helpers' attention toward key information and assist them in reading and responding to long stories. We first examine the stories posted on Reddit for the prevalence of these parts. Second, we develop a computational model to highlight these parts of a story. On ten-fold cross-validation of a dataset, our model achieves a macro F1 score of 0.82. In addition, we contribute METHREE, a dataset comprising 8,947 labeled sentences for these parts from Reddit stories. A survey of users who are helpers on some relevant subreddits shows that the parts highlighted by our tool represent important information and assist them while reading and responding to long stories. We find that these tool-generated highlights statistically significantly reduce the demandingness of long stories. Moreover, almost all helpers felt that highlighted stories are helpful and easier to read, understand, and respond to than nonhighlighted ones. In particular, on a 4-point Likert scale, there is about 0.7 point reduction in demandingess when stories were presented with highlights.
The fragmentation of social media challenges how we might efficiently and effectively identify, understand, and counter harmful content. Prior work establishes frameworks for measuring problematic narratives, evaluating harms, and leveraging interdisciplinary theories and findings to design mitigating solutions. However, there is little understanding of these phenomena outside of mainstream platforms. This is particularly concerning given that alt-tech users have been observed to include insurrectionists, active shooters, and other extremists – often driven from mainstream platforms due to deplatforming and content moderation . Our work aims to characterize online narratives across alt-tech platforms. In particular, we highlight how rumoring and conspiracy theory narratives in the context of the 2022 U.S. elections impact social dynamics and inspire collective action. We gather a unique dataset of over 7,000 social media posts from Gab, Gettr, Parler, and Truth Social from which we derive prevalent narratives using natural language processing techniques. We then examine how the platform, affect, and engagement differ across context through the lens of narrative, social identity, and mobilization potential using mixed methods. Findings from our analyses show variation between how narratives support social identity conceptions of power and mobilization potential.
Despite Arabic being one of the most widely spoken languages, there is a scarcity of available dialectal Arabic data. In this paper, we address this challenge by proposing a novel approach to data collection through the main use of video captions from TikTok, and other resources such as dictionaries and articles, resulting in the creation of the ArDia dataset. To the best of our knowledge, the ArDia dataset is the largest labeled dialectal Arabic dataset, containing over 900,000 examples, each labeled with its respective dialect. We further leverage this dataset to pretrain transformer-based models, ArDiaBERT and ArDiaGPT. Due to a lack of research on the Arabic models, we present a comprehensive study of Arabic dialect identification using the ArDia dataset on the dialect identification task.