
Demographic classification is essential in fairness assessment in recommender systems or in measuring unintended bias in online networks and voting systems. Important fields like education and politics, which often lay a foundation for the future of equality in society, need scrutiny to design policies that can better foster equality in resource distribution constrained by the unbalanced demographic distribution of people in the country. We collect three publicly available datasets to train state-of-the-art classifiers in the domain of gender and caste classification. We train the models in the Indian context, where the same name can have different styling conventions (Jolly Abraham/Kumar Abhishikta in one state may be written as Abraham Jolly/Abishikta Kumar in the other). Finally, we also perform cross-testing (training and testing on different datasets) to understand the efficacy of the above models. We also perform an error analysis of the prediction models. Finally, we attempt to assess the bias in the existing Indian system as case studies and find some intriguing patterns manifesting in the complex demographic layout of the sub-continent across the dimensions of gender and caste.
The emergence of COVID-19 and its associated containment strategies, such as lockdowns and social distancing, are expected to impact mental health, which could be more severe among people with pre-existing mental health disorders. In this research, we aim to better understand the changes in mental health during the COVID-19 pandemic by analysing data from mental health-related communities. We have collected data from the 15th of February 2020 to the 15th of July 2020 and analysed these data using interaction, linguistic structure, and interpersonal awareness measures. Our findings show that early in the lockdown, individuals showed selflessness, solidarity, and low rates of seeking help, but they also showed a negative mental health state. Moreover, considering the importance of social support in mental illness, we also aim to explore what derives social support in mental health communities. We found that receiving high social support was hindered by the use more swearing and negative emotional or self-referent words. Furthermore, not receiving social support may push actual help seekers to repeat their posts, which may be considered spamming in Reddit forums. Hence, we investigated the characteristics of duplicate posts authored by real help seekers to build a spam classifier. Our investigation showed that actual help seekers tend to show different levels of mental health when they repeat their posts for seeking immediate help.
Online social media have become an important forum for exchanging political opinions. In response to COVID measures citizens expressed their policy preferences directly on these platforms. Quantifying political preferences in online social media remains challenging: The vast amount of content requires scalable automated extraction of political preferences – however fine grained political preference extraction is difficult with current machine learning (ML) technology, due to the lack of data sets. Here we present a novel data set of tweets with fine grained political preference annotations. A text classification model trained on this data is used to extract policy preferences in a German Twitter corpus ranging from 2019 to 2022. Our results indicate that in response to the COVID pandemic, expression of political opinions increased. Using a well established taxonomy of policy preferences we analyse fine grained political views and highlight changes in distinct political categories. These analyses suggest that the increase in policy preference expression is dominated by the categories pro-welfare, pro-education and pro-governmental administration efficiency. All training data and code used in this study are made publicly available to encourage other researchers to further improve automated policy preference extraction methods. We hope that our findings contribute to a better understanding of political statements in online social media and to a better assessment of how COVID measures impact political preferences.
Social media platforms like Twitter play a pivotal role in public debates. Recent studies showed that users online tend to join groups of like-minded peers, called echo chambers, in which they frame and reinforce a shared narrative. Such a polarized configuration may trigger heated debates and foment misinformation spreading. In this work, we explore the interplay between the systematic spreading of misinformation and the emergence of toxic conversations. We perform a thorough quantitative analysis on 3.3 M comments by more than 1 M unique users from 25 K conversations involving 60 news outlets active on Twitter from January 2020 to April 2022. By tagging the news triggering the conversation with a specific reliability score provided by an independent fact-checking organization (NewsGuard), we perform a network-based analysis of the structure of toxic conversations on Twitter. We find that users using toxic language are few and evenly distributed over the entire reliability score range, showing no significant evidence for the interplay between toxic speech and misinformation spreading.
Twitter was widely used during the 2020 U.S. election to disseminate claims of election fraud. As a result, a number of works have examined this phenomenon from a variety of perspectives. However, none of them focus on analyzing topics behind the general fraud claims and associating them with user communities. To fill this gap, we propose to uncover and characterize groups of Twitter users engaging in discussions about election fraud claims during the 2020 U.S. election using a large dataset that spans seven weeks during this period. To accomplish this, we model a sequence of co-retweet networks and employ a backbone extraction method that controls for inherent traits of social media applications, particularly, user activity levels and the popularity of tweets (which together generate many spurious edges in the network), thus allowing us to reveal topics of tweets that lead users to retweet them. After extracting the backbones, we identify user groups representative of the communities present in the network backbones and finally analyze the topics behind the retweeted tweets to understand how they contributed to the spread of fraud claims at that time. Our main results show that (i) our approach uncovers better-structured communities than the original network in terms of users spreading discussions about fraud; and (ii) these users discuss 25 topics with specific psycholinguistic and temporal characteristics.
Social integration is known to be beneficial for mental health. However, it is not clear whether this applies to online as well as offline relationships. In this paper, we explore the association between online friendship and symptoms of depression among adolescents. We combine data from the popular social networking site with survey data on high school students ( $$N = 144$$ ) and find that integration into the online network is a protective factor against depression. We also find that not all online connections are equally important: friendship ties with students from the same schools are stronger associated with depression than outside ties. In addition to friendship ties, we explore the effect of online interaction (“likes”). Overall, our results suggest that online relationships are associated with depression as well as offline friendship. However, the effect of more distant online connections is limited, while immediate social environment and peer relationships at school are more important.
This work proposes a heterophily-based metric for quantifying polarization in social networks where multiple ideological, antagonistic communities coexist. This metric captures node-level polarization and is built on user’s affinity towards other communities rather than their own. Node-level values can then be aggregated at the community, network, or sub-network level, providing a more detailed map of polarization. We tested our metric on the Polblogs network, White Helmets Twitter interaction network with two communities and the VoterFraud2020 domain network with five communities. We also tested our metric on dK-random graphs to verify that it results in low polarization scores, as expected. Finally, we compared our metric with two widely used polarization measures: Guerra’s polarization index and RWC.
The 311 system has been deployed in many U.S. cities to manage non-emergency civic issues such as noise and illegal parking. To assess the performance of 311-mediated public service provision, researchers developed models based on execution time and the status of execution. However, research on user satisfaction suggests that the level of individuals’ perception is asymmetric with respect to the quality of services, because negative experiences have a stronger impact on people’s dissatisfaction than positive experiences do for satisfaction. Informed by the uneven nature of human satisfaction regarding positive and negative service quality, we propose an expectation-based model that measures the quality of public services by adapting the asymmetric function that reflects different perceptions of positive and negative experiences. Our preliminary analysis using the NYC 311 and Census data provides an initial assessment of the model’s validity.
With the rapid growth in the number of users on social networking sites (SNS), harmful content has also been fueled enormously over time. The purpose of harmful content is usually to mislead or harm or deceive an individual or a group of users. This study focuses on two types of harmful content: hate speech and misinformation. Alongside existing methods to detect misinformation and hate speech, users are still not well informed about the content description. This study proposes an interactive user interface, 'TweetInfo', for analysing information consumption towards harmful content by providing metainformation about social media posts. The study aims to explore the consumption of harmful content by users by providing an interactive user interface which flags harmful content and provides metadata about the post by doing content analysis and standard platform without the above features. The effectiveness of the proposed TweetInfo is measured using a user study conducted with 30 participants. TweetInfo reduces the spread of harmful content compared to the standard platform. While there are still interactions with bots on both platforms, verified accounts are also involved in propagating harmful content.
In this era of information explosion, deceivers use different domains or mediums of information to exploit the users, such as News, Emails, and Tweets. Although numerous research has been done to detect deception in all these domains, information shortage in a new event necessitates these domains to associate with each other to battle deception. To form this association, we propose a feature augmentation method by harnessing the intermediate layer representation of neural models. Our approaches provide an improvement over the self-domain baseline models by up to 6.60