Human annotation of data, including texts and images, is a bedrock of political science research. Yet, we often fail to consider how the identities of our labelers may systematically affect their annotations and our downstream applications. Collecting annotator demographic information, regardless of task type, can help us establish measurement validity and better appreciate variation in interrater reliability. We may also discover things about our topic that we did not previously appreciate. We demonstrate the benefits of collecting labeler characteristics with two annotation cases, one using images from the United States and the second using text from the Netherlands. For both cases on a range of tasks, we find that annotator gender and political identity are associated with significantly different annotations. We consider three main approaches to addressing labeler characteristic issues: adjusting labels based on labeler identity, weighting composite labels based on target population demographics, and intentionally modeling subgroup variation.
The Spanish Legislative Amendments Dataset provides comprehensive information on all amendments proposed to all executive bills introduced to the Spanish parliament from 1996 to 2023. At the bill level, the dataset provides information on the total number of proposed amendments to each bill (a total of 117.575), the parliamentary group proposing the amendments; and the main characteristics of each bill to which amendments were proposed. These characteristics include the bill title, date of proposal, the legislative committee that debated it, the bill topic, the type of law (if the bill is passed), the number of parliamentary appearances organized along the bill's debate, and the bill's regional and European content. This is the first systematic dataset of legislative amendments created in Spain. It is comprehensive (not sample-based) and was developed using web scraping and text-parsing techniques. The dataset is a valuable resource for researchers in political science, particularly those focusing on legislative studies, parliamentary behaviour, party dynamics, and political institutions. It is also relevant for teachers, interest groups, political parties, and other stakeholders (including public administrations and consulting firms) with an interest in, or affected by, legislation passed in the Spanish parliament.
Who shapes the issue-attention cycle of state legislators? Although state governments make critical policy decisions, data and methodological constraints have limited researchers' ability to study state-level agenda setting. For this article, we collect more than 122 million Twitter messages sent by state and national actors in 2018 and 2021. We then employ supervised machine learning and time series techniques to study how the issue attention of state lawmakers evolves vis-& agrave;-vis various local- and national-level actors. Our findings suggest that state legislators operate at the confluence of national and local influences. In line with arguments highlighting the nationalization of state politics, we find that state legislators are consistently responsive to policy debates among members of Congress. However, despite growing nationalization concerns, we also find strong evidence of issue responsiveness by legislators to members of the public in their states and moderate responsiveness to regional media sources.
In this report we provide an overview of the kinds of data academics need in order to conduct independent research into political online safety matters on social media platforms, and the challenges they currently face. Additionally, we put forward ideas regarding novel governance structures that would enable high-quality independent research, while protecting users’ rights and data privacy, in the United Kingdom.
Using the Spanish case, this paper explores whether the European and regional content of legislation debated in national parliaments influences parties' legislative behaviour, and more specifically their decision to propose legislative amendments. Based on regression analysis and multilevel modelling for hypothesis testing and an original dataset including information on more than 90,000 amendments, results illustrate that parties are more likely to propose amendments on bills having regional content. This is explained because self-government between the state and regional authorities remains open to renegotiation and due to the persistent territorial cleavage. Yet, against the niche party literature, results show that both regional and statewide parties propose a significant number of amendments on bills having regional content, especially when territorial affairs are politicised. Regarding the European dimension, despite the absence of hard Eurosceptic parties in Spain and the lack of a strong political cleavage over European integration, our findings challenge the idea that in Europhile countries EU-related bills are rarely amended. Overall, while previous literature has primarily focused on governmental factors to explain parties' decisions to propose amendments on national legislation, this paper highlights the importance of considering multi-level dynamics.
Language models like BERT or GPT are becoming increasingly popular measurement tools, but are the measurements they produce valid? Literature suggests that there is still a relevant gap between the ambitions of computational text analysis methods and the validity of their outputs. One prominent threat to validity is hidden biases in the training data, where models learn group-specific language patterns instead of the concept researchers want to measure. This paper investigates to what extent these biases impact the validity of measurements created with language models. We conduct a comparative analysis across nine group types in four datasets with three types of classification models, focusing on the robustness of models against biases and on the validity of their outputs. While we find that all types of models learn biases, the effects on validity are surprisingly small. In particular when models receive instructions as an additional input, they become more robust against biases from the fine-tuning data and produce more valid measurements across different groups. An instruction-based model (BERT-NLI) sees its average test-set performance decrease by only 0.4% F1 macro when trained on biased data and its error probability on groups it has not seen during training increases only by 0.8%.
Lawmaking is not about enacting bills. It is about enacting policies. We portray bills as necessary vehicles for the advancement of policy ideas. To survive, a policy must be incorporated into a successful bill. But bills and policies are not the same thing. Central to our argument is the legislative hitchhiker – a policy proposal first proposed in one bill that is subsequently incorporated into another bill. Drawing on text as data methods, we demonstrate the value and feasibility of viewing lawmaking from a policy progress rather than bill progress perspective.
Supervised machine learning is an increasingly popular tool for analyzing large political text corpora. The main disadvantage of supervised machine learning is the need for thousands of manually annotated training data points. This issue is particularly important in the social sciences where most new research questions require new training data for a new task tailored to the specific research question. This paper analyses how deep transfer learning can help address this challenge by accumulating “prior knowledge” in language models. Models like BERT can learn statistical language patterns through pre-training (“language knowledge”), and reliance on task-specific data can be reduced by training on universal tasks like natural language inference (NLI; “task knowledge”). We demonstrate the benefits of transfer learning on a wide range of eight tasks. Across these eight tasks, our BERT-NLI model fine-tuned on 100 to 2,500 texts performs on average 10.7 to 18.3 percentage points better than classical models without transfer learning. Our study indicates that BERT-NLI fine-tuned on 500 texts achieves similar performance as classical models trained on around 5,000 texts. Moreover, we show that transfer learning works particularly well on imbalanced data. We conclude by discussing limitations of transfer learning and by outlining new opportunities for political science research.
Social media companies increasingly play a role in regulating freedom of speech. Debates over ideological motivations behind suspension policies of major platforms are on the rise. This study contributes to this ongoing debate by looking at content moderation from a geopolitical perspective. The starting premise is that US-based social media companies may be inclined to moderate content on their platforms in compliance with US sanctions laws, especially those concerned with the Specially Designated Nationals and Blocked Persons List. Despite the release of transparency reports by social media companies, we know little about the scope of the problem and the impact of suspensions on political conversations. I tracked 600,000 users who follow Iranian elites on Twitter. After accounting for alternative explanations, the results show that Principlist (conservative) users and those supportive of the Iranian government are significantly more likely to be suspended. Further analyses uncover the types of discussions that are being suppressed as a result of these suspensions. Although the exact mechanism at hand cannot be decisively isolated, this paper contributes to building a better understanding of how governments can influence conversations of geopolitical relevance, and how social media suspensions shape political conversations online.
Americans view their in-party members positively and out-party members negatively. It remains unclear, however, whether in-party affinity (i.e., positive partisanship) or out-party animosity (i.e., negative partisanship) more strongly influences political attitudes and behaviors. Unlike past work, which relies on survey self-reports or experimental designs among ordinary citizens, this pre-registered project examines actual social media expressions of an exhaustive list of American politicians as well as citizens’ engagement with these posts. Relying on 1,195,844 tweets sent by 564 political elites (i.e., members of US House and Senate, Presidential and Vice-Presidential nominees from 2000 to 2020, and members of the Trump Cabinet) and machine learning to reliably classify the tone of the tweets, we show that elite expressions online are driven by positive partisanship more than negative partisanship. Although politicians post many tweets negative toward the out-party, they post more tweets positive toward their in-party. However, more ideologically extreme politicians and those in the opposition (i.e., the Democrats) are more negative toward the out-party than those ideologically moderate and whose party is in power. Furthermore, examining how Twitter users react to these posts, we find that negative partisanship plays a greater role in online engagement: users are more likely to like and share politicians’ tweets negative toward the out-party than tweets positive toward the in-party. This project has important theoretical and democratic implications, and extends the use of trace data and computational methods in political behavior.
Based on a new comprehensive dataset containing information on 93,722 amendments, this article explores the circumstances under which Spanish legislators propose amendments to executive bills. Our results show that legislators respond to variations in both governmental factors and bargaining dynamics. In single-party minority governments, ad hoc legislature agreements translate into more amendments. However, legislators do not introduce significantly fewer amendments under absolute majority governments, when the chances of their proposals being accepted fall. After controlling for many confounders, the results show that amending activity reacts to attention allocation dynamics - mediatised bills receive more amendments - but not to variations in contextual factors - the number of amendments does not significantly increase when the economic situation deteriorates. Finally, bills associated with a greater number of committee appearances from interest groups, experts and public officials are more often the target of amendments, signalling that an informational logic is also at play.
Generative Large Language Models (LLMs) have become the mainstream choice for fewshot and zeroshot learning thanks to the universality of text generation. Many users, however, do not need the broad capabilities of generative LLMs when they only want to automate a classification task. Smaller BERT-like models can also learn universal tasks, which allow them to do any text classification task without requiring fine-tuning (zeroshot classification) or to learn new tasks with only a few examples (fewshot), while being significantly more efficient than generative LLMs. This paper (1) explains how Natural Language Inference (NLI) can be used as a universal classification task that follows similar principles as instruction fine-tuning of generative LLMs, (2) provides a step-by-step guide with reusable Jupyter notebooks for building a universal classifier, and (3) shares the resulting universal classifier that is trained on 33 datasets with 389 diverse classes. Parts of the code we share has been used to train our older zeroshot classifiers that have been downloaded more than 55 million times via the Hugging Face Hub as of December 2023. Our new classifier improves zeroshot performance by 9.4%.
Human annotation of data, including text and image materials, is a bedrock of political science research. Yet we often overlook how the identities of our annotators may systematically affect their labels. We call the sensitivity of labels to annotator identity "labeler-characteristic bias" (LCB). We demonstrate the persistence and risks of LCB for downstream analyses in two examples, first with image data from the United States and second with text data from the Netherlands. In both examples we observe significant differences in annotations based on annotator gender and political identity. After laying out a general typology of annotator biases and their relationship to inter-rater reliability, we provide suggestions and solutions for how to handle LCB. The first step to addressing LCB is to recruit a diverse labeler corps and test for LCB. Where LCB is found, solutions are modeling subgroup effects or generating composite labels based on target population demographics.
This study investigates the potential role both untrustworthy and partisan websites play in misinforming audiences by testing whether actual exposure to these sites is associated with political misperceptions. Using a sample of American adult social media users, we match data from individuals' Internet browser histories with a survey measuring the accuracy of political beliefs. We find that visits to partisan websites are at times related to misperceptions consistent with the political bias of the site. However, we do not find strong evidence that untrustworthy websites consistently relate to false beliefs. There is also little evidence that visits to less partisan, centrist news sites are associated with more accurate political beliefs about these issues, suggesting that exposure to politically neutral news is not necessarily the antidote to misinformation. Results suggest that focusing on partisan news sites-rather than untrustworthy sites-may be fruitful to understanding how media contribute to political misperceptions.
The social science toolkit for computational text analysis is still very much in the making. We know surprisingly little about how to produce valid insights from large amounts of multilingual texts for comparative social science research. In this paper, we test several recent innovations from deep transfer learning to help advance the computational toolkit for social science research in multilingual settings. We investigate the extent to which ‘prior language and task knowledge’ stored in the parameters of modern language models is useful for enabling multilingual research; we investigate the extent to which these algorithms can be fruitfully combined with machine translation; and we investigate whether these methods are not only accurate but also practical and valid in multilingual settings – three essential conditions for lowering the language barrier in practice. We use two datasets with texts in 12 languages from 27 countries for our investigation. Our analysis shows, that, based on these innovations, supervised machine learning can produce substantively meaningful outputs. Our BERT-NLI model trained on only 674 or 1674 texts in only one or two languages can validly predict political party families’ stances towards immigration in eight other languages and ten other countries.
We offer comprehensive evidence of preferences for ideological congruity when people engage with politicians, pundits, and news organizations on social media. Using 4 years of data (2016-2019) from a random sample of 1.5 million Twitter users, we examine three behaviors studied separately to date: (i) following of in-group versus out-group elites, (ii) sharing in-group versus out-group information (retweeting), and (iii) commenting on the shared information (quote tweeting). We find that the majority of users (60%) do not follow any political elites. Those who do follow in-group elite accounts at much higher rates than out-group accounts (90 versus 10%), share information from in-group elites 13 times more frequently than from out-group elites, and often add negative comments to the shared out-group information. Conservatives are twice as likely as liberals to share in-group versus out-group content. These patterns are robust, emerge across issues and political elites, and exist regardless of users' ideological extremity.
Democratic theorists and the public emphasize the centrality of news media to a well-functioning society. Yet, there are reasons to believe that news exposure can have a range of largely overlooked detrimental effects. This preregistered project examines news exposure effects on desirable outcomes, i.e., political knowledge, participation, and support for compromise, and detrimental outcomes, i.e., attitude and affective polarization, negative system perceptions, and worsened individual well-being. We rely on two complementary over-time experiments that combine participants' survey self-reports and their behavioral browsing data: one that incentivized participants to take a 'news vacation' for a week (N = 803; 6M visits) in the US, the other to 'news binge' for 2 weeks (N = 939; 4M visits) in Poland. Across both experiments, we demonstrate that reducing or increasing news exposure has no impact on the positive or negative outcomes tested. These null effects emerge irrespective of participants' prior levels of news consumption and whether prior news diet was like-minded, and regardless of compliance levels. We argue that these findings reflect the reality of limited news exposure in the real world, with news exposure comprising on average roughly 3% of citizens' online information diet.
Amsterdam University Press is a leading publisher of academic books, journals and textbooks in the Humanities and Social Sciences. Our aim is to make current research available to scholars, students, innovators, and the general public. AUP stands for scholarly excellence, global presence, and engagement with the international academic community.
State governments are the focus of important policy decisions in the United States. How do state legislators use their public communications ---particularly social media---to engage with policy debates? Due to previous data limitations, we lack systematic information about whether and how state legislators publicly discuss policy and how this behavior varies across contexts. Using Twitter data and state of the art topic modeling techniques, we introduce a method to study state legislator policy priorities and apply the method to fifteen U.S. states in 2018. We show that we are able to capture the policy issues discussed by state legislators with substantially more accuracy than existing methods. We then present initial findings that validate the method and speak to debates in the literature. For example, state legislators in competitive districts are more likely to discuss policy than those in less competitive districts, and legislators from more professional legislatures discuss policy at similar rates to those in less professional legislatures. We conclude by discussing promising avenues for future state politics research using this new approach.
A narrow information diet may be partly to blame for the growing political divides in the United States, suggesting exposure to dissimilar views as a remedy. These efforts, however, could be counterproductive, exacerbating attitude and affective polarization. Yet findings on whether such boomerang effect exists are mixed and the consequences of dissimilar exposure on other important outcomes remain unexplored. To contribute to this debate, we rely on a preregistered longitudinal experimental design combining participants' survey self-reports and their behavioral browsing data, in which one should observe boomerang effects. We incentivized liberals to read political articles on extreme conservative outlets (Breitbart, The American Spectator, and The Blaze) and conservatives to read extreme left-leaning sites (Mother Jones, Democracy Now, and The Nation). We maximize ecological validity by embedding the treatment in a larger project that tracks over time changes in online exposure and attitudes. We explored the effects on attitude and affective polarization, as well as on perceptions of the political system, support for democratic principles, and personal well-being. Overall we find little evidence of boomerang effects.