Manual annotation of the policy content of political texts forms the basis for one of the most widely used empirical measures in comparative politics: left-right policy positions. Bridging automated “text as data” approaches and qualitative content analysis, we apply statistical scaling to this data to learn more about the association of specific policy dimensions to the left-right super-dimension, in a way that minimizes ex ante assumptions about the substantive content of left-right policy. We apply a Bayesian negative binomial variant of Slapin and Proksch’s (2008) “wordfish” model to category counts from party manifestos coded by the Manifesto Project, providing a data-driven approach that offers new insights into the policy content of left and right. We demonstrate how this method also works with content not originally designed for measuring positions. In addition, we show how the approach can be extended to measure the policy content of two latent dimensions, with some categories contributing to both.
and the Career Mobility of Bureaucrats. Career mobility is conceptualized in terms of the amount and bias of organizational, occupational, vertical and geographical movement. An operational measure is offered for the amount of mobility (not the bias) which permits comparative generalizations to be made about officials based on simple sampling procedures. For illustrative purposes, the amount of organizational mobility is measured for two samples of federal officials: David T. Stanley's sample (N = 557) of higher civil servants and a sample (N = 300) of foreign affairs officials. Several dichotomous background variables are included for both samples of officials including: age, rank, and education. Varying background characteristics is found to make only a small mobility difference for higher civil servants and a much greater difference for foreign affairs officials. Several speculations are offered to explain the political significance of the empirical findings, and a comparative typology of career mobility is developed. of the effects of the new rules on the selection of delegates to the Democratic National Convention on the strategic environment faced by presidential candidates. The article outlines the basic aspects of those rule changes, describes the problems involved in the implementation of the rules, and attempts to evaluate some of the consequences faced by the presidential campaign of Senator George McGovern in the context of the California primary of 1972. 45 Social Mobilization and the Russification of Soviet Nationalities. This paper examines certain major demographic bases of ethnic identity change among the mass populations of the non-Russian nationalities of the USSR. Ethnic identity is defined here as an individual's affective attachment to certain core symbols of his nationality group: the group name and its historic language. The hypoth-eses tested concern the impact of social mobilization, contact with Russians, and traditional religion on ethnic identity change. The levels of Russification of 46 indigenous nationalities whose official national homelands have Autonomous Oblast status or higher are examined on the basis of 1959 Soviet census materials. By use of regression analysis it is shown that: (a) social mobilization is strongly conducive to the Russification of non-Russian nationalities residing in their official areas; (b) exposure to Russians is conducive to the Russification of both mobilized and unmobilized local populations, but the Russification effect of exposure to Russians is much smaller for the unmobilized than for the mobilized populations; (c) even where exposure to Russians is extensive and enduring, both socially mobilized and unmobilized Muslim ethnic groups are much less likely to be Russified than non-Muslims; it is proposed that a Muslim ethnic ideology mediates between the dynamic demographic influences on Russification and the actual manifestation of Russification. 67 The Divisive Primary Revisited: Party Activists in Iowa. This study was conducted to test the frequently made assertion that primary elections are divisive among party activists who participate time-series cohort provides considerable support for an historically based generational explanation for age-group differences and per- mits examination of one process through which partisan realignments may occur. show number of first-term increase in membership stability. computation of such data is first step in and consequences ; of the rightwing theories of the and Soviet foreign policy, and left wing explanations of the United and foreign policy. The conclusion suggests that both theories are fundamentally flawed in two (1) As employed by their proponents, the theories appear incapable of falsified; and (2) studies them serious methodological flaws that the canons of systematic inquiry. A on Data Handbooks in Political Science. Four major new com- pilations of macropolitical data are compared and evaluated. Each summarizes a large-scale research effort to code or to collect data suitable for theoretically relevant, cross-national comparisons. As a group the new handbooks incorporate many improvements and innovations on earlier handbooks, which con-centrated mainly on cross-sectional, aggregate data or simplistically coded judgments about nation-states. About a third of their measures consist of "made" data, derived by coding journalistic and historical sources. All provide some measures for cross-time comparisons; one is devoted exclusively to time-series data. Many of their measures denote properties of internal and international conflict and of international transactions. All but one are painfully self-conscious about problems of reliability and comparability of data. One criticism is the reliance of several of the handbooks on "counts" of conflict events rather than assessment of more theoretically relevant properties of conflict. A second is the paucity of indicators of inequality and, more generally, of measures which give a "view from the bottom" of political systems.
An abstract is not available for this content. As you have access to this content, full HTML content is provided on this page. A PDF of this content is also available in through the ‘Save PDF’ action button.
: We analyze the engagement of citizens, media and politicians on social media during the campaign period of Brexit referendum in the UK in June 2016. We focus on the social media conversation on Twitter, the largest micro-blogging service with more than 300 million active users and approximately 500 million new messages generated per day. We analyse the networks of Tweets focusing on both the topics users have discussed during the referendum campaign and the network of followerships. Using machine learning, we classify users as favouring Leave or Remain, and use this information to compare the esti-mates of the topics they discuss in the debate. The topic analysis reveals that the structure of topics by Leave users are more densely related than those by Remain users. Furthermore, by analyzing the network analysis of followership, we reveal that the important accounts in the Brexit debate such as media and politicians are categorically different between Leave and Remain, indicating a fundamentally divided messaging communication structure, reinforcing work done elsewhere on how social networks form ideological “echo chambers” reinforcing one’s pre-existing attitudes.
Computational text analysis has become an exciting research field with many applications in communication research. It can be a difficult method to apply, however, because it requires knowledge of various techniques, and the software required to perform most of these techniques is not readily available in common statistical software packages. In this teacher's corner, we address these barriers by providing an overview of general steps and operations in a computational text analysis project, and demonstrate how each step can be performed using the R statistical software. As a popular open-source platform, R has an extensive user community that develops and maintains a wide range of text analysis packages. We show that these packages make it easy to perform advanced text analytics.
Political scientists lack domain-specific measures for the purpose of measuring the sophistication of political communication. We systematically review the shortcomings of existing approaches, before developing a new and better method along with software tools to apply it. We use crowdsourcing to perform thousands of pairwise comparisons of text snippets and incorporate these results into a statistical model of sophistication. This includes previously excluded features such as parts of speech and a measure of word rarity derived from dynamic term frequencies in the Google Books data set. Our technique not only shows which features are appropriate to the political domain and how, but also provides a measure easily applied and rescaled to political texts in a way that facilitates probabilistic comparisons. We reanalyze the State of the Union corpus to demonstrate how conclusions differ when using our improved approach, including the ability to compare complexity as a function of covariates.
Computational text analysis has become an exciting research field with many applications in communication research. It can be a difficult method to apply, however, because it requires knowledge of various techniques, and the software required to perform most of these techniques is not readily available in common statistical software packages. In this teacher’s corner, we address these barriers by providing an overview of general steps and operations in a computational text analysis project, and demonstrate how each step can be performed using the R statistical software. As a popular open-source platform, R has an extensive user community that develops and maintains a wide range of text analysis packages. We show that these packages make it easy to perform advanced text analytics. With the increasing importance of computational text analysis in communication research (Boumans & Trilling, 2016; Grimmer & Stewart, 2013), many researchers face the challenge of learning how to use advanced software that enables this type of analysis. Currently, one of the most popular environments for computational methods and the emerging field of “data science” 1 is the R statistical software (R Core Team, 2017). However, for researchers that are not well-versed in programming, learning how to use R can be a challenge, and performing text analysis in particular can seem daunting. In this teacher’s corner, we show that performing text analysis in R is not as hard as some might fear. We provide a step-bystep introduction into the use of common techniques, with the aim of helping researchers get acquainted with computational text analysis in general, as well as getting a start at performing advanced text analysis studies in R. R is a free, open-source, cross-platform programming environment. In contrast to most programming languages, R was specifically designed for statistical analysis, which makes it highly suitable for data science applications. Although the learning curve for programming with R can be steep, especially for people without prior programming experience, the tools now available for carrying out text analysis in R make it easy to perform powerful, cutting-edge text analytics using only a few simple commands. One of the keys to TEXT ANALYSIS IN R 2 R’s explosive growth (Fox & Leanage, 2016; TIOBE, 2017) has been its densely populated collection of extension software libraries, known in R terminology as packages, supplied and maintained by R’s extensive user community. Each package extends the functionality of the base R language and core packages, and in addition to functions and data must include documentation and examples, often in the form of vignettes demonstrating the use of the package. The best-known package repository, the Comprehensive R Archive Network (CRAN), currently has over 10,000 packages that are published, and which have gone through an extensive screening for procedural conformity and cross-platform compatibility before being accepted by the archive.2 R thus features a wide range of inter-compatible packages, maintained and continuously updated by scholars, practitioners, and projects such as RStudio and rOpenSci. Furthermore, these packages may be installed easily and safely from within the R environment using a single command. R thus provides a solid bridge for developers and users of new analysis tools to meet, making it a very suitable programming environment for scientific collaboration. Text analysis in particular has become well established in R. There is a vast collection of dedicated text processing and text analysis packages, from low-level string operations (Gagolewski, 2017) to advanced text modeling techniques such as fitting Latent Dirichlet Allocation models (Blei, Ng, & Jordan, 2003; Roberts et al., 2014)—nearly 50 packages in total at our last count. Furthermore, there is an increasing effort among developers to cooperate and coordinate, such as the rOpenSci special interest group.3 One of the main advantages of performing text analysis in R is that it is often possible, and relatively easy, to switch between different packages or to combine them. Recent efforts among the R text analysis developers’ community are designed to promote this interoperability to maximize flexibility and choice among users.4 As a result, learning the basics for text analysis in R provides access to a wide range of advanced text analysis features. Structure of this Teacher’s Corner This teacher’s corner covers the most common steps for performing text analysis in R, from data preparation to analysis, and provides easy to replicate example code to perform each step. The example code is also digitally available in our online appendix, which is updated over time.5 We focus primarily on bag-of-words text analysis approaches, meaning that only the frequencies of words per text are used and word positions are ignored. Although this drastically simplifies text content, research and many real-world applications show that word frequencies alone contain sufficient information for many types of analysis (Grimmer & Stewart, 2013). Table 1 presents an overview of the text analysis operations that we address, categorized in three sections. In the data preparation section we discuss five steps to prepare texts for analysis. The first step, importing text, covers the functions for reading texts from various types of file formats (e.g., txt, csv, pdf) into a raw text corpus in R. The steps string operations and preprocessing cover techniques for manipulating raw texts and processing them into tokens (i.e., units of text, such as words or word stems). The tokens are then used for creating the document-term matrix (DTM), which is a common format for representing a bag-of-words type corpus, that is used by many R text analysis packages. Other nonbag-of-words formats, such as the tokenlist, are briefly touched upon in the advanced topics section. Finally, it is a common step to filter and weight the terms in the DTM. These TEXT ANALYSIS IN R 3 Table 1 An overview of text analysis operations, with the R packages used in this teacher’s corner Operation R packages example alternatives Data preparation importing text readtext jsonlite, XML, antiword, readxl, pdftools string operations stringi stringr preprocessing quanteda stringi, tokenizers, snowballC, tm, etc. document-term matrix (DTM) quanteda tm, tidytext, Matrix filtering and weighting quanteda tm, tidytext, Matrix Analysis dictionary quanteda tm, tidytext, koRpus, corpustools supervised machine learning quanteda RTextTools, kerasR, austin unsupervised machine learning topicmodels quanteda, stm, austin, text2vec text statistics quanteda koRpus, corpustools, textreuse Advanced topics advanced NLP spacyr coreNLP, cleanNLP, koRpus word positions and syntax corpustools quanteda, tidytext, koRpus Figure 1 . Order of text analysis operations for data preparation and analysis. dtm files web pages etc. R text corpus tokens tokenlist
Borrowing from automated “text as data” approaches, we show how statistical scaling models can be applied to hand-coded content analysis to improve estimates of political parties’ leftright policy positions. We apply a Bayesian item-response theory (IRT) model to category counts from coded party manifestos, treating the categories as “items” and policy positions as a latent variable. This approach also produces direct estimates of how each policy category relates to left-right ideology, without having to decide these relationships in advance based on out of sample fitting, political theory, assertion, or guesswork. This approach not only prevents the misspecification endemic to a fixed-index approach, but also works well even with items that are not specifically designed to measure ideological positioning.
Social media play an increasingly important part in the communication strategies of political campaigns by reflecting information about the policy preferences and opinions of political actors and their public followers. In addition, the content of the messages provides rich information about the political issues and the framing of those issues during elections, such as whether contested issues concern Europe or rather extend pre-existing national debates. In this study, we survey the European landscape of social media using tweets originating from and referring to political actors during the 2014 European Parliament election campaign. We describe the language and national distribution of the messages, the relative volume of different types of communications, and the factors that determine the adoption and use of social media by the candidates. We also analyze the dynamics of the volume and content of the communications over the duration of the campaign with reference to both the EU integration dimension of the debate and the prominence of the most visible list-leading candidates. Our findings indicate that the lead candidates and their televised debate had a prominent influence on the volume and content of communications, and that the content and emotional tone of communications more reflects preferences along the EU dimension of political contestation rather than classic national issues relating to left-right differences.
An abstract is not available for this content. As you have access to this content, full HTML content is provided on this page. A PDF of this content is also available in through the ‘Save PDF’ action button.
Abstract Hand-coded party manifestos have formed the largest source of comparative, over-time data for estimating party policy positions and emphases, based on the fundamental assumption that left-right ideological positions can be measured by comparing the relative emphasis of predefined policy categories. We critically challenge this approach by showing that left-right ideology can be better measured from specific policy emphasis using an inductive approach, and by demonstrating that there is no single a priori definition of left-right policy that outperforms the inductive approach across contexts. To estimate party positions, we apply a Bayesian measurement model to category counts from coded party manifestos, treating treating the categories as “items” and policy positions as a latent variable. This approach also produces direct estimates of how each policy category relates to left-right ideology, without having to decide these relationships in advance based on political theory, exploratory analysis, or guesswork. We also demonstrate that the IRT approach can work even when the items are not specifically designed to measure ideological positions. A big advantage of our framework lies in its flexibility: here, we specifically show how two infer policy positions in two dimensions , but there are numerous extensions for future research, such as examining coder effects or adding covariates to predict the model parameters.
An abstract is not available for this content. As you have access to this content, full HTML content is provided on this page. A PDF of this content is also available in through the ‘Save PDF’ action button.