The Russian state-funded international broadcaster RT (formerly Russia Today) has attracted much attention as a purveyor of Russian propaganda. To date, most studies of RT have focused on its broadcast, website, and social media content, with little research on its audiences. Through a data-driven application of network science and other computational methods, we address this gap to provide insight into the demographics and interests of RT's Twitter followers, as well as how they engage with RT. Building upon recent studies of Russian state-sponsored media, we report three main results. First, we find that most of RT's Twitter followers only very rarely engage with its content and tend to be exposed to RT's content alongside other mainstream news channels. This indicates that RT is not a central part of their online news media environment. Second, using probabilistic computational methods, we show that followers of RT are slightly more likely to be older and male than average Twitter users, and they are far more likely to be bots. Third, we identify thirty-five distinct audience segments, which vary in terms of their nationality, languages, and interests. This audience segmentation reveals the considerable heterogeneity of RT's Twitter followers. Accordingly, we conclude that generalizations about RT's audience based on analyses of RT's media content, or on vocal minorities among its wider audiences, are unhelpful and limit our understanding of RT and its appeal to international audiences.
In the context of deteriorating relations with ‘Western’ states, Russia’s state-funded international broadcasters are often understood as malign propaganda rather than as agents of soft power. Subsequently, there is a major credibility gap between how Russian state media represents itself to the world and how it is actually perceived by overseas publics. However, based on the study of RT’s coverage of the Russian hosted FIFA 2018 World Cup and the audience reactions this prompted, we find that this credibility gap was partially bridged. By analysing over 700 articles published by RT, alongside social media and focus group research, we find that RT’s World Cup coverage created an unusually positive vision of Russia that appealed to international audiences. Our study demonstrates how state-funded international broadcaster coverage of sports mega-events can generate a soft power effect with audiences, even when the host state – such as Russia – has a poor international reputation.
Since the early years of the past century, many scholars have focused their efforts towards designing models to better understand the way listeners perceive musical tension. From the existing models, Lerdahl’s has shown strong correlations against tension judgements provided by human listeners and has been used to make accurate predictions of musical tension. However, a full automation of Lerdahl’s model of tension has not yet been made available. This paper presents a computational approach to automatically calculate musical tension according to Lerdahl’s model, with a publicly available implementation.
In the world of public misinformation, there are many cases where the information is not false or fabricated, but rather has been manipulated using more subtle techniques such as word replacements, selection of details, omissions and argument distortion. These techniques can have the effect of influencing the reader’s frame of mind towards the events reported. We currently lack the necessary tools to uncover such manipulations automatically. In this position paper, we propose an integrated analysis framework and pipeline to identify various narrative signals in news articles; such as structural roles, framing, and subjectivity. By comparing these at the document level and sentence level, it will be possible to highlight differences of narrative techniques used to report the same news events.
A basic step in any annotation effort is the measurement of the Inter Annotator Agreement (IAA). An important factor that can affect the IAA is the presence of annotator bias. In this paper we introduce a new interpretation and application of the Item Response Theory (IRT) to detect annotators’ bias. Our interpretation of IRT offers an original bias identification method that can be used to compare annotators’ bias and characterise annotation disagreement. Our method can be used to spot outlier annotators, improve annotation guidelines and provide a better picture of the annotation reliability. Additionally, because scales for IAA interpretation are not generally agreed upon, our bias identification method is valuable as a complement to the IAA value which can help with understanding the annotation disagreement.
Written communication skills are considered to be highly desirable in computing graduates. However, many computing students do not have a background in which these skills have been developed, and the skills are often not well addressed within a computing curriculum. For some multidisciplinary areas, such as data science, the range of potential stakeholders makes the need for communications skills all the greater. As interest in data science increases and the technical skills of the area are in ever higher demand, understanding effective teaching and learning of these interdisciplinary aspects is receiving significant attention by academics, industry and government in an effort to address the digital skills gap. In this paper, we report on the experience of adapting a final year data science module in an undergraduate computing curriculum to help develop the skills needed for writing extended reports. From its inception, the module has used Jupyter notebooks to develop the students' skills in the coding aspects of the module. However, over several presentations, we have investigated how the cell-based structure of the notebooks can be exploited to improve the students' understanding of how to structure a report on a data investigation. We have increasingly designed the assessment for the module to take advantage of the learning affordances of Jupyter notebooks to support both raw data analysis and effective report writing. We reflect on the lessons learned from these changes to the assessment model, and the students' responses to the changes.
Throughout 2017, the Russian state broadcaster, RT (formerly Russia Today), commemorated the centenary of the 1917 revolution with a social media re-enactment. Centred on Twitter, the 1917LIVE project involved over 90 revolution-era characters tweeting in real time as if the 1917 revolution was happening live on social media. This article is based on an analysis of a sample of tweets by users who engaged with 1917LIVE, alongside focus group discussions with its followers. We argue that a cultural studies perspective can shed important light on the political significance of RT’s social media re-enactment in ways that current studies of public diplomacy as a soft power resource often fail to do. It can advance soft power theory by offering a more nuanced, dynamic analysis of how state media mobilise, and how audiences engage with, social media re-enactments as commemorative events. We find that rather than promoting a unitary propagandistic narrative about Russia, 1917LIVE served instead to soften attitudes towards RT itself – encouraging audiences to view RT as an educator and entertainer as well as a news broadcaster – normalising its presence as a Russian public diplomacy resource in the international news media landscape. Our analysis of audience interactions with and interpretations of 1917LIVE affords insights into how the 1917 re-enactment worked as didactic entertainment eliciting affective identification with the characters of the revolution. Such public diplomacy projects contribute in the short term to a strengthening of the engagement required to create longer-term soft power effects.
Rating and Likert scales are widely used in evaluation experiments to measure the quality of Natural Language Generation (NLG) systems. We review the use of rating and Likert scales for NLG evaluation tasks published in NLG specialized conferences over the last ten years (135 papers in total). Our analysis brings to light a number of deviations from good practice in their use. We conclude with some recommendations about the use of such scales. Our aim is to encourage the appropriate use of evaluation methodologies in the NLG community.
Inter-Annotator Agreement (IAA) is used as a means of assessing the quality of NLG evaluation data, in particular, its reliability. According to existing scales of IAA interpretation – see, for example, Lommel et al. (2014), Liu et al. (2016), Sedoc et al. (2018) and Amidei et al. (2018a) – most data collected for NLG evaluation fail the reliability test. We confirmed this trend by analysing papers published over the last 10 years in NLG-specific conferences (in total 135 papers that included some sort of human evaluation study). Following Sampson and Babarczy (2008), Lommel et al. (2014), Joshi et al. (2016) and Amidei et al. (2018b), such phenomena can be explained in terms of irreducible human language variability. Using three case studies, we show the limits of considering IAA as the only criterion for checking evaluation reliability. Given human language variability, we propose that for human evaluation of NLG, correlation coefficients and agreement coefficients should be used together to obtain a better assessment of the evaluation data reliability. This is illustrated using the three case studies.
In the last few years Automatic Question Generation (AQG) has attracted increasing interest. In this paper we survey the evaluation methodologies used in AQG. Based on a sample of 37 papers, our research shows that the systems’ development has not been accompanied by similar developments in the methodologies used for the systems’ evaluation. Indeed, in the papers we examine here, we find a wide variety of both intrinsic and extrinsic evaluation methodologies. Such diverse evaluation practices make it difficult to reliably compare the quality of different generation systems. Our study suggests that, given the rapidly increasing level of research in the area, a common framework is urgently needed to compare the performance of AQG systems and NLG systems more generally.
Tacit knowledge in requirements documents can lead to miscommunication between software engineers and other stakeholders. One way in which the presence of tacit knowledge is signalled in text is by linguistic presuppositions. In this paper, we present a brief introduction to tacit knowledge, presuppositions and the links between them. Our aim is to build a theoretically grounded system which is able to automatically highlight all the presuppositions that might have a negative impact on communication through requirements documents.
Human evaluations are broadly thought to be more valuable the higher the inter-annotator agreement. In this paper we examine this idea. We will describe our experiments and analysis within the area of Automatic Question Generation. Our experiments show how annotators diverge in language annotation tasks due to a range of ineliminable factors. For this reason, we believe that annotation schemes for natural language generation tasks that are aimed at evaluating language quality need to be treated with great care. In particular, an unchecked focus on reduction of disagreement among annotators runs the danger of creating generation goals that reward output that is more distant from, rather than closer to, natural human-like language. We conclude the paper by suggesting a new approach to the use of the agreement metrics in natural language generation evaluation tasks.
This paper pays attention to the immense and febrile field of digital image files which picture the smart city as they circulate on the social media platform Twitter. The paper considers tweeted images as an affective field in which flow and colour are especially generative. This luminescent field is territorialised into different, emergent forms of becoming ‘smart’. The paper identifies these territorialisations in two ways: firstly, by using the data visualisation software ImagePlot to create a visualisation of 9030 tweeted images related to smart cities; and secondly, by responding to the affective pushes of the image files thus visualised. It identifies two colours and three ways of affectively becoming smart: participating in smart, learning about smart, and anticipating smart, which are enacted with different distributions of mostly orange and blue images. The paper thus argues that debates about the power relations embedded in the smart city should consider the particular affective enactment of being smart that happens via social media. More generally, the paper concludes that geographers must pay more attention to the diverse and productive vitalities of social media platforms in urban life and that this will require experiment with methods that are responsive to specific digital qualities.
Recent research has shown that the performance of search personalization depends on the richness of user profiles which normally represent the user’s topical interests. In this paper, we propose a new embedding approach to learning user profiles, where users are embedded on a topical interest space. We then directly utilize the user profiles for search personalization. Experiments on query logs from a major commercial web search engine demonstrate that our embedding approach improves the performance of the search engine and also achieves better search performance than other strong baselines.
Recent research has shown the usefulness of using collective user interaction data (e.g., query logs) to recommend query modification suggestions for Intranet search. However, most of the query suggestion approaches for Intranet search follow an ``one size fits all'' strategy, whereby different users who submit an identical query would get the same query suggestion list. This is problematic, as even with the same query, different users may have different topics of interest, which may change over time in response to the user's interaction with the system. We address the problem by proposing a personalised query suggestion framework for Intranet search. For each search session, we construct two temporal user profiles: a click user profile using the user's clicked documents and a query user profile using the user's submitted queries. We then use the two profiles to re-rank the non-personalised query suggestion list returned by a state-of-the-art query suggestion method for Intranet search. Experimental results on a large-scale query logs collection show that our personalised framework significantly improves the quality of suggested queries.
We study the problem of detecting sentences describing adverse drug reactions (ADRs) and frame the problem as binary classification. We investigate different neural network (NN) architectures for ADR classification. In particular, we propose two new neural network models, Convolutional Recurrent Neural Network (CRNN) by concatenating convolutional neural networks with recurrent neural networks, and Convolutional Neural Network with Attention (CNNA) by adding attention weights into convolutional neural networks. We evaluate various NN architectures on a Twitter dataset containing informal language and an Adverse Drug Effects (ADE) dataset constructed by sampling from MEDLINE case reports. Experimental results show that all the NN architectures outperform the traditional maximum entropy classifiers trained from n-grams with different weighting strategies considerably on both datasets. On the Twitter dataset, all the NN architectures perform similarly. But on the ADE dataset, CNN performs better than other more complex CNN variants. Nevertheless, CNNA allows the visualisation of attention weights of words when making classification decisions and hence is more appropriate for the extraction of word subsequences describing ADRs.
Abstract Stylistic composition is a creative musical activity, in which students as well as renowned composers write according to the style of another composer or period. We describe and evaluate two computational models of stylistic composition, called Racchman-Oct2010 (random constrained chain of Markovian nodes, October 2010) and Racchmaninof-Oct2010 (Racchman with inheritance of form). The former is a constrained Markov model, and the latter embeds this model in an analogy-based design system. Racchmaninof-Oct2010 applies a pattern discovery algorithm called SIACT and a perceptually validated formula for rating pattern importance, to guide the generation of a new target design from an existing source design. A listening study is reported concerning human judgments of music excerpts that are, to varying degrees, in the style of mazurkas by Frédéric Chopin (1810–1849). The listening study acts as an evaluation of the two computational models and a third, benchmark system, called Experiments in Musical Intelligence. Judges' responses indicate that some aspects of musical style, such as phrasing and rhythm, are being modeled effectively by our algorithms. Judgments are also used to identify areas for future improvements. We discuss the broader implications of this work for the fields of engineering and design, where there is potential to make use of our models of hierarchical repetitive structure.
We study the problem of detecting sentences describing adverse drug reactions (ADRs) and frame the problem as binary classification. We investigate different neural network (NN) architectures for ADR classification. In particular, we propose two new neural network models, Convolutional Recurrent Neural Network (CRNN) by concatenating convolutional neural networks with recurrent neural networks, and Convolutional Neural Network with Attention (CNNA) by adding attention weights into convolutional neural networks. We evaluate various NN architectures on a Twitter dataset containing informal language and an Adverse Drug Effects (ADE) dataset constructed by sampling from MEDLINE case reports. Experimental results show that all the NN architectures outperform the traditional maximum entropy classifiers trained from n-grams with different weighting strategies considerably on both datasets. On the Twitter dataset, all the NN architectures perform similarly. But on the ADE dataset, CNN performs better than other more complex CNN variants. Nevertheless, CNNA allows the visualisation of attention weights of words when making classification decisions and hence is more appropriate for the extraction of word subsequences describing ADRs.
Current game music systems typically involve the playback of prerecorded audio tracks which are crossfaded in response to game events such as level changes. However, crossfading can limit the expressive power of musical transitions, and can make fine grained structural variations difficult to achieve. We therefore describe an alternative approach in which music is algorithmically generated based on a set of high-level musical features that can be controlled in real-time according to a player’s progression through a game narrative. We outline an implementation of the approach in an actual game, focusing primarily on how the music system traces the game’s emotional narrative by periodically querying certain narrative parameters and adjusting the musical features of its output accordingly.
Current game music systems typically involve the playback of prerecorded audio tracks which are crossfaded in response to game events such as level changes. However, crossfading can limit the expressive power of musical transitions, and can make fine grained structural variations difficult to achieve. We therefore describe an alternative approach in which music is algorithmically generated based on a set of high-level musical features that can be controlled in real-time according to a player’s progression through a game narrative. We outline an implementation of the approach in an actual game, focusing primarily on how the music system traces the game’s emotional narrative by periodically querying certain narrative parameters and adjusting the musical features of its output accordingly.
Vincenzo Gervasi合作论文数Computer Science Department of the University of Pisa, Italy4