Within the online media universe, there are many underlying communities. These may be defined, for example, through politics, location, health, occupation, extracurricular interests or retail habits. Government departments, charities and commercial organisations can benefit greatly from insights about the structure of these communities; the move to customer-centred practices requires knowledge of the customer base. Motivated by this issue, we address the fundamental question of whether a sub-network looks like a collection of individuals who have effectively been picked at random from the whole, or instead forms a distinctive community with a new, discernible structure. In the former case, to spread a message to the intended user base it may be best to use traditional broadcast media (TV, billboard), whereas in the latter case a more targeted approach could be more effective. In this work, we therefore formalise a concept of testing for sub-structure and apply it to social interaction data. First, we develop a statistical test to determine whether a given sub-network (induced sub-graph) is likely to have been generated by sampling nodes from the full network uniformly at random. This tackles an interesting inverse alternative to the more widely studied “forward” problem. We then apply the test to a Twitter reciprocated mentions network where a range of brand name based sub-networks are created via tweet content. We correlate the computed results against the independent views of 16 digital marketing professionals. We conclude that there is great potential for social media based analytics to quantify, compare and interpret online brand allegiances systematically, in real time and at large scale.
We propose a novel mathematical model for the activity of microbloggers during an external, event-driven spike. The model leads to a testable prediction of who would become most active if a spike were to take place. This type of insight into human behaviour has many applications, as it identifies key players who can be targeted with information in real time when the network is most receptive. The model takes account of the fact that dynamic interactions evolve over an underlying, static network that records "who listens to whom". Our fundamental assumption is that, in the case where the entire community has become aware of an external news event, a key driver of activity is the motivation to participate by responding to incoming messages. We validate the resulting algorithm on a large scale Twitter conversation concerning the appointment of a UK Premier League football club manager. We also find that the half-life of a spike in activity can be quantified in terms of the network size and the typical response rate.
Online human interactions take place within a dynamic hierarchy, where social influence is determined by qualities such as status, eloquence, trustworthiness, authority and persuasiveness. In this work, we consider topic-based twitter interaction networks, and address the task of identifying influential players. Our motivation is the strong desire of many commercial entities to increase their social media presence by engaging positively with pivotal bloggers and tweeters. After discussing some of the issues involved in extracting useful interaction data from a twitter feed, we define the concept of an active node subnetwork sequence. This provides a time-dependent, topic-based, summary of relevant twitter activity. For these types of transient interactions, it has been argued that the flow of information, and hence the influence of a node, is highly dependent on the timing of the links. Some nodes with relatively small bandwidth may turn out to be key players because of their prescience and their ability to instigate follow-on network activity. To simulate a commercial application, we build an active node subnetwork sequence based on key words in the area of travel and holidays. We then compare a range of network centrality measures, including a recently proposed version that accounts for the arrow of time, with respect to their ability to rank important nodes in this dynamic setting. The centrality rankings use only connectivity information (who tweeted whom, when), without requiring further information about the account type or message content, but if we post-process the results by examining account details, we find that the time-respecting, dynamic approach, which looks at the follow-on flow of information, is less likely to be ‘misled’ by accounts that appear to generate large numbers of automatic tweets with the aim of pushing out web links. We then benchmark these algorithmically derived rankings against independent feedback from five social media experts, given access to the full tweet content, who judge twitter accounts as part of their professional duties. We find that the dynamic centrality measures add value to the expert view, and can be hard to distinguish from an expert in terms of who they place in the top ten. These algorithms, which involve sparse matrix linear system solves with sparsity driven by the underlying network structure, can be applied to very large-scale networks. We also test an extension of the dynamic centrality measure that allows us to monitor the change in ranking, as a function of time, of the twitter accounts that were eventually deemed influential.
Last year, two of us (PG and DJH) contributed to a two-part article in SIAM News about the growth of social network analysis in business and government [5]. Joined in the present article by Peter Laflin, head of Data Insight at the UK-based digital marketing agency Bloom [1], we describe some developments that followed directly from the SIAM News article. Bloom clients typically wish to monitor and improve their online social media presence. The SIAM News article alerted Bloom’s Insight team to a discussion of time-dependent networks in [4]. The matrixbased algorithms described there proved to be useful for identifying key players in the large-scale online conversations taking place on topics of interest to Bloom clients. Following this initial success, Bloom made a good old-fashioned telephone call, which has led to the mutually beneficial collaboration briefly reported here. On the academic side, researchers at the Universities of Reading and Strathclyde have advised on cutting-edge developments; Bloom’s role has been to provide examples of large data sets, along with some current and future challenges [6]. As an example of knowledge exchange driving new research, our collaboration on a Twitter data set [8,9] flagged the need to identify and categorize spam accounts that generate automated Tweets. That case study also provided the university researchers with a rare benchmarking opportunity— Twitter accounts deemed influential by the computational algorithms could be compared with those picked out from the same data set by a team of social-media experts with day-to-day experience in hand-curating this type of information. Our study found the best computational measures to be essentially indistinguishable from the selections made by human experts. This fall, Bloom presented an overview of its work at Londata, a regular “big data” event in London. Before the talk, the team ran a preprocessing step in which it assessed the influence of the people who had registered for the event, based on their Twitter footprints. The ten most influential attendees, ranked in terms of their “dynamic broadcast” score from [4], are shown in Table 1. Where available, the table also shows some other measures of influence: number of followers for each account, Klout score [7], and Peer Index score [10]. This initial analysis wasn’t conversation-specific: It simply took the 150 people registered for Londata and considered how readily messages could flow between them. The @LondataEvent account is of interest here; its third-place ranking, in our opinion, shows the account to be highly influential. It has a small following, but those people typically amplify messages received from @LondataEvent. This indicates that the dynamic broadcast measure of influence goes far deeper than a simple counting of followers. The measure can also be applied in real time, so the Insight team—at the risk of disappearing into an infinitely recursive puff of smoke—used its analytical tools to calculate influence and visualise the interactions by listening to Twitter activity during the course of Bloom’s own presentation on the topic. The team thus recorded and analysed the 290 Tweets made by participants (the audience was made up of digitally savvy types!), updating the influence scores every 5 minutes. Figure 1 shows a snapshot of part of the Collaboration Blooms from SIAM News Article
A novel way of calculating online influence has been proposed in [2,1]. Bloom Agency have created new online software capable of collecting social data and calculating these new influence metrics in real time. A demonstration of this software will be given at the conference. Delegates will be encouraged to Tweet using the #socinfo2012 hashtag and the influence of the top ten Tweeters will be shown, along with a visualisation of the evolving conversation.
Figure 1. A snapshot of part of the interaction network during the recording and analysis of 290 Tweets made during a presentation at Londata. The visualisation reveals distinct communities involved in the Londata conversation and highlights influential Twitter accounts.
We describe the results of a new computational experiment on Twitter data. By listening to Tweets on a selected topic, we generate a dynamic social interaction network. We then apply a recently proposed dynamic network analysis algorithm that ranks Tweeters according to their ability to broadcast information. In particular, we study the evolution of importance rankings over time. Our presentation will also describe the outcome of an experiment where results from automated ranking algorithms are compared with the views of social media experts.