Social networks are quickly becoming the primary medium for discussing what is happening around real-world events. The information that is generated on social platforms like Twitter can produce rich data streams for immediate insights into ongoing matters and the conversations around them. To tackle the problem of event detection, we model events as a list of clusters of trending entities over time. We describe a real-time system for discovering events that is modular in design and novel in scale and speed: it applies clustering on a large stream with millions of entities per minute and produces a dynamically updated set of events. In order to assess clustering methodologies, we build an evaluation dataset derived from a snapshot of the full Twitter Firehose and propose novel metrics for measuring clustering quality. Through experiments and system profiling, we highlight key results from the offline and online pipelines. Finally, we visualize a high profile event on Twitter to show the importance of modeling the evolution of events, especially those detected from social data streams.
Social curation is a new trend which has emerged following on the heels of the information glut created by user-generated content revolution. Rather than create new content, social curation allows users to categorise content created by others, and thereby creating and resharing their personal taxonomies of the Web. In this dissertation, we collect a large dataset from Pinterest, arguably the most popular image curation service, and seek to understand the trend on three levels: content, friends and crowds. We first take an empirical look at social curation by mining its content usage. Our data reveals that curation tends to focus on niche items that may not rank highly in popularity and search rankings. Yet, curated items exhibit their own skewed popularity, although most users, or curators, act for personal reasons. At the same time, it also shows that curators with consistent activity and diversity of interests show more social value in attracting followers. This drives us to explore the role of social networks on social curation. We find that social users are more active and are more likely to return soon in Pinterest, indicating a bonding effect enabled by social networks. Then we divide the social network into two subgraphs, according to whether they are created natively or copied from some other established social networks (e.g., Facebook) via a social bootstrapping method. It shows that, when users just join the service, copied network can promote more social interaction, as it initiates a stronger and denser social structure than native network. However, social networks are not critical for information seeking, as a non-trivial number of users’ content are curated from strangers with high interest matching. In fact, this trend also holds for social interaction: Users tend to wean from copied friends to interact more with interest-based native friends over a long-term view. Finally, we understand social curation as a distributed computation process, and examine the relationship between curators and crowds. We show that despite being categorised by individual actions, there is generally a global agreement in implicitly assigning content into a coarse-grained global taxonomy of categories, and furthermore, users tend to specialise in a handful of categories. By exploiting these characteristics, and augmenting with image-related features drawn from a state-of-the-art deep convolutional neural network, we develop a cascade of predictors that together automate a large fraction of curation actions with an end-to-end accuracy of 0.69 (Accuracy@5 of 0.75).
This paper investigates when users create profiles in different social networks, whether they are redundant expressions of the same persona, or they are adapted to each platform. Using the personal webpages of 116,998 users on About.me, we identify and extract matched user profiles on several major social networks including Facebook, Twitter, LinkedIn, and Instagram. We find evidence for distinct site-specific norms, such as differences in the language used in the text of the profile self-description, and the kind of picture used as profile image. By learning a model that robustly identifies the platform given a user’s profile image (0.657–0.829 AUC) or self-description (0.608–0.847 AUC), we confirm that users do adapt their behaviour to individual platforms in an identifiable and learnable manner. However, different genders and age groups adapt their behaviour differently from each other, and these differences are, in general, consistent across different platforms. We show that differences in social profile construction correspond to differences in how formal or informal the platform is.
The aim of this article is to provide an understanding of social networks as a useful addition to the standard toolbox of techniques used by system designers. To this end, we give examples of how data about social links have been collected and used in different application contexts. We develop a broad taxonomy-based overview of common properties of social networks, review how they might be used in different applications, and point out potential pitfalls where appropriate. We propose a framework, distinguishing between two main types of social network-based user selection— personalised user selection, which identifies target users who may be relevant for a given source node, using the social network around the source as a context, and generic user selection or group delimitation, which filters for a set of users who satisfy a set of application requirements based on their social properties. Using this framework, we survey applications of social networks in three typical kinds of application scenarios: recommender systems, content-sharing systems (e.g., P2P or video streaming), and systems that defend against users who abuse the system (e.g., spam or sybil attacks). In each case, we discuss potential directions for future research that involve using social network properties.
On most current websites untrustworthy or spammy identities are easily created. Existing proposals to detect untrustworthy identities rely on reputation signals obtained by observing the activities of identities over time within a single site or domain; thus, there is a time lag before which websites cannot easily distinguish attackers and legitimate users. In this paper, we investigate the feasibility of leveraging information about identities that is aggregated across multiple domains to reason about their trustworthiness. Our key insight is that while honest users naturally maintain identities across multiple domains (where they have proven their trustworthiness and have acquired reputation over time), attackers are discouraged by the additional effort and costs to do the same. We propose a flexible framework to transfer trust between domains that can be implemented in today's systems without significant loss of privacy or significant implementation overheads. We demonstrate the potential for inter-domain trust assessment using extensive data collected from Pinterest, Facebook, and Twitter. Our results show that newer domains such as Pinterest can benefit by transferring trust from more established domains such as Facebook and Twitter by being able to declare more users as likely to be trustworthy much earlier on (approx. one year earlier).
This paper seeks to answer the question of whether social ties are important on interest-driven social networks, by analysing 4-years of activities of 50,000 randomly sampled users on Pinterest, a social image discovery website. We find that a non-trivial number of users’ images are copied or repinned from strangers instead of friends, suggesting that social-based information exploration is not important. However, social interactions and social repins are critical for user retention: users interacting with friends are more likely to return Pinterest soon. These results suggest that the real role of social ties on interest-driven social networks is to enable bonding of users rather than seeking information.
Everyday, millions of users save content items for future use on sites like Pinterest, by "pinning" them onto carefully categorised personal pinboards, thereby creating personal taxonomies of the Web. This paper seeks to understand Pinterest as a distributed human computation that categorises images from around the Web. We show that despite being categorised onto personal pinboards by individual actions, there is a generally a global agreement in implicitly assigning images into a coarse-grained global taxonomy of 32 categories, and furthermore, users tend to specialise in a handful of categories. By exploiting these characteristics, and augmenting with image-related features drawn from a state-of-the-art deep convolutional neural network, we develop a cascade of predictors that together automate a large fraction of Pinterest actions. Our end-to-end model is able to both predict whether a user will repin an image onto her own pinboard, and also which pinboard she might choose, with an accuracy of 0.69 (Accuracy@5 of 0.75).
How does one develop a new online community that is highly engaging to each user and promotes social interaction? A number of websites offer friend-finding features that help users bootstrap social networks on the website by copying links from an established network like Facebook or Twitter. This paper quantifies the extent to which such social bootstrapping is effective in enhancing a social experience of the website. First, we develop a stylised analytical model that suggests that copying tends to produce a giant connected component (i.e., a connected community) quickly and preserves properties such as reciprocity and clustering, up to a linear multiplicative factor. Second, we use data from two websites, Pinterest and Last.fm, to empirically compare the subgraph of links copied from Facebook to links created natively. We find that the copied subgraph has a giant component, higher reciprocity and clustering, and confirm that the copied connections see higher social interactions. However, the need for copying diminishes as users become more active and influential. Such users tend to create links natively on the website, to users who are more similar to them than their Facebook friends. Our findings give new insights into understanding how bootstrapping from established social networks can help engage new users by enhancing social interactivity.
How does one develop a new online community that is highly engaging to each user and promotes social interaction? A number of websites offer friend-finding features that help users bootstrap social networks on the website by copying links from an established network like Facebook or Twitter. This paper quantifies the extent to which such social bootstrapping is effective in enhancing a social experience of the website. First, we develop a stylised analytical model that suggests that copying tends to produce a giant connected component (i.e., a connected community) quickly and preserves properties such as reciprocity and clustering, up to a linear multiplicative factor. Second, we use data from two websites, Pinterest and Last.fm, to empirically compare the subgraph of links copied from Facebook to links created natively. We find that the copied subgraph has a giant component, higher reciprocity and clustering, and confirm that the copied connections see higher social interactions. However, the need for copying diminishes as users become more active and influential. Such users tend to create links natively on the website, to users who are more similar to them than their Facebook friends. Our findings give new insights into understanding how bootstrapping from established social networks can help engage new users by enhancing social interactivity.
Copying, sharing and linking have always been important for the functioning and the growth of the World Wide Web. Two recent copying trends which have emerged are social content curation, and social logins. Social curation involves the copying, categorization and sharing of links and images from third party websites on the social curation website. Social logins enable the copying of user identities and their friends from an established social network such as Facebook or Twitter, onto third party websites. In this article, we chronicile our ongoing work on Pinterest, a popular image sharing website and social network. The highly active user community on Pinterest has been instrumental in making social curation a mainstream phenomenon. Interestingly, a large fraction (nearly 60%) of the users have also linked their Pinterest accounts with Facebook and have copied their Facebook friends over onto the new website. Thus, using a large dataset crawled from Pinterest, we uncover both the practices used for sharing content, as well as how the copying of friends has helped the content sharing. We find that social curation tends to copy and share hard-to-find niche interest content from websites with a low Alexa Rank or Google Page Rank, and curators with consistent updates and a diversity of interests are popular and attract more followers. On the other hand, Pinterest users can also copy friends from Faceebook, or Twitter. We find that this copying of friends create a community with higher levels of social interaction; thus social logins serve as a social bootstrapping tool. But beyond bootstrapping, we also find a weaning process, where active and influential users tend to form more links natively on Pinterest interact with native friends rather than copied friends.
How does one develop a new online community that is highly engaging to each user and promotes social interaction? A number of websites offer friend-finding features that help users bootstrap social networks on the website by copying links from an established network like Facebook or Twitter. This paper quantifies the extent to which such social bootstrapping is effective in enhancing a social experience of the website. First, we develop a stylised analytical model that suggests that copying tends to produce a giant connected component (i.e., a connected community) quickly and preserves properties such as reciprocity and clustering, up to a linear multiplicative factor. Second, we use data from two websites, Pinterest and Last.fm, to empirically compare the subgraph of links copied from Facebook to links created natively. We find that the copied subgraph has a giant component, higher reciprocity and clustering, and confirm that the copied connections see higher social interactions. However, the need for copying diminishes as users become more active and influential. Such users tend to create links natively on the website, to users who are more similar to them than their Facebook friends. Our findings give new insights into understanding how bootstrapping from established social networks can help engage new users by enhancing social interactivity.
This paper looks at how and why users categorise and curate content into collections online, using datasets containing nearly all the relevant activities from Pinterest.com during January 2013, and Last.fm in December 2012. In addition, a user survey of over 25 Pinterest and 250 Last.fm users is used to obtain insights into the motivations for content curation and corroborate results. The data reveal that curation tends to focus on items that may not rank highly in popularity and search rankings. Yet, curated items exhibit their own skewed popularity, with the top few items receiving most of the attention; indicative of a synchronised community. We distinguish structured curation by active categorisation from a more passive bookmarking by `liking' an item, and find the former more prevalent for popularly curated items. Likes, however, are initially accumulated at a faster pace. Finally, we study the social value of content curation and show that curators attract more followers with consistent activity, and diversity of interests. Interestingly, our user study indicates a divided opinion on the relevance of the social network.