
Access to diverse perspectives nurtures an informed citizenry. Google and Bing have emerged as the duopoly that largely arbitrates which English language documents are seen by web searchers. We present our empirical study over the search results produced by Google and Bing that shows a large overlap. Thus, citizens may not gain different perspectives by simultaneously probing them for the same query. Fortunately, our study also shows that by mining Twitter data one can obtain search results that are quite distinct from those produced by Google and Bing. Additionally, the users found those results to be quite informative.
Many systems involve the allocation of rewards for achievements, and these rewards produce a set of incentives that in turn guide behavior. Such effects are visible in many domains from everyday life, and they are increasingly forming a designed aspect of participatory on-line sites through the use of badges and other reward systems. We consider several aspects of the interaction between rewards and incentives in the context of collective effort, including a method for reasoning about on-line user activity in the presence of milestones and badges; and a graph-theoretic framework for analyzing procrastination and other forms of behavior that are inconsistent over time. The talk includes joint work with Ashton Anderson, Dan Huttenlocher, Jure Leskovec, and Sigal Oren
Data brokers have traditionally collected data from businesses, government records, and other publicly available offline sources. While each data source may provide only a few elements about a person's activities, data brokers combine these elements to form a detailed, composite view of the consumer's life. The emergence of social media gives data brokers unprecedented opportunities to enhance their profiles. Data brokers are increasingly interested in combining the information collected from offline sources with information publicly available in social networks to profile not only adults but also children. In this paper, we show how data brokers and other third parties can combine online and offline data sources -- namely, public Facebook profiles and voter registration records -- to create detailed profiles of adults, teens, and children in any target city in the US. We outline and execute an approach that leverages a Facebook user's social ties combined with the city's voter registration records to infer the Facebook users who reside in the city. These inferences enable a data broker to create detailed user profiles, which not only include information publicly available from Facebook but also the user's exact residential address, date and year of birth, and political affiliation. We further show how additional inferences can be made from the combined data. We then discuss how this city attack can be extended to create detailed profiles of minors and children. Finally, we make recommendations to Facebook, municipal authorities, and individuals to decrease the risk of this large-scale privacy breach.
Users today access a multitude of online services---among the most popular of which are online social networks (OSNs)---via both web sites and dedicated mobile applications (apps), using a range of devices (traditional PCs, tablets, and smartphones) that are connected via a variety of networks. The resulting infrastructure makes these services conveniently available anytime and anywhere, enabling them to become an integral part of daily life. As a consequence, users explicitly and implicitly provide a wealth of Personal Information (PI) that reflects several aspects of their life. Service providers monetize this information by selling to third parties (e.g., advertisers). Unfortunately, today, it remains difficult for end users to fully understand the amount and nature of the collected data. Our goal in this paper is to bring visibility into PI collected when accessing online services such as online social networks. This is a major challenge because PI is transferred in a proprietary way by each service. We develop a novel method that can automatically discover various types of PI carried within protocol fields of network traffic; the method includes techniques to filter out potential "containers" that do not actually carry PI and extend the set of containers initially found with additional ones. We evaluate the false positive/negative rates of our proposed method and show examples of interesting findings, including what kind of web sites or apps are more likely to transmit PI and which types of PI are most commonly collected.
Information sharing dynamics of social networks rely on a small set of influencers to effectively reach a large audience. Our recent results and observations demonstrate that the shape and identity of this elite, especially those contributing original content, is difficult to predict. Information acquisition is often cited as an example of a public good. However, this emerging and powerful theory has yet to provably offer qualitative insights on how specialization of users into active and passive participants occurs. This paper bridges, for the first time, the theory of public goods and the analysis of diffusion in social media. We introduce a non-linear model of perishable public goods, leveraging new observations about sharing of media sources. The primary contribution of this work is to show that shelf time, which characterizes the rate at which content get renewed, is a critical factor in audience participation. Our model proves a fundamental dichotomy in information diffusion: While short-lived content has simple and predictable diffusion, long-lived content has complex specialization. This occurs even when all information seekers are ex ante identical and could be a contributing factor to the difficulty of predicting social network participation and evolution.
Diffusion in social networks has been studied extensively in the past few years. Most previous work assumes that the underlying network is a static object that remains unchanged as the diffusion process progresses. However, there are several real-life networks that change dynamically over time. In this paper, we study diffusion on such evolving networks and extend the popular Independent Cascade and Linear Threshold models to account for network evolution. In particular, we introduce two natural variations, a persistent and a transient one, to capture diffusions of different types. We consider the problem of influence maximization where the goal is to select a few influential nodes to initiate a diffusion with maximum spread. We show that, surprisingly, when considering evolving networks the diffusion function is no longer submodular for the transient models, and not even monotone for the transient Independent Cascade model. We also show that, depending on the model, delaying the activation of the initiators may improve diffusion. Our experiments, using three real datasets, demonstrate the effect of network evolution on the diffusion process, and highlight the importance of timing in the selection process.
Recently, graph matching algorithms have been successfully applied to the problem of network de-anonymization, in which nodes (users) participating in more than one social network are identified only by means of the structure of their links to other members. This procedure exploits an initial set of seed nodes large enough to trigger a percolation process which correctly matches almost all other nodes across the different social networks. Our main contribution is to show the crucial role played by clustering, which is a ubiquitous feature of realistic social network graphs (and many other systems). Clustering has both the effect of making matching algorithms more vulnerable to errors, and the potential to dramatically reduce the number of seeds needed to trigger percolation, thanks to a wave-like propagation effect. We demonstrate these facts by considering a fairly general class of random geometric graphs with variable clustering level, and showing how clever algorithms can achieve surprisingly good performance while containing matching errors.
We investigate the relation between the sentiment of a message on social media and its virality, defined as the volume and the speed of message diffusion. We analyze 4.1 million messages (tweets) obtained from Twitter. Although factors affecting message diffusion on social media have been studied previously, we focus on message sentiment, and reveal how the polarity of message sentiment affects its virality. The virality of a message is measured by the number of message repostings (retweets) and the time elapsed from the original posting of a message to its Nth reposting (N-retweet time). Through extensive analyses, we find that negative messages are likely to be reposted more rapidly and frequently than positive and neutral messages. Specifically, the reposting volume of negative messages is 1.2--1.6-fold that of positive and neutral messages, and negative messages spread at 1.25 times the speed of positive and neutral messages when the diffusion volume is large.
Measurement studies of online social networks show that all sociallinks are not equal, and the strength of each link is best characterized by the frequency of interactions between the linked users. To date, few studies have been able to examine detailed interaction data over time, and studied the problem of modeling user interactions. A generative model can shed light on the fundamental processes that underlie user interactions. In this paper, we analyze the first complete record of full interaction and network dynamics in a large online social network. Our dataset covers all wall posts, new user events, and new social link events during the first full year of Renren, the largest social network in China, including 623K new users, 8.2 million new links, and 29 million wall posts. Our analysis provides surprising insights into the evolution of user interactions over time. We find that users invite new friends to interact at a nearly constant rate, prefer to interact with friends with whom they share significant overlaps in social circles, and most social links drop in interaction frequency over time. We also validate our findings on Facebook, and show that they do generalize across OSNs. We use our insights to derive a generative model of social interactions that accurately captures both our new results and previously observed network properties. Our model captures the inherently heterogeneous strengths of social links, and has broad implications on the design of social network algorithms such as friend recommendation, information diffusion and viral marketing.
One of the defining features of online social networks is that users’ actions are visible to other users. In this paper, we argue that such social visibility has a detrimental effect on users’ willingness to gift digital goods. The gift giving process often generates substantial anxiety, and social visibility exacerbates this anxiety to the point that it can deter gifting altogether. To study the effect of social visibility on the decision to gift, we analyze a unique dataset from a large online social network that offers users the option of buying a digital gifting service. We find that purchase rates of the service increased with the number of social ties that users kept on the network, but decreased with the extent to which those ties were tied to each other. We argue that the latter effect is due to the fact that, when a user’s ties are tied themselves, any gift sent between the user and one tie is visible to their mutual contacts. This argument is bolstered by a stronger negative effect of social visibility for users with larger, less intimate, and categorically diverse networks.
Popular social and e-commerce sites increasingly rely on crowd computing to rate and rank content, users, products and businesses. Today, attackers who create fake (Sybil) identities can easily tamper with these computations. Existing defenses that largely focus on detecting individual Sybil identities have a fundamental limitation: Adaptive attackers can create hard-to-detect Sybil identities to tamper arbitrary crowd computations. In this paper, we propose Stamper, an approach for detecting tampered crowd computations that significantly raises the bar for evasion by adaptive attackers. Stamper design is based on two key insights: First, Sybil attack detection gains strength in numbers: we propose statistical analysis techniques that can determine if a large crowd computation has been tampered by Sybils, even when it is fundamentally hard to infer which of the participating identities are Sybil. Second, Sybil identities cannot forge the timestamps of their activities as they are recorded by system operators; Stamper analyzes these unforgeable timestamps to foil adaptive attackers. We applied Stamper to detect tampered computations in Yelp and Twitter. We not only detected previously known tampered computations with high accuracy, but also uncovered tens of thousands of previously unknown tampered computations in these systems.
Information extracted from social media streams has been leveraged to forecast the outcome of a large number of real-world events, from political elections to stock market fluctuations. An increasing amount of studies demonstrates how the analysis of social media conversations provides cheap access to the wisdom of the crowd. However, extents and contexts in which such forecasting power can be effectively leveraged are still unverified at least in a systematic way. It is also unclear how social-media-based predictions compare to those based on alternative information sources. To address these issues, here we develop a machine learning framework that leverages social media streams to automatically identify and predict the outcomes of soccer matches. We focus in particular on matches in which at least one of the possible outcomes is deemed as highly unlikely by professional bookmakers. We argue that sport events offer a systematic approach for testing the predictive power of social media conversations, and allow to compare such power against the rigorous baselines set by external sources. Despite such strict baselines, our framework yields above 8% marginal profit when used to inform simple betting strategies. The system is based on real-time sentiment analysis and exploits data collected immediately before the game start, allowing for bets informed by its predictions. We first discuss the rationale behind our approach, then describe the learning framework, its prediction performance and the return it provides as compared to a set of betting strategies. To test our framework we use both historical Twitter data from the 2014 FIFA World Cup games (10% sample), and real-time Twitter data (full stream) collected by monitoring the conversations about all soccer matches of the four major European tournaments (FA Premier League, Serie A, La Liga, and Bundesliga), and the 2014 UEFA Champions League, during the period between October, 25th 2014 and November, 26th 2014.
From network topologies to online social networks, many of today's most sensitive datasets are captured in large graphs. A significant challenge facing the data owners is how to share sensitive graphs with collaborators or authorized users, e.g. ISP's network topology graphs with a third party networking equipment vendor. Current tools can provide limited node or edge privacy, but significantly modify the graph reducing its utility. In this work, we propose a new alternative in the form of graph watermarks . Graph watermarks are small graphs tailor-made for a given graph dataset, a secure graph key, and a secure user key. To share a sensitive graph G with a collaborator C , the owner generates a watermark graph W using G , the graph key, and C 's key as input, and embeds W into G to form G' . If G' is leaked by C , its owner can reliably determine if the watermark W generated for C does in fact reside inside G' , thereby proving C is responsible for the leak. Graph watermarks serve both as a deterrent against data leakage and a method of recourse after a leak. We provide robust schemes for embedding and extracting watermarks, and use analysis and experiments on large real graphs to show that they are unique and difficult to forge. We study the robustness of graph watermarks against both single and powerful colluding attacker models, then propose and evaluate mechanisms to dramatically improve resilience.
Pinterest provides a social curation service where people can collect, organize, and share content (pins in Pinterest) that reflect their interests. This paper investigates (1) the differences in pinning (i.e., the act of posting a pin) and repinning (i.e., the act of sharing other user's pin) behaviors by topics and user gender, and (2) the relations among topics in Pinterest. We conduct a measurement study using a large-scale dataset (1.6 M pins shared by 1.1 M users) in Pinterest. We show that there is a notable discrepancy between pinning and repinning behaviors on different topics. We also show that male and female users show different behaviors on different topics in terms of dedication, responsiveness, and sentiment. By introducing the notion of a Topic Network (TN) whose nodes are topics and are linked if they share common users, we analyze how topics are related to one another, which can give a valuable implication on topic demand forecasting or cross-topic advertisement. Lastly, we explore the implications of our findings for predicting a user's interests and behavioral patterns in Pinterest.
Online interactions are increasingly involving images, especially those containing human faces, which are naturally attention grabbing and more effective at conveying feelings than text. To understand this new convention of digital culture, we study the collective behavior of sharing \textit{selfies} on Instagram and present how people appear in selfies and which patterns emerge from such interactions. Analysis of millions of photos shows that the amount of selfies has increased by 900 times from 2012 to 2014. Selfies are an effective medium to grab attention; they generate on average 1.1--3.2 times more likes and comments than other types of content on Instagram. Compared to other content, interactions involving selfies exhibit variations in homophily scores (in terms of age and gender) that suggest they are becoming more widespread. Their style also varies by cultural boundaries in that the average age and majority gender seen in selfies differ from one country to another. We provide explanations of such country-wise variations based on cultural and socioeconomic contexts.
Online reviews are an important source for consumers to evaluate products/services on the Internet (e.g. Amazon, Yelp, etc.). However, more and more fraudulent reviewers write fake reviews to mislead users. To maximize their impact and share effort, many spam attacks are organized as campaigns, by a group of spammers. In this paper, we propose a new two-step method to discover spammer groups and their targeted products. First, we introduce NFS (Network Footprint Score), a new measure that quantifies the likelihood of products being spam campaign targets. Second, we carefully devise GroupStrainer to cluster spammers on a 2-hop subgraph induced by top ranking products. We demonstrate the efficiency and effectiveness of our approach on both synthetic and real-world datasets from two different domains with millions of products and reviewers. Moreover, we discover interesting strategies that spammers employ through case studies of our detected groups.
Access to diverse perspectives nurtures an informed citizenry. Google and Bing have emerged as the duopoly that largely arbitrates which English language documents are seen by web searchers. We present our empirical study over the search results produced by Google and Bing that shows a large overlap. Thus, citizens may not gain different perspectives by simultaneously probing them for the same query. Fortunately, our study also shows that by mining Twitter data one can obtain search results that are quite distinct from those produced by Google and Bing. Additionally, the users found those results to be quite informative.
Location-Based Social Networks (LBSN) such as Foursquare allow users to indicate venue visits via check-ins. This results in much fine grained context-rich data, useful for studying user mobility. In this work, we use check-ins to characterize trips and visitors to two cities, where visitors are defined as having their home cities elsewhere. First, we divide trips into two duration types: long and short. We then show that trip types differ in check-in distributions over venue categories, time slots, as well as check-in intensity. Based on the trip types, we then divide visitors into long-term and short-term visitors. We compare visitor types in terms of popularities of check-in venues and proximities to friends' check-ins. Our findings indicate that short-term visitors are more biased towards popular venues. As for proximity to friends' check-ins, the effect is more consistently observed for long-term visitors. These findings also illustrate that locations of incoming visitors can effectively be analyzed using LBSN data in addition to conducting user surveys which are relatively costlier. Lastly, we investigate the importance of visitor type information in models for venue prediction. We apply models including a state of the art kernel density estimation technique and ranking based on venue popularity. For each model, we consider two settings where visitor type information is absent/present. For long-term visitors, we observed little differences in accuracies. However, for short-term visitors, predictions are significantly more accurate by using type information. These findings suggest that venue prediction or recommender systems should consider visitor type to improve accuracy.
The recent growth of anonymous messaging services -- such as 4chan, Whisper, and Yik Yak -- has brought online anonymity into the spotlight. Ideally, anonymous posts in principle allow for fully open discussion, and facilitate the sharing of controversial ideas. In this research, we attempt to exploit the location-based feature of Yik Yak to determine the location from which a yak is posted. We conducted our experiment across Missoula, home of the University of Montana by using six computers to probe from the 2,880 locations and record the yaks seen. Our preliminary results show that Yik Yak, and location-based services like it, potentially have a major privacy flaw. By knowing the general vicinity of the yak, other side information -- such as the content of the yak along with the locations of college dorms, off-campus residences and local coffee shops -- may allow one to further narrow down where the yak is coming from. A more in-depth experiment and analysis, possibly using a denser honeycomb and machine learning, may be able to predict yak locations with much greater accuracy.
In everyday life, we often observe unusually frequent interactions among people before or during important events, e.g., we receive/send more greetings from/to our friends on Christmas Day, than usual. We also observe that some videos suddenly go viral through people's sharing in online social networks (OSNs). Do these seemingly different phenomena share a common structure? All these phenomena are associated with sudden surges of user activities in networks, which we call "bursts" in this work. We find that the emergence of a burst is accompanied with the formation of triangles in networks. This finding motivates us to propose a new method to detect bursts in OSNs. We first introduce a new measure, "triadic cardinality distribution", corresponding to the fractions of nodes with different numbers of triangles, i.e., triadic cardinalities, within a network. We demonstrate that this distribution changes when a burst occurs, and is naturally immunized against spamming social-bot attacks. Hence, by tracking triadic cardinality distributions, we can reliably detect bursts in OSNs. To avoid handling massive activity data generated by OSN users, we design an efficient sample-estimate solution to estimate the triadic cardinality distribution from sampled data. Extensive experiments conducted on real data demonstrate the usefulness of this triadic cardinality distribution and the effectiveness of our sample-estimate solution.