We use de-identified data from Facebook Groups to study and provide a descriptive analysis of local gift-giving communities, in particular Buy Nothing groups. These communities allow people to give items they no longer need, reduce waste, and connect to local community. Millions of people have joined Buy Nothing groups on Facebook, with an increasing pace through the COVID-19 pandemic. Buy Nothing groups are more popular in dense and urban US counties with higher educational attainment. Compared to other local groups, Buy Nothing groups have lower Facebook friendship densities, suggesting they bring together people who are not already connected. The interaction graphs in Buy Nothing groups form larger strongly connected components, indicative of norms of generalized reciprocity. The interaction patterns in Buy Nothing groups are similar to other local online gift-giving groups, with names containing terms such as `free stuff" and `pay it forward". This points to an interaction signature for local online gift-giving communities.
We present a de-identified and aggregated dataset based on geographical patterns of Facebook Groups usage and demonstrate its association with measures of social capital. The dataset is aggregated at United States county level. Established spatial measures of social capital are known to vary across US counties. Their availability and recency depends on running costly surveys. We examine to what extent a dataset based on usage patterns of Facebook Groups, which can be generated at regular intervals, could be used as a partial proxy by capturing local online associations. We identify four main latent factors that distinguish Facebook group engagement by county, obtained by exploratory factor analysis. The first captures small and private groups, dense with friendship connections. The second captures very local and small groups. The third captures non-local, large, public groups, with more age mixing. The fourth captures partially local groups of medium to large size. Only two of these factors, the first and third, correlate with offline community level social capital measures, while the second and fourth do not. Together and individually, the factors are predictive of offline social capital measures, even controlling for various demographic attributes of the counties. To our knowledge this is the first systematic test of the association between offline regional social capital and patterns of online community engagement in the same regions. By making the dataset available to the research community, we hope to contribute to the ongoing studies in social capital.
In this work, we study the use of Twitter by House, Senate and gubernatorial candidates during the midterm (2010) elections in the U.S. Our data includes almost 700 candidates and over 690k documents that they produced and cited in the 3.5 years leading to the elections. We utilize graph and text mining techniques to analyze differences between Democrats, Republicans and Tea Party candidates, and suggest a novel use of language modeling for estimating content cohesiveness. Our findings show significant differences in the usage patterns of social media, and suggest conservative candidates used this medium more effectively, conveying a coherent message and maintaining a dense graph of connections. Despite the lack of party leadership, we find Tea Party members display both structural and language-based cohesiveness. Finally, we investigate the relation between network structure, content and election results by creating a proof-of-concept model that predicts candidate victory with an accuracy of 88.0%.
People assume different and important roles within social networks. Some roles have received extensive study: that of influencers who are well-connected, and that of brokers who bridge unconnected parts of the network. However, very little work has explored another potentially important role, that of creating opportunities for people to interact and facilitating conversation between them. These individuals bring people together and act as social catalysts. In this paper, we test for the presence of social catalysts on the online social network Facebook. We first identify posts that have spurred conversations between the poster's friends and summarize the characteristics of such posts. We then aggregate the number of catalyzed comments at the poster level, as a measure of the individual's "catalystness." The top 1% of such individuals account for 31% of catalyzed interactions, although their network characteristics do not differ markedly from others who post as frequently and have a similar number of friends. By collecting survey data, we also validate the behavioral measure of catalystness: a person is more likely to be nominated as a social catalyst by their friends if their posts prompt discussions between other people more frequently. The measure, along with other conversation-related features, is one of the most predictive of a person being nominated as a catalyst. Although influencers and brokers may have gotten more attention for their network positions, our findings provide converging evidence that another important role exists and is recognized in online social networks.
The present disclosure relates to systems and methods for increasing messaging activity in a messaging system. Using the interactions of users with each other and/or with the messaging system, the disclosed systems and methods can predict how likely a pairing of two or more users are to engage in a highly active messaging thread. Based on this prediction, the disclosed methods and systems can, for example, more effectively organize contact lists and conduct promotional efforts associated with messaging features.
Anecdotally, social connections made in university have life-long impact. Yet knowledge of social networks formed in college remains episodic, due in large part to the difficulty and expense involved in collecting a suitable dataset for comprehensive analysis. To advance and systematize insight into college social networks, we describe a dataset of the largest online social network platform used by college students in the United States. We combine de-identified and aggregated Facebook data with College Scorecard data, campus-level information provided by U.S. Department of Education, to produce a dataset covering the 2008-2015 entry year cohorts for 1,159 U.S. colleges and universities, spanning 7.6 million students. To perform the difficult task of comparing these networks of different sizes we develop a new methodology. We compute features over sampled ego-graphs, train binary classifiers for every pair of graphs, and operationalize distance between graphs as predictive accuracy. Social networks of different year cohorts at the same school are structurally more similar to one another than to cohorts at other schools. Networks from similar schools have similar structures, with the public/private and graduation rate dimensions being the most distinguishable. We also relate school types to specific outcomes. For example, students at private schools have larger networks that are more clustered and with higher homophily by year. Our findings may help illuminate the role that colleges play in shaping social networks which partly persist throughout people's lives.
In the U.S., a significant portion of many people's life-long social networks is formed in college. Yet our understanding of many aspects of this formation process, such as the role of time variation, heterogeneity between educational contexts, and the persistence of ties formed during college, is incomplete. In order to help fill some of these gaps, we use a population-level dataset of the social networks of 1,181 U.S. institutions of higher education, ranging from 2008 to 2019, to provide a detailed view of how the structure of college networks changes over time. The most prominent feature in the evolution of these networks is the burst in friending activity when students first enter college. Ties formed during this period play a strong role in shaping the structure of the networks overall and the students' position within them. Subsequent starts and breaks from instruction further affect the volume of new tie formation. Homophily in tie formation likewise shows variation in time. Same-gender ties are more likely to form when students settle into housing, while sharing a major spurs friendships as students progress through their degree. Properties of the college, such as whether many students live on campus, also modulate these effects. Ties that form in different contexts and at different points in students' college lives vary in their likelihood of remaining close years after graduation. Together, these findings suggest that educational context mediates network formation in multiple different ways.
In the classic “influence-maximization” (IM) problem, people influence one another to adopt a product and the goal is to identify people to “seed” with the product so as to maximize long-term adoption. Many influence-maximization models suggest that, if the number of people who can be seeded is unconstrained, then it is optimal to seed everyone at the start of the IM process. In a recent paper, we argued that this is not necessarily the case for social products that people use to communicate with their friends (Iyer and Adamic, The costs of overambitious seeding of social products. In: International Workshop on Complex Networks and Their Applications_273–286, 2018). Through simulations of a model in which people repeatedly use such a product and update their rate of subsequent usage depending upon their satisfaction, we showed that overambitious seeding can result in people adopting in suboptimal contexts, having bad experiences, and abandoning the product before more favorable contexts for adoption arise. Here, we extend that earlier work by showing that the costs of overambitious seeding also appear in more traditional threshold models of collective behavior, once the possibility of permanent abandonment of the product is introduced. We further demonstrate that these costs can be mitigated by using conservative seeding approaches besides those that we explored in the earlier paper. Synthesizing these results with other recent work in this area, we identify general principles for when overambitious seeding can be of concern in the deployment of social products.
Social ties form the bedrock of the global economy and international political order. Understanding the nature of these ties is thus a focus of social science research in fields including economics, sociology, political science, geography, and demography. Yet prior empirical studies have been constrained by a lack of granular data on the interconnections between individuals; most existing work instead uses indirect proxies for international ties such as levels of international trade or air passenger data. In this study, using several billion domestic and international Facebook friendships, we explore in detail the relationship between international social ties and human mobility. Our findings suggest that long-term migration accounts for roughly 83% of international ties on Facebook. Migrants play a critical role in bridging international social networks.
Product-adoption scenarios are often theoretically modeled as "influence-maximization" (IM) problems, where people influence one another to adopt and the goal is to find a limited set of people to "seed" so as to maximize long-term adoption. In many IM models, if there is no budgetary limit on seeding, the optimal approach involves seeding everybody immediately. Here, we argue that this approach can lead to suboptimal outcomes for "social products" that allow people to communicate with one another. We simulate a simplified model of social-product usage where people begin using the product at low rates and then ramp their usage up or down depending upon whether they are satisfied with their experiences. We show that overambitious seeding can result in people adopting in suboptimal contexts, where their friends are not active often enough to produce satisfying experiences. We demonstrate that gradual seeding strategies can do substantially better in these regimes.
Large cascades can develop in online social networks as people share information with one another. Though simple reshare cascades have been studied extensively, the full range of cascading behaviors on social media is much more diverse. Here we study how diffusion protocols, or the social exchanges that enable information transmission, affect cascade growth, analogous to the way communication protocols define how information is transmitted from one point to another. Studying 98 of the largest information cascades on Facebook, we find a wide range of diffusion protocols - from cascading reshares of images, which use a simple protocol of tapping a single button for propagation, to the ALS Ice Bucket Challenge, whose diffusion protocol involved individuals creating and posting a video, and then nominating specific others to do the same. We find recurring classes of diffusion protocols, and identify two key counterbalancing factors in the construction of these protocols, with implications for a cascade's growth: the effort required to participate in the cascade, and the social cost of staying on the sidelines. Protocols requiring greater individual effort slow down a cascade's propagation, while those imposing a greater social cost of not participating increase the cascade's adoption likelihood. The predictability of transmission also varies with protocol. But regardless of mechanism, the cascades in our analysis all have a similar reproduction number (≈ 1.8), meaning that lower rates of exposure can be offset with higher per-exposure rates of adoption. Last, we show how a cascade's structure can not only differentiate these protocols, but also be modeled through branching processes. Together, these findings provide a framework for understanding how a wide variety of information cascades can achieve substantial adoption across a network.
Systems, methods, and non-transitory computer-readable media can receive a plurality of comments to a posted content item. Each of the plurality of comments is associated with at least one category of a plurality of categories based on a machine learning model. A first comment of the plurality of comments is selected for inclusion in a comment sample to be presented in a graphical user interface based on the first comment being associated with a first category of the plurality of categories. A second comment of the plurality of comments is selected for inclusion in the comment sample based on the second comment being associated with a second category of the plurality of categories.
In traditional models for word-of-mouth recommendations and viral marketing, the objective function has generally been based on reaching as many people as possible. However, a number of studies have shown that the indiscriminate spread of a product by word-of-mouth can result in overexposure, reaching people who evaluate it negatively. This can lead to an effect in which the over-promotion of a product can produce negative reputational effects, by reaching a part of the audience that is not receptive to it. How should one make use of social influence when there is a risk of overexposure? In this paper, we develop and analyze a theoretical model for this process; we show how it captures a number of the qualitative phenomena associated with overexposure, and for the main formulation of our model, we provide a polynomial-time algorithm to find the optimal marketing strategy. We also present simulations of the model on real network topologies, quantifying the extent to which our optimal strategies outperform natural baselines.
In this paper, we analyse the time series of 12,000+ networks of traders in the E-mini S&P 500 stock index futures contract and we empirically link network variables with financial variables more commonly used to describe market conditions. We show that network variables lead trading volume, intertrade duration, effective spreads, trade imbalances and other market liquidity measures. Network variables reflect information, information asymmetry and market liquidity and significantly presage future market conditions prior to volume or liquidity measures. We also find two-way Granger-causality between network variables and both returns and volatility, highlighting strong feedback between market conditions and trading behaviour.
Detecting large reshare cascades is an important problem in online social networks. There are a variety of attempts to model this problem, from using time series analysis methods to stochastic processes. Most of these approaches heavily depend on the underlying network features and use network information to detect the virality of cascades. In most cases, however, getting such detailed network information can be hard or even impossible. In contrast, in this paper, we propose SANSNET, a network-agnostic approach instead. Our method can be used to answer two important questions: (1) Will a cascade go viral? and (2) How early can we predict it? We use techniques from survival analysis to build a supervised classifier in the space of survival probabilities and show that the optimal decision boundary is a survival function. A notable feature of our approach is that it does not use any network-based features for the prediction tasks, making it very cheap to implement. Finally, we evaluate our approach on several real-life data sets, including popular social networks like Facebook and Twitter, on metrics like recall, F-measure and breakout coverage. We find that network agnostic SANSNET classifier outperforms several non-trivial competitors and baselines which utilize network information.
When people share updates with their friends on Facebook they have varying expectations for the feedback they will receive. In this study, we quantitatively examine the factors contributing to feedback expectations and the potential outcomes of expectation fulfillment. We conducted two sets of surveys: one asking people about their feedback expectations immediately after posting on Facebook and the other asking how the amount of feedback received on a post matched the participant's expectations. Participants were more likely to expect feedback on content they evaluated as more important, and to a lesser extent more personal. Expectations also depended on participants' age, gender, and level of activity on Facebook. When asked about feedback expectations from specific friends, participants were more likely to expect feedback from closer friends, but expectations varied considerably based on recency of communication, geographical proximity, and the type of relationship (e.g. family, co-worker). Finally, receiving more feedback relative to expectations correlated with a greater feeling of connectedness to one's Facebook friends. The findings suggest implications for the theory and the design of social network sites.
Online activity is characterized by regularities such as diurnal and weekly patterns, reflecting human circadian rhythms and work and leisure schedules. Using data from the online social networking site Facebook, we uncover temporal patterns at a much smaller time scale: within individual sessions. Longer sessions have different characteristics than shorter ones, and this distinction is already visible in the first minute of a person's session activity. This allows us to predict the ultimate length of his or her session and how much content the person will see. The length of the session and other factors are in turn predictive of when the individual will return. Within a session, the amount of time a person spends on different kinds of content depends on both the person's demographic attributes, such as age and the number of Facebook friends, and the length of the time elapsed since the start of the session. We also find that liking and commenting is very non-uniformly distributed between sessions. Predictions of session duration and activity can potentially be leveraged to more efficiently cache content, especially to mobile devices in places with poor communications infrastructure, in order to improve user online experience.
Cascades of information-sharing are a primary mechanism by which content reaches its audience on social media. In this talk, I will describe three large-scale analyses of reshare cascades on Facebook, which were performed in aggregate using de-identified data. The first study aims to understand how predictable the growth of cascades is. We formulate the problem as one of predicting whether a cascade will double in size, and find that the prediction accuracy increases the longer a cascade has been observed. Furthermore, temporal and structural features of the cascade, as well as properties of its origin and content, along with the characteristics of those participating, are all useful in predicting how much more a cascade will grow. If we examine these cascades over significantly longer time scales, we find that many large cascades recur, exhibiting multiple bursts of popularity with periods of quiescence in between. We characterize recurrence by measuring the time elapsed between bursts, their overlap and proximity in the social network, and the diversity in the demographics of individuals participating in each peak. We discover that content virality, as revealed by its initial popularity, is a main driver of recurrence, with the availability of multiple copies of that content helping to spark new bursts. Still, beyond a certain popularity of content, the rate of recurrence drops as cascades start exhausting the population of interested individuals. We reproduce these observed patterns in a simple model of content recurrence simulated on a real social network. Using only characteristics of a cascade’s initial burst, we demonstrate strong performance in predicting whether it will recur in the future. Finally, I will discuss not just how information is transmitted perfectly, but how it evolves as changes are made as it is copied. Using a dataset of thousands of memes collectively replicated hundreds of millions of times, we find that the information undergoes an evolutionary process that exhibits several regularities. A meme’s mutation rate characterizes the population distribution of its variants, in accordance with the Yule process. Variants further apart in the diffusion cascade have greater edit distance, as would be expected in an iterative, imperfect replication process. Some text sequences can confer a replicative advantage; these sequences are abundant and transfer “laterally” between different memes. Subpopulations of the social network can preferentially transmit a specific variant of a meme if the variant matches their beliefs or culture. Understanding the mechanism driving change in diffusing information has important implications for how we interpret and harness the information that reaches us through our social networks.
Social networks readily transmit information, albeit with less than perfect fidelity. We present a large-scale measurement of this imperfect information copying mechanism by examining the dissemination and evolution of thousands of memes, collectively replicated hundreds of millions of times in the online social network Facebook. The information undergoes an evolutionary process that exhibits several regularities. A meme's mutation rate characterizes the population distribution of its variants, in accordance with the Yule process. Variants further apart in the diffusion cascade have greater edit distance, as would be expected in an iterative, imperfect replication process. Some text sequences can confer a replicative advantage; these sequences are abundant and transfer "laterally" between different memes. Subpopulations of the social network can preferentially transmit a specific variant of a meme if the variant matches their beliefs or culture. Understanding the mechanism driving change in diffusing information has important implications for how we interpret and harness the information that reaches us through our social networks.
Cascades of information-sharing are a primary mechanism by which content reaches its audience on social media, and an active line of research has studied how such cascades, which form as content is reshared from person to person, develop and subside. In this paper, we perform a large-scale analysis of cascades on Facebook over significantly longer time scales, and find that a more complex picture emerges, in which many large cascades recur, exhibiting multiple bursts of popularity with periods of quiescence in between. We characterize recurrence by measuring the time elapsed between bursts, their overlap and proximity in the social network, and the diversity in the demographics of individuals participating in each peak. We discover that content virality, as revealed by its initial popularity, is a main driver of recurrence, with the availability of multiple copies of that content helping to spark new bursts. Still, beyond a certain popularity of content, the rate of recurrence drops as cascades start exhausting the population of interested individuals. We reproduce these observed patterns in a simple model of content recurrence simulated on a real social network. Using only characteristics of a cascade's initial burst, we demonstrate strong performance in predicting whether it will recur in the future.
Ching-Yung Lin (林清詠)合作论文数IBM T. J. Watson Rsearch Center2