Q&A forums pool massive amounts of crowd expertise from a broad spectrum of geographical, cultural, and disciplinary knowledge toward specific, user-posed questions. Existing studies on these forums focus on how to route questions to the best answerers based on content or predict whether a question will be answered, but few of them investigated the inherent knowledge sharing relationship among users. We study knowledge sharing among users of StackOverflow, a popular Q&A forum, where the knowledge sharing process is related to the time elapsed since a question was posted, the reputation of the questioner, and the content of the posted text. Taking these factors into consideration, the paper proposes time-based information sharing model (TISM), where the likelihood a user will share or provide knowledge to another is modeled as a continuous function of time, reputation, and post length. With the resulting knowledge sharing network learned by TISM, we are able to predict for a given question the number of responses over time, who will answer the question and who will provide the accepted answer. Our experiments show that predictions using TISM outperform NetRate, query likelihood language, random forest, and linear regression models.
Determining cascade size and the factors affecting cascade size are two fundamental research problems in social network analysis. The commonly considered independent cascade model, when applied to social networks such as Digg, produces a phase-transition phenomenon where the cascade is either very small or very large. This phenomenon can be explained based on the concept of Giant Propagation Component (GPC). The GPC is defined as a maximally connected component, such that, by applying the independent cascade model, once any node of the component is infected, most of the remaining nodes in the component will eventually become infected with a high probability. While GPC exists in social networks, the phase-transition phenomenon, is not observed in the actual cascade size distribution when the information propagation is due to actions such as ``like'' or ``dig''. This paper hypothesizes that the cascade process, i.e., the likeliness of a node being infected changes over time and depends on how far away the node is from the seed. Furthermore, each node will not be exactly independently considered for infection from each of its infected friends, because the chance of information propagation through ``like'' or ``dig'' does not necessarily increase when there are more friends like/dig the information. To this end, we develop and simulate a new non-independent infection cascade process. The experiment results show that the proposed cascade process generates power-law like cascade size distribution without phase transition, which resembles much better the real-world cascade distribution observed in the Digg social network.
Selective Forwarding (SF) attacks impact the data transmission integrity by not forwarding a subset of received packets from time to time. The `selective' characteristic makes SF attacks hard to be distinguished from the normal packet drops or poor receptions in a volatile wireless environment. To understand this stealthy attack, an analytical model is developed to estimate the wellness of a node's forwarding behavior. Further analysis examines the cases where multiple nodes launch SF attacks in a Wireless Sensor Network (WSN), where we borrow the idea of the PageRank algorithm to estimate the most susceptible nodes to SF attacks in a network. Based on the analyses, we develop a novel reactive routing scheme that bypass suspicious nodes by estimating parent node's reliability and link quality in an integrated manner. The proposed scheme is compared to traditional approaches that also use the Node Reliability Estimator (NRE). The simulation results show that the bypass scheme provides resilience to SF attacks by achieving over 95% data delivery ratio (DDR) consistently and signif cantly outperforms the baseline Collection Tree Protocol and other algorithms without incurring additional overhead across a comprehensive set of test scenarios.
When a novel research topic emerges, we are interested in discovering how the topic will propagate over the bibliography network, i.e., which author will research and publish about this topic. Inferring the underlying influence network among authors is the basis of predicting such topic adoption. Existing works infer the influence network based on past adoption cascades, which is limited by the amount and relevance of cascades collected. This work hypothesizes that the influence network structure and probabilities are the results of many factors including the social relationships and topic popularity. These heterogeneous information shall be optimized to learn the parameters that define the homogeneous influence network that can be used to predict future cascade. Experiments using DBLP data demonstrate that the proposed method outperforms the algorithm based on traditional cascade network inference in predicting novel topic adoption.
Twitter has become a key social media for sharing information, not only for casual conversations but also for business and technologies. As the Twitter community continues to grow, an intriguing question is to determine how to obtain most valuable information the earliest by following fewest Tweeters or Tweets. This multi-criteria optimization problem exhibits similar features as in the information cascade problem for blogs. This work revises an information cascade outbreak detection algorithm to find critical Twitter accounts that disseminate the most cyber vulnerabilities the earliest. Three award functions are defined to evaluate every account’s contribution per topic from three aspects: timeliness, originality and influence. Critical users are selected according to their total contribution on a specific security category. Experiments were conducted using Tweets containing CVE information over a five-week period, to compare the proposed algorithm with account selections based on the number of followers and based on the PageRank algorithm. The results show that with the same number of users and tweets, our algorithm outperforms in both information coverage and timeliness.
Predictive analytics in situation awareness requires an element to comprehend and anticipate potential adversary activities that might occur in the future. Most work in high level fusion or predictive analytics utilizes machine learning, pattern mining, Bayesian inference, and decision tree techniques to predict future actions or states. The emergence of social computing in broader contexts has drawn interests in bringing the hypotheses and techniques from social theory to algorithmic and computational settings for predictive analytics. This paper aims at answering the question on how influence and attitude (some interpreted such as intent) of adversarial actors can be formulated and computed algorithmically, as a higher level fusion process to provide predictions of future actions. The challenges in this interdisciplinary endeavor include drawing existing understanding of influence and attitude in both social science and computing fields, as well as the mathematical and computational formulation for the specific context of situation to be analyzed. The study of ‘influence’ has resurfaced in recent years due to the emergence of social networks in the virtualized cyber world. Theoretical analysis and techniques developed in this area are discussed in this paper in the context of predictive analysis. Meanwhile, the notion of intent, or ‘attitude’ using social theory terminologies, is a relatively uncharted area in the computing field. Note that a key objective of predictive analytics is to identify impending/planned attacks so their ‘impact’ and ‘threat’ can be prevented. In this spirit, indirect and direct observables are drawn and derived to infer the influence network and attitude to predict future threats. This work proposes an integrated framework that jointly assesses adversarial actors’ influence network and their attitudes as a function of past actions and action outcomes. A preliminary set of algorithms are developed and tested using the Global Terrorism Database (GTD). Our results reveals the benefits to perform joint predictive analytics with both attitude and influence. At the same time, we discover significant challenges in deriving influence and attitude from indirect observables for diverse adversarial behavior. These observations warrant further investigation of optimal use of influence and attitude for predictive analytics, as well as the potential inclusion of other environmental or capability elements for the actors.