We propose a new method of constructing a variable bin width histogram that can accommodate the unbalanced distribution of the samples yet retaining, as a whole, the good aspect of both equal width (EW) and equal-area (EA) histograms that are being used popularly for data visualization and analysis. We formulate this as an optimal change point detection problem in which the bin boundaries are determined by minimizing the sum of the absolute error or the squared error in each bin. The former is based on Distance Minimization (DM) and new, and the latter is based on Variance Minimization (VM) and is considered the state-of-the-art. The constructed histograms can effectively be used to detect and visualize hidden outliers/anomalies by applying the interquartile range method in each bin. The final histograms are obtained by adjusting bin boundaries and heights accordingly after removing the detected outliers/anomalies. We further propose a method to annotate the constructed bins if the data for annotation is given for each sample as a set of nominal variables, using z -score with respect to their distribution within each bin. We applied our method to both real vinyl greenhouse datasets and two different sets of three synthetic datasets, and confirmed that both DM and VM methods work as intended, both can represent the sample distribution with a smaller number of bins than those by EW and EA methods, The use of interquartile range method can detect anomalies as well as outliers, and the terms selected for annotation are interpretable and reasonable. EW and EA methods have contrasting properties. DM and VM methods lie in between, but the former is closer to EA method and the latter to EW method. DM method runs substantially faster than VM method and performs slightly better than VM method in outlier detection and annotation tasks.
Finding building blocks buried in real-world networks is an important task not only in network science but also in biology, chemistry, sociology, and other fields. Many attempts have been conducted to efficiently search for these building blocks under different settings. In this study, we take up the challenge of counting motifs from uncertain graphs, i.e., computing the expected frequency of motifs. In general, analysis for uncertain graphs is computationally expensive even with a small graph because there exists a vast number of possible worlds and a large number of sample graphs are needed to accurately cover the possible worlds. To alleviate such inefficiency coming from sampling which many existing studies rely on, we propose an analytical computation method that gives an exact expected frequency and is not based on costly sampling. The key idea of our method is to marginalize the probability of each possible world on a candidate motif, which can drastically reduce the number of possible worlds. We introduce matrices on the number of transition patterns and the transition probabilities among motifs to achieve further acceleration if the edge-existence probability is uniform and constant. We conduct experimental evaluations in the task of computing the frequency significance of directed 3- and 4-node and undirected 4-node motifs under both the uniform and non-uniform probability settings. The results confirm that the proposed method is effective and efficient. The sampling-based state-of-the-art method needs 2–4 orders of magnitude more time than the proposed method to achieve the same accuracy. In addition, the accelerated version of our method can achieve further acceleration, i.e., it runs about 1 order of magnitude faster than the general version of ours under the uniform probability setting.
Counting motifs in an uncertain graph for which each link is associated with a connection probability is computationally expensive when the graph is huge due to the extremely large number of possible worlds. Natural approach is to rely on sampling-based approximation methods, but this still needs many sample graphs for obtaining accurate results. We propose a novel method that analytically computes the expected frequency of motif without relying on expensive sampling. Marginalizing the probability of each possible world on a candidate motif can drastically reduce the number of possible worlds that need be considered when the size of motif is small. Experiments using real-world data confirm that the proposed method is effective and efficient. It is far better than the state-of-the-art sampling-based method. The accuracy is guaranteed and the running time is about 4 order of magnitude faster. It runs at a speed that does not depend on the connection probability.
We challenge the problem of efficiently identifying critical links that substantially degrade network performance if they do not function under a realistic situation where each link is probabilistically disconnected, e.g., unexpected traffic accident in a road network and unexpected server down in a communication network. To solve this problem, we utilize the bridge detection technique in graph theory and efficiently identify critical links in case the node reachability is taken as the performance measure.To be more precise, we define a set of target nodes and a new measure associated with it, Target-oriented latent link Criticalness Centrality (TCC), which is defined as the marginal loss of the expected number of nodes in the network that can reach, or equivalently can be reached from, one of the target nodes, and compute TCC for each link by use of detected bridges. We apply the proposed method to two real-world networks, one from social network and the other from spatial network, and empirically show that the proposed method has a good scalability with respect to the network size and the links our method identified possess unique properties. They are substantially more critical than those obtained by the others, and no known measures can replace the TCC measure.
We propose a general framework of opening and closing shops in group competitive environment, i.e., shops in the same group work cooperatively and those in different groups competitively, based on stochastic utility which is given by a function of shop distance and attractiveness with an explicit traveling time-bound imposed. The framework allows to derive a specific prediction model by choosing a specific function form of utility, including the one we propose, its variants and the conventional state-of-the-art gravity model which we chose as a reference. We compute a marginal gain of the market share which is derived from the utility function and the consumers buying power as a measure to rank the candidate location. Using the real dataset of three convenience store groups in four cities in Japan, we analyzed how the derived models behave with respect to the time-bound and the other parameters and how each model compare with others. We confirm that, despite the simplification we made in the model, inferred rankings of the shops newly opened in the real data are shown to be high implying that our prediction model and other variants are reasonable. We show that our model gives much more realistic results than the gravity model, which indicates that our group competitive mechanism with the time-bounded stochastic utility is vital and promising. Inclusion of the time-bound constraint is crucially important. Analyses of the dynamics of opening and closing shops indicate that competition indeed affects the market share of each group over time, and the total share eventually increases although small, and the difference of the share within each group gradually becomes smaller, revealing that the spatial distribution of the shops in each group becomes more uniform.
We focus on a class of link injection problem of spatial network, i.e., finding best places to construct k new roads that save as many people as possible in a time-critical emergency situation. We quantify the network performance by node coverage under the presence of time constraint and propose an efficient algorithm that maximizes the marginal gain by use of lazy evaluation making the best of time constraint. We apply our algorithm to three problem scenarios (disaster evacuation, ambulance call, fire engine dispatch) using real-world road network and geographical information of actual facilities and demonstrate that 1) use of lazy evaluation can achieve nearly two orders of magnitude reduction of computation time compared with the straightforward approach and 2) the location of new roads is intuitively explainable and reasonable.
The ADMAS 2020 conference details original research, case study results, and experienced-based insights into advanced data mining and applications.
We address a problem of efficiently estimating the influence of a node in information diffusion over a social network. Since the information diffusion is a stochastic process, the influence degree of a node is quantified by the expectation, which is usually obtained by very time-consuming many runs of simulation. Our contribution is that we proposed a framework for predictive simulation based on the leave-N-out cross-validation technique that well approximates the error from the unknown ground truth for two target problems: one to estimate the influence degree of each node, and the other to identify top-K influential nodes. The method we proposed for the first problem estimates the approximation error of the influence degree of each node, and the method for the second problem estimates the precision of the derived top-K nodes, both without knowing the true influence degree. We experimentally evaluate the proposed methods using the three real-world networks and show that they can serve as a good measure to solve the target problems with far fewer runs of simulation ensuring the accuracy when the leave-half-out cross-validation, i.e., N is the half of the number of runs, is used, which means that one can identify the influential nodes without knowing exactly their influence degree in good accuracy.
We address the problem of opening and closing shops in group competitive environment, i.e., shops in the same group work cooperatively and those in different groups competitively, and analyze how the market share and location changes over time. We formulate a stochastic utility of each shop as a function of shop distance and attractiveness from which a market share is computed by weighting consumers buying power. We further place a constraint on the traveling time, which is crucial to reduce the computation time, and use a marginal gain of the market share as a measure to rank the candidate location. Using the real dataset of three convenience stores in four cities in Japan, we confirm that, despite the simplification we made in the model, rankings of the existing shops are shown to be high which implies that our model is reasonable. Further, comparison with the baseline gravity model shows that our model gives much more realistic results. Analyses of the dynamics of opening and closing shops indicate that the reasonable time-bound for walking is about 10 min., the market share of each group, thus total share, eventually increases although small, and the difference of the share within each group gradually becomes smaller, revealing that the spatial distribution of the shops in each group becomes more uniform.
Ranking nodes in uncertain graph is computationally expensive when the graph is huge due to the extremely large number of possible worlds. Some approximation is needed in general. We focus on PageRank centrality measure to rank and propose a method that does not use any approximation for uncertain graph in which all the links can be uncertain. We first compute the expected transition matrix over all the possible graphs accurately and then run PageRank algorithm only once to rank the nodes (p-avg approach). This is not the same as computing the scores for each individual graph first and then rank the nodes by taking their average (s-avg approach). Exact computation of the latter is not possible because of the heavy computational load and only the approximate scores are obtained by limiting the number of graphs by sampling. We have tested the performance from various angles using three real world networks. We show that the proposed method (p-avg approach) gives very high precision to the s-avg approach for highly ranked nodes and can be a good alternative to it. Pactically, the p-avg approach runs orders of magnitude, i.e., sample size, faster than the s-avg approach.
We address a problem of efficiently estimating value of a centrality measure for a node in a large network, and propose a sampling-based framework in which only a small number of nodes that are randomly selected are used to estimate the measure. The error estimator we derived is an unbiased estimator of the approximation error defined as the expectation of the difference between the true and the estimated values of the centrality. We experimentally evaluate the fundamental performance of the proposed framework using the closeness and betweenness centralities on six real world networks from different domains, and show that it allows us to estimate the approximation error more tightly and more precisely than the standard error estimator traditionally used based on i.i.d. sampling, i.e., with the confidence level of 95% for a small number of sampling, say 20% of the total number of nodes.
In this paper, we focus on an emergency situation in the real-world such as disaster evacuation and propose an algorithm that can efficiently identify critical links in a spatial network that substantially degrade network performance if they fail to function. For that purpose, we quantify the network performance by node reachability from/to one of target facilities within the prespecified time limitation, which corresponds to the number of people who can safely evacuate in a disaster. Using a real-world road network and geographical information of actual facilities, we demonstrated that the proposed method is much more efficient than the method based on the betweenness centrality that is one of the representative centrality measures and that the critical links detected by our method cannot be identified by using a straightforward extension of the betweenness centrality.
The problem of efficiently identifying critical nodes that substantially degrade network performance if they do not function is crucial and essential in analyzing a large complex network such as social networks on the Web, and it is still challenging. In this paper, we tackle this problem under a realistic situation where each link is probabilistically disconnected reflecting that an information path between two persons in a social network is not always open to pass on a message, rather than assuming that every information path is always open and passes on any information from one to the other. To solve this problem, we focus on the articulation point and utilize the bridge detection technique in graph theory to efficiently identify critical nodes in case the node reachability is taken as the performance measure. This corresponds to the total number of people who can receive information issued by every single person in a social network. Using two real-world social networks, we empirically show that the proposed method has a good scalability with respect to the network size and the nodes our method identified possesses unique properties and they are difficult to be identified by using conventional centrality measures.
The overarching context of work consists of activities, in which people are actors within a “choreographed” social interaction. Goals and problems arise within this conceptual, social context, in which technical, product-oriented tasks and their associated methods and evaluation criteria are defined. Brahms is a simulation tool for modeling this interactive social context, represented as the activities or “practice” of located agents. Rather than modeling cognition in detail, Brahms models focus on what people do where, when, and with whom. This entails modeling social knowledge—what people know about each other’s activities and capabilities, by which collaboration is possible. Rather than modeling problem solving as disembodied puzzle manipulation, Brahms models focus on circumstantial, interactional influences on how work actually gets done, especially how information is shared and how participation (and hence a problem solving method) is determined. Brahms is suitable for use in work systems design, instruction, implementing software agents, and as a workbench for studying social systems. INTRODUCTION TO BRAHMS Brahms is a multi-agent simulation framework for modeling work practice, incorporating state-of-the-art methods from artificial intelligence research and insights about work and learning from the social sciences. Brahms was developed for use in work systems design, instruction, and as a language for software agents: • Brahms models consist of groups of agents with context-sensitive, interactive behaviors. Agents are located, mobile, and have knowledge and changing beliefs. Groups may define job functions, teams, people at a certain location, or people with certain knowledge and beliefs. • Brahms enables modeling activities of people during the day—how people spend their time—emphasizing information processing, communication in different modalities (phone, fax, voice mail, face-to-face, databases), and location-specific interaction (meetings, chance conversations, teamwork). Thus, Brahms allows modeling a community of practice—a group of people who participate in some shared, choreographed interaction, usually involving collaboration between individuals with different roles and experience.
Social media allows people to post widely and evaluate diverse information including ideas, news and opinions. Once such an online item is posted on a social media site, it can be appreciated and shared by many people and become popular. This kind of phenomenon can have a large influence on people’s daily life and social trends. Thus, studies on modeling the arrival process of shares to an individual item have recently attracted a great deal of interest in the field of social media mining. In this paper, we propose, by combining a Dirichlet process with a Hawkes process in a novel way, a probabilistic model, called cooperative Hawkes process (CHP) model, to discover the cooperative structure among all the items involved. The proposed model takes into account all the arrival processes of shares for those items. We develop an efficient method of inferring the CHP model from the observed sequences of share-events, and present an effective framework for predicting the future popularity of each of these items. Using synthetic and real data, we demonstrate that the CHP model outperforms the Hawkes process model without interaction among items (HP model) and the multivariate Hawkes process model (MHP model) in terms of popularity prediction. Moreover, for real data from a cooking-recipe sharing site, we discover the cooperative structure among cooking-recipes in view of popularity dynamics by applying the CHP model.
At its heart the act of reviewing is very subjective, but in reality many factors would influence user's decision. This can be called social influence bias. We pick two factors, "Who" and "When" and discuss which factor is more influential when a user posts his/her own rate in an online review system. We consider two kinds of users: real and virtual. In the former each user has its own metric, but in the latter the metric is assigned to the order of review posting actions (rating). We propose a weighted multinomial generative model that can learn the factor metric quite efficiently from a vast amount of data already available in many online review systems. If the model can explain the data well enough, this implies that such a social bias does exist. We evaluate the proposed method and confirm its effectiveness by five review datasets, and empirically clarify that there is no universal solution, but the social bias does exist. In reality the influential factor depends on each dataset, the majority of users is normal (average), and there are two small groups of users, each with high metric value and low metric value.
We address the problem of efficiently detecting critical links in a large network. Critical links are such links that their deletion exerts substantial effects on the network performance such as the average node reachability. We tackle this problem by proposing a new method which consists of three acceleration techniques: redundant-link skipping (RLS), marginal-node pruning (MNP) and burn-out following (BOF). All of them are designed to avoid unnecessary computation and work both in combination and in isolation. We tested the effectiveness of the proposed method using two real-world large networks and two synthetic large networks. In particular, we showed that the proposed method can estimate the performance degradation by link removal without introducing any approximation within a computation time comparable to that needed by the bottom-k sketch which is a summary of dataset and can efficiently process approximate queries, i.e., reachable nodes, on the original dataset, i.e., the given network. Further, we confirmed that the measures easily composed by the well known existing centralities, e.g. in/out-degree, betweenness, PageRank, authority/hub, are not able to detect critical links. Links detected by these measures do not reduce the average reachability at all, i.e., not critical at all.
Efficiently identifying critical links that substantially degrade network performance if they fail to function is challenging for a large complex network. In this paper, we tackle this problem under a more realistic situation where each link is probabilistically disconnected as if a road is blocked in a natural disaster than assuming that any road is never blocked in a disaster. To solve this problem, we utilize the bridge detection technique in graph theory and efficiently identify critical links in case the node reachability is taken as the performance measure, which corresponds to the number of people who can reach at least one evacuation facility in a disaster. Using two real-world road networks, we empirically show that the proposed method is much more efficient than the other methods that are based on traditional centrality measures and the links our method detected are substantially more critical than those by the others.
We address the problem of efficiently detecting critical links in a large network in order to maintain network performance, e.g., in case of disaster evacuation, for which a probabilistic link disconnection model plays an essential role. Here, critical links are such links that their disconnection exerts substantial effects on the network performance such as the average node reachability. We tackle this problem by proposing a new method consisting of two new acceleration techniques: reachability condition skipping (RCS) and distance constraints skipping (DCS). We tested the effectiveness of the proposed method by using three real-world spatial networks. In particular, we show that the proposed method achieves the efficiency gain of around 10(4) compared with a naive method in which every single link is blindly tested as a critical link candidate.
Kouzou Ohara合作论文数The Institute of Scientific and Industrial Research76