Socio-technical systems usually consist of many intertwined networks, each connecting different types of objects or actors through a variety of means. As these networks are co-dependent, one can take advantage of this entangled structure to study interaction patterns in a particular network from the information provided by other related networks. A method is, hence, proposed and tested to recover the weights of missing or unobserved links in heterogeneous information networks (HIN)—abstract representations of systems composed of multiple types of entities and their relations. Given a pair of nodes in a HIN, this work aims at recovering the exact weight of the incident link to these two nodes, knowing some other links present in the HIN. To do so, probability distributions resulting from path-constrained random walks, i.e., random walks where the walker is forced to follow only a specific sequence of node types and edge types, capable to capture specific semantics and commonly called a meta-path, are combined in a linearly fashion to approximate the desired result. This method is general enough to compute the link weight between any types of nodes. Experiments on Twitter and bibliographic data show the applicability of the method.
Diversity is a concept relevant to numerous domains of research varying from ecology, to information theory, and to economics, to cite a few. It is a notion that is steadily gaining attention in the information retrieval, network analysis, and artificial neural networks communities. While the use of diversity measures in network-structured data counts a growing number of applications, no clear and comprehensive description is available for the different ways in which diversities can be measured. In this article, we develop a formal framework for the application of a large family of diversity measures to heterogeneous information networks (HINs), a flexible, widely-used network data formalism. This extends the application of diversity measures, from systems of classifications and apportionments, to more complex relations that can be better modeled by networks. In doing so, we not only provide an effective organization of multiple practices from different domains, but also unearth new observables in systems modeled by heterogeneous information networks. We illustrate the pertinence of our approach by developing different applications related to various domains concerned by both diversity and networks. In particular, we illustrate the usefulness of these new proposed observables in the domains of recommender systems and social media studies, among other fields.
Graph compression is a data analysis technique that consists in the replacement of parts of a graph by more general structural patterns in order to reduce its description length. It notably provides interesting exploration tools for the study of real, large-scale, and complex graphs which cannot be grasped at first glance. This article proposes a framework for the compression of temporal graphs, that is for the compression of graphs that evolve with time. This framework first builds on a simple and limited scheme, exploiting structural equivalence for the lossless compression of static graphs, then generalises it to the lossy compression of link streams, a recent formalism for the study of temporal graphs. Such generalisation relies on the natural extension of (bidimensional) relational data by the addition of a third temporal dimension. Moreover, we introduce an information-theoretic measure to quantify and to control the information that is lost during compression, as well as an algebraic characterisation of the space of possible compression patterns to enhance the expressiveness of the initial compression scheme. These contributions lead to the definition of a combinatorial optimisation problem, that is the Lossy Multistream Compression Problem, for which we provide an exact algorithm.
In social network Twitter, users can interact with each other and spread information via retweets. These millions of interactions may result in media events whose influence goes beyond Twitter framework. In this paper, we thoroughly explore interactions to provide a better understanding of the emergence of certain trends. First, we consider an interaction on Twitter to be a triplet $(s,a,t)$ meaning that user $s$, called the spreader, has retweeted a tweet of user $a$, called the author, at time $t$. We model this set of interactions as a data cube with three dimensions: spreaders, authors and time. Then, we provide a method which builds different contexts, where a context is a set of features characterizing the circumstances of an event. Finally, these contexts allow us to find relevant unexpected behaviors, according to several dimensions and various perspectives: a user during a given hour which is abnormal compared to its usual behavior, a relationship between two users which is abnormal compared to all other relationships, \textit{etc.} We apply our method to a set of retweets related to the 2017 French presidential election and show that one can build interesting insights regarding political organization on Twitter.
This paper aims at precisely detecting and identifying anomalous events in IP traffic. To this end, we adopt the link stream formalism which properly captures temporal and structural features of the data. Within this framework, we focus on finding anomalous behaviours with respect to the degree of IP addresses over time, i.e. the number of distinct IP addresses with which they interact over time. Due to diversity in IP profiles, this feature is typically distributed heterogeneously, preventing us to directly find anomalies. To deal with this challenge, we design a method to detect outliers as well as precisely identify their cause in a sequence of similar heterogeneous distributions. We apply it to several IP traffic captures and we show that it succeeds in detecting relevant patterns in terms of anomalous network activity.
Heterogeneous information networks (HINs) are abstract representations of systems composed of multiple types of entities and their relations. Given a pair of nodes in a HIN, this work aims at recovering the exact weight of the incident link to these two nodes, knowing some other links present in the HINs. Actually, this weight is approximated by a linear combination of probabilities, results of path-constrained random walks, i.e., random walks where the walker is forced to follow only a specific sequence of node types and edge types which is commonly called a meta path, performed on the HINs. This method is general enough to compute the link weight between any types of nodes. Experiments on Twitter data show the applicability of the method.
We introduce a method which aims at getting a better understanding of how millions of interactions may result in global events. Given a set of dimensions and a context, we find different types of outliers: a user during a given hour which is abnormal compared to its usual behavior, a relationship between two users which is abnormal compared to all other relationships, etc. We apply our method on a set of retweets related to the 2017 French presidential election and show that one can build interesting insights regarding political organization on Twitter.
Precise detection and identification of anomalous events in IP traffic are crucial in many applications. This paper intends to address this task by adopting the link stream formalism which properly captures temporal and structural features of the data. Within this framework we focus on finding anomalous behaviours with the degree of IP addresses over time. Due to diversity in IP profiles, this feature is typically distributed heterogeneously, preventing us to find anomalies. To deal with this challenge, we design a method to detect outliers as well as precisely identify their cause in a sequence of similar heterogeneous distributions. We apply it to a MAWI capture of IP traffic and we show that it succeeds at detecting relevant patterns in terms of anomalous network activity.
Social research on public opinion has been affected by the recent deluge of new digital data on the Web, from blogs and forums to Facebook pages and Twitter accounts. This fresh type of information useful for mining opinions is emerging as an alternative to traditional techniques, such as opinion polls. Firstly, by building the state of the art of studies of political opinion based on Twitter data, this paper aims at identifying the relationship between the chosen data analysis method and the definition of political opinion implied in these studies. Secondly, it aims at investigating the feasibility of performing multiscale analysis in digital social research on political opinion by addressing the merits of several methodological techniques, from content-based to interaction-based methods, from statistical to semantic analysis, from supervised to unsupervised approaches. The end result of such an approach is to identify future trends in social science research on political opinion.
Twitter is now an integral part of means of communication used by political leaders to disseminate information to the public. A politician may use it sporadically to merely broadcast to his followers or on the contrary employ it regularly and tweet at strategic moments. Likewise, their followers may be occasional spreaders or real online activists retweeting primarily a particular political figure. The complex processes formed by interactions between users through a retweet may result in a media event whose influence goes beyond Twitter framework. We propose a multidimensional and multilevel analysis method to describe structural and temporal relational patterns in this retweet network as well as to find unexpected behaviors related to this political strategy of communication. The retweet temporal network we use has been obtained from observing the Twitter accounts of nearly 3, 500 political actors such as individuals, organizations and institutions. In this network, two users are connected at time t if one of the two – the spreader – has retweeted a tweet of the other – the author – at time t. Our method consists firstly in evaluating the quantity of interactions between an author and a spreader during a given time period. After this step, we have access to local information: the finest scale at which we can observe interactions. One then uses data aggregation to obtain the total quantity of interactions of an entity such as an author, a spreader, an hour or a couple obtained by combining those three dimensions (marginal values). This step gives us access to global information that can be used to provide more context to local data. Afterwards, in order to detect irregularities in users and temporal behaviors we compare the previously obtained quantities of interactions between them. Here again, we propose to decompose comparison in multiple levels. For instance, one can look at the quantities of interactions by hours, comparing them all at global scale and find an hour with unexpected activity level compared to all the others. Then, one can look at the couples (author, hour), comparing them at local scale after having fixed the hour previously found abnormal, i.e.
Notre image du monde depend dans une large mesure des flux d'information que nous recevons de l'etranger via les medias de masse. Nous proposons dans cet article un cadre d'analyse quantitative des flux mediatiques – marqueurs possibles des dynamiques contemporaines de mondialisation et de regionalisation – reposant sur le concept d'« agenda geomediatique », c'est-a-dire le processus de selection des unites territoriales qui sont portees a l'attention du public par les medias. Nous presentons pour cela trois modeles permettant d'identifier les ressemblances et les specificites geographiques et temporelles de differents medias, et ainsi d'analyser la formation de l'actualite internationale selon trois perspectives distinctes.
Les evenements mediatiques constituent un objet de recherche empirique, methodologique et theorique d’un grand interet pour la creation d’une science des territoires. Cette communication propose trois variations de complexite croissante autour de l’application possible des notions de « territoire », « territorialite » et « territorialisation » a la description des evenements mediatiques. Chacune de ces variations est illustree par des resultats de recherche recents du projet ANR Geomedia, sur la base d’un corpus de flux RSS internationaux de journaux de langues francaise, anglaise et espagnole localises dans differents pays du monde.
Notre image du monde dépend dans une large mesure des flux d’information que nous recevons de l’étranger via les médias de masse. Nous proposons dans cet article un cadre d’analyse quantitative des flux médiatiques – marqueurs possibles des dynamiques contemporaines de mondialisation et de régionalisation – reposant sur le concept d’« agenda géomédiatique », c’est-à-dire le processus de sélection des unités territoriales qui sont portées à l’attention du public par les médias. Nous présentons pour cela trois modèles permettant d’identifier les ressemblances et les spécificités géographiques et temporelles de différents médias, et ainsi d’analyser la formation de l’actualité internationale selon trois perspectives distinctes.
Because the dynamics of complex systems is the result of both decisive local events and reinforced global effects, the prediction of such systems could not do without a genuine multilevel approach. This paper proposes to found such an approach on information theory. Starting from a complete microscopic description of the system dynamics, we are looking for observables of the current state that allows to efficiently predict future observables. Using the framework of the information bottleneck (IB) method, we relate optimality to two aspects: the complexity and the predictive capacity of the retained measurement. Then, with a focus on agent-based models (ABMs), we analyze the solution space of the resulting optimization problem in a generic fashion. We show that, when dealing with a class of feasible measurements that are consistent with the agent structure, this solution space has interesting algebraic properties that can be exploited to efficiently solve the problem. We then present results of this general framework for the voter model (VM) with several topologies and show that, especially when predicting the state of some sub-part of the system, multilevel measurements turn out to be the optimal predictors.
Science studies have shown that a large number of exogenous factors influence the conduct of scientific activities, such as the social, institutional, economical, political, and geographical context of academic research. Hence, the description and explanation of scientific practices cannot go without a clear understanding of the impact of these macroscopic contexts on the individuals. In this communication, we propose a multilevel analysis method to describe coauthorship networks at various scales and thus address these challenging issues. Our method exploits information-theoretic data aggregation to summarise the information that is contained at the individual level in a given data set and to highlight consistent patterns that arises at higher-levels [1, 2, 3]. Applied to relational data, this method consists in partitioning the adjacency matrix of co-authorship networks into “rectangular tiles”. Each such tile represents an aggregate edge between two aggregate nodes, that is the aggregation of several collaboration relations between two groups of researchers. Similarly to blockmodeling methods – introduced for the analysis of social networks – the quality of such an aggregate depends on the homogeneity of the relations it contains, that is the homogeneity of collaborations between the two groups of researchers. We propose to use information-based measures, such as Kullback-Leibler divergence, to quantify and optimise this homogeneity criterion. Data aggregation thus provides an abstraction method to identify homogeneous macroscopic patterns within the network. However, one can require additional constraints to apply on the set of feasible abstractions depending on the exogenous factors one is interested in. For example, one could only consider as valid abstractions the groups of researchers that are defined at some institutional level: e.g., individuals aggregated as research teams, as research departments, as institutes, and so on. By introducing such constraints within the aggregation process, the network is then macroscopically described on the
Notre image du monde depend dans une large mesure des flux d’information que nous recevons de l’etranger via les medias de masse. Nous proposons dans cet article un cadre d’analyse quantitative des flux mediatiques – marqueurs possibles des dynamiques contemporaines de mondialisation et de regionalisation – reposant sur le concept d’« agenda geomediatique », c’est-a-dire le processus de selection des unites territoriales qui sont portees a l’attention du public par les medias. Nous presentons pour cela trois modeles permettant d’identifier les ressemblances et les specificites geographiques et temporelles de differents medias, et ainsi d’analyser la formation de l’actualite internationale selon trois perspectives distinctes.
The design and the debugging of large-scale MAS require abstraction tools in order to work at a macroscopic level of description. Agent aggregation provides such abstractions by reducing the complexity of the system’s microscopic representation. Since it leads to an information loss, such a key process may be extremely harmful for the analysis if poorly executed. This paper presents measures inherited from information theory to evaluate abstractions and to provide the experts with feedback regarding the quality of generated representations. Several evaluation techniques are applied to the spatial and temporal aggregation of an agent-based model of international relations. The information from on-line newspapers constitutes a complex microscopic representation of the agent states. Our approach is able to evaluate geographical abstractions used by the domain experts in order to provide efficient and meaningful macroscopic representations of the world global state.
Guillaume Huard合作论文数Laboratoire d'Informatique de Grenoble2