We use barycentric coordinates to embed a graph and its given edge-core partition into a stratified collection of disks. We illustrate this embedding with graphs from social networks (a Friendster graph with 1.8 billion edges), paper citation networks (a Microsoft academic graph with 1.6 billion edges), and social media posts (Parler with 1.1 million edges, French election data with 185 thousand edges, and COVID-19 with 18 thousand edges). The proposed embedding aggregates vertices with similar participation in the input edge-core partition, and this similarity can be used to support subset queries useful for making sense of billion-edge graph data.
Recently, Graph Cities [1, 2] have been proposed as scalable 3d visual representations of graph edge partitions where each subgraph in the partition is a “fixed point of degree peeling”. In this work, we propose “intuitive” primitives to extract language semantics from the topology of these fixed points aided by provided graph vertex labels. The main approach is to view the collection of data labels as a set system derived from the graph topology and to derive “intuitive” language semantics from a specially derived set system intersection meta-graph. Exploration primitives include a glyph grid map of the distribution of all fixed points in the data set and a textual summary tool. We illustrate our approach with a variety of fixed points subgraphs extracted from “large” datasets that include a patent citation network (16.5 million edges) [3], a movie keywords co-occurrence network derived from the Internet Movie Database (5 million edges), a paper citation network derived from arXiv Computer Science papers (1.5 million edges), and a Parler dataset [4].
The ATU tale type index and the Motif Index of Folk-Literature have formed the basis for many comparative folktale studies. While the indices have been used extensively for the study of small groups of folktales and their associated motifs, there have been few attempts of describing a large linguistically and culturally unified corpus through its indexing. The study corpus consists of 2,606 folktales collected by Evald Tang Kristensen in nineteenth century Denmark, which were later indexed according to the second revised edition of the Aarne-Thompson index. We adjust this older index to align with the current ATU index. By creating linked network representations of the ATU index and the MI, as well as updating the Brandt indexing of the Danish folktales, we generate a network with 19,738 nodes and 28,292 edges, where nodes can be ATU numbers, MI numbers, Danish folktales, storytellers, or places of collection. By embedding all the Danish stories in this network, we provide a large-scale overview of the Danish folktale tradition. We introduce two novel interrelated network decomposition methods for the study of folktale collections at corpus scale: fixed points of degree peeling and graph fragments. The resulting analysis of the Danish corpus supports comparison with other traditions. Any collection that is similarly indexed can be embedded in this ATU+MI network and then subjected to the same interrelated graph decompositions.
We present efficient algorithmic mechanisms to partition graphs with up to 1.8 billion edges into subgraphs which are fixed points of degree peeling. We are able to process graphs with tens of millions of edges in seconds and hundreds of millions of edges in minutes. For fixed points that turn out to be larger than a desired interactivity parameter we further decompose them with a novel (to our knowledge) linear algorithm into what we call “graph waves and fragments”. This decomposition is used to create spanning views of fixed points that we call “DAG Covers”. We illustrate these decompositions by presenting intuitive and interactive visualizations of the meta-structures of a variety of publicly available data sets including social, web, and citation networks.
Graph Cities are the 3-D visual representations of partitions of a graph edge set into maximal connected subgraphs, each of which is called a fixed point of degree peeling. Each such connected subgraph is visually represented as a Building. A polylog bucketization of the size distribution of the subgraphs represented by the buildings generates a 2-D position for each bucket. The Delaunay triangulation of the bucket building locations determines the street network. We illustrate Graph Cities for the Friendster social network (1.8 billion edges), a co-occurrence keywords network derived from the Internet Movie Database (115 million edges), and a patent citation network (16.5 million edges). Up to 2 billion edges, all the elements of their corresponding Graph Cities are built in a few minutes (excluding I/O time). Our ultimate goal is to provide tools to build humanly interpretable descriptions of any graph, without being constrained by the graph size.
Graphs are everywhere, growing increasingly complex, and still lack scalable, interactive tools to support sensemaking. To address this problem, we present Atlas, an interactive graph exploration system that adapts scalable edge decomposition to enable a new paradigm for large graph exploration, generating explorable multilayered representations. Atlas simultaneously reveals peculiar subgraph structures, (e.g., quasi-cliques) and possible vertex roles in connecting such subgraph patterns. Atlas decomposes million-edge graphs in seconds, scaling to graphs with up to 117 million edges. We present the results from a think-aloud user study with three graph experts and highlight discoveries made possible by Atlas when applied to graphs from multiple domains, including suspicious yelp reviews, insider trading, and word embeddings. Atlas runs in-browser and is open-sourced.
We are developing an interactive graph exploration system called Graph Playground for making sense of large graphs. Graph Playground offers a fast and scalable edge decomposition algorithm, based on iterative vertex-edge peeling, to decompose million-edge graphs in seconds. Graph Playground introduces a novel graph exploration approach and a 3D representation framework that simultaneously reveals (1) peculiar subgraph structure discovered through the decomposition's layers, (e.g., quasi-cliques), and (2) possible vertex roles in linking such subgraph patterns across layers.
Project dates: June – July 2018 Graphs are an excellent form for the visualization of data because they clearly shows how individual data points are connected. However, for large sets of data, a direct visualization of a graph is indecipherable to the human eye. Separating sets and k-connected components are interesting structures in graphs that highlight critical data points and clusters of highly connected points respectively. In this paper, we develope an algorithm that uses minimum separating sets to decompose a graph into a hierarchy of k-connected components. The complexity of the algorithm depends linearly on the number of kconnected components in the graph. For each k-connected component K = (V,E), however, the complexity of finding its minimum separating set is O(nm ( n 2 ) ) where n = |V | and m = |E|. By using different approximate procedures this complexity can be improved to O(nm) or at more cost to accuracy to O(n + m). Performing the separating set decomposition creates a tree-structured map of the decomposed graph that more easily shows the connectivity of the graph and consequently the data it represents. I. PROJECT DESCRIPTION With the ever expanding availability of data, good visualization tools are more important than ever to aid people in analyzing data quickly. A popular way to visualize data is through the use of graphs as they are excellent sturctures for showing connections among a set of data. Fortunately, the connectivity of graphs is a heavily studied topic and many algorithms exist to analyze and manipulate the structures of graphs. The graph structures that we will be focusing in this paper are minimum separating sets and k-connected components. There structures are insteresting to visualize as they often reveal critical data points or groupings of data points in the data that the graph represents. For example, minimum separating sets could show critical failure points in a computer network, bottlenecks in distribution system, or even influential individuals in a social network. On the other hand, k-connected components offer a way to measure the redundancy in a computer network and distribution system and the closeness of groups in a social network. For the purpose of this paper we will specifically focus on these structures within Fixed Points of Degree Peeling. The peeling process and an algorithm to perform a decomposition of a graph into Fixed Points of Degree Peeling is described in detail in [1]. In this paper we first define basic graph-theortic concepts as well as k-connected components and minimum separating sets. Then, we introduce an algorithm for a graph decomposition into a hierarchy of k-connected components. Next, we propose some relaxations that allow for a faster algorithm implementation which still produces a decomposition of similar macro-structure to the ideal version. Finally we show the decomposition applied to a real data set and offer some interpretations for what it shows. II. BASIC DEFINITIONS An undirected graph G = (V,E) is a tuple of a vertex set V and an edge set E, which consists of unordered pairs of elements in V . A path in G is a sequence of vertices such that each consecutive pair of vertices is in E. A graph is said to be connected if there exists a path between any two vertices in V . A directed graph is very similar to a graph except that each element of E is an ordered pair of elements in V . Each edge is often described as (and represented visualy as) pointing from its source vertex to its target vertex. The degree of a vertex v ⊂ V denoted by deg(v) is the number of edges containing v. The diameter of a graph is the longest path of shortest distance between any two vertices. A tree is an undirected graph in which any two vertices are connected by exactly one path. A rooted tree is simply a tree with one vertex chosen as the root. The choice of root is only significant in context. One such instance is for a directed rooted tree where all the edges point either towards the root or away from the root. A leaf of a tree is any vertex v ⊂ V such that deg(v) = 1. A node of a tree is any vertex that is not a leaf. A separating set V ′ ⊂ V is a distinct set of vertices that, when removed from G, induces a subgraph that is not connected. A minimum separating set is any such V ′ where |V ′| ≤ |Si|,∀Si ∈ S the set of separating sets of G. A graph G is said to be k-connected if |V | > k and there doesn’t exist a minimum separating set V ′ with |V ′| < k. III. SEPARATING SET DECOMPOSITION Fig. 1. Separating Set Decomposition of Peel Layer 11 of Rutgers MS students’ study plans. A. Decomposition Algorithm The Separating Set Decomposition produces a directed rooted tree with the edges pointing towards the root. Each node in the Separating Set Decomposition represents a kconnected component of the graph, k being the label of the vertex (see Fig. 1). The children of each node are the connected components resulting from spliting the graph along the minimum separating set of the given k-connected component (see Ftn. split). The leaves of the tree either represent kconnected components that don’t have a minimum separating set (complete graphs) or trivial components (trees).
We use an iterative edge decomposition approach, derived from the popular iterative vertex peeling strategy, to globally split each vertex egonet (subgraph induced by a vertex and its neighbors) into a collection of edge-disjoint layers. Each layer is an edge maximal induced subgraph of minimum degree k that determines the layer density. This edge decomposition is derived completely from the overall network topology, and since each vertex can appear in multiple layers, we can associate to each vertex a vector profile that can be used to identify its different “roles” across the network. This allows us to explore a network’s topology at different levels of granularity, e.g., per layer and across layers. This is only feasible by mapping simultaneously a vertex to a set of 3D coordinates (x, y, and z) where the third coordinate encodes the different layers a vertex belongs to. This is one of the few instances where 3D visualization enhances graph exploration and navigation in an arguably “natural” way: a graph now becomes a 3D graph playground where †abello@dimacs.rutgers.edu *Authors contributed equally ‡fredhohman@gatech.edu §polo@gatech.edu a vertex plays a certain “role” per layer that is determined by the overall network topology. Our approach helps disentangle “hairball” looking embeddings produced by conventional 2D graph drawings.
12.1 Abstract . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1 12.2 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 12.3 Related Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 12.4 General Definitions Data Model and Statistics . . . . . . . . . . . . . . . . 5 12.5 Recency and Top K Filters . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 12.5.1 Recency Filtering. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 12.5.2 Top-K Tape Filtering. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 12.6 Graph Stream Visualization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 12.6.1 Recency Graph Visualization . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9 12.6.2 Top-K Tape Visualization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10 12.7 Life Cycle in a Graph Stream . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10 12.8 Top K Edge Group Patterns . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 12.8.1 Trending and Untrending . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 12.8.2 Herding and Straying . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 12.8.3 Pattern Identification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 12.8.4 Verification and Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 12.8.5 Holes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 12.9 Twitter Data Sample Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 12.9.1 Sample Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 12.10 Degree-of-Interest-based Visual Exploration . . . . . . . . . . . . . . . . . . . . . 18 12.11 Conclusions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22 12.12 Acknowledgments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22
Dimacs Discrete Methods in Epidemiology , Dimacs Discrete Methods in Epidemiology , کتابخانه مرکزی دانشگاه علوم پزشکی تهران
In many real-world networks, interactions between entities are observed at specific moments in continuous time, such as email, SMS messaging, and IP traffic. The majority of methods for analyzing such data first aggregate communication over designated time blocks, resulting in one or more discrete time series, to which existing tools can be applied. However, regardless of how the block lengths are chosen, discretizing time inherently introduces information loss and biases analysis towards patterns occurring at the designated time scale, effects which can be especially pronounced in networks with a high degree of temporal variability. Due to this, there has been increasing interest in using stochastic point processes to model network activity. We present a novel approach based on such models to detect times and sets of entities with temporally correlated recent activity. We develop efficient algorithms and compare our approach to existing and baseline methods through experiments on synthetic and real-world data.
Making sense of large graph datasets is a fundamental and challenging process that advances science, education and technology. We survey research on graph exploration and visualization approaches aimed at addressing this challenge. Different from existing surveys, our investigation highlights approaches that have strong potential in handling large graphs, algorithmically, visually, or interactively; we also explicitly connect relevant works from multiple research fields — data mining, machine learning, human-computer ineraction, information visualization, information retrieval, and recommender systems — to underline their parallel and complementary contributions to graph sensemaking. We ground our discussion in sensemaking research; we propose a new graph sensemaking hierarchy that categorizes tools and techniques based on how they operate on the graph data (e.g., local vs global). We summarize and compare their strengths and weaknesses, and highlight open challenges. We conclude with future research directions for graph sensemaking.
Visualization is a powerful paradigm for exploratory data analysis. Visualizing large graphs, however, often results in a meaningless hairball. In this paper, we propose a different approach that helps the user adaptively explore large million-node graphs from a local perspective. For nodes that the user investigates, we propose to only show the neighbors with the most subjectively interesting neighborhoods. We contribute novel ideas to measure this interestingness in terms of how surprising a neighborhood is given the background distribution, as well as how well it fits the nodes the user chose to explore. We are currently designing and developing AdaptiveNav, a fast and scalable method for visually exploring large graphs. By implementing our above ideas, it allows users to look into the forest through its trees.
The Federal Railroad Administration grade crossing accident database contains numerous interrelated variables. Understanding of how the variables are interrelated can be enhanced using modern visualization techniques. These techniques can allow managers from railroads and government agencies to find complex variables relationships not usually provided by routine statistical analyses. For this research we have developed several dashboards of linked visualizations using the Weave data visualization software [5]. Our visualizations explore various accident types of concern to the railroad industry including trespassing and pedestrian accidents, passenger train accidents, actions of highway users involved in accidents, and the effect of different types of warning devices on grade crossing accidents. In addition, we are currently developing an advanced visualization system that views the accident data as time varying events occurring over a fixed grade crossings topology. This view allows the application of a recent network data abstraction termed "Graph Cards." We present initial examples of the advanced system that provides a variety of filtering mechanisms to view statistical distributions and their time varying behavior over the grade crossings topology.
Visualization is a powerful paradigm for exploratory data analysis. Visualizing large graphs, however, often results in a meaningless hairball. In this paper, we propose a different approach that helps the user adaptively explore large million-node graphs from a local perspective. For nodes that the user investigates, we propose to only show the neighbors with the most subjectively interesting neighborhoods. We contribute novel ideas to measure this interestingness in terms of how surprising a neighborhood is given the background distribution, as well as how well it fits the nodes the user chose to explore. We introduce FACETS, a fast and scalable method for visually exploring large graphs. By implementing our above ideas, it allows users to look into the forest through its trees. Empirical evaluation shows that our method works very well in practice, providing rankings of nodes that match interests of users. Moreover, as it scales linearly, FACETS is suited for the exploration of very large graphs.
Visualization is a powerful paradigm for exploratory data analysis. Visualizing large graphs, however, often results in a meaningless hairball. In this paper, we propose a different approach that helps the user adaptively explore large million-node graphs from a local perspective. For nodes that the user investigates, we propose to only show the neighbors with the most subjectively interesting neighborhoods. We contribute novel ideas to measure this interestingness in terms of how surprising a neighborhood is given the background distribution, as well as how well it fits the nodes the user chose to explore. We are currently designing and developing AdaptiveNav, a fast and scalable method for visually exploring large graphs. By implementing our above ideas, it allows users to look into the forest through its trees.
Networks that evolve over time, or dynamic graphs, have been of interest to the areas of information visualization and graph drawing for many years. Typically, the structure of the dynamic graph evolves as vertices and edges are added or removed from the graph. In a multivariate scenario, however, attributes play an important role and can also evolve over time. In this chapter, we characterize and survey methods for visualizing temporal multivariate networks. We also explore future applications and directions for this emerging area in the fields of information visualization and graph drawing.
Thomas Shermer合作论文数Graph Theory and Computer Graphics;Computational Geometry2