What is Panther
Panther is a method that can estimate vertex structural similarity in a very large network very quickly. Panther calculate two kinds of vertex similarities:
- The first one is based on the principle that two vertices are considered structurally equivalent if they have many common neighbors in a network.
- The second one is based on the principle that two vertices are considered structurally equivalent if they play the same structural role-this can be further quantified by degree closeness centrality, betweennes and other network centrality metrics.
How does Panther work
The algorithm is baed on a novel idea of random path. Specifically, given a network, we perform R random walks, each starting from a randomly picked vertex and walking T steps. The idea behind this is that two vertices have a high similarity if they frequently appear on the same paths. To improve the efficiency, we build an inverted index of vertex-to-path. Using the index, we can retrieve all paths that contain a specific vertex v with a complexity of O(1). Figure 1 illustrates the process of random path sampling.

Figure1. Illustration of random path sampling
We provide theoretical proofs for the error-bound and confidence of the proposed algorithm. In general, the path similarity can be viewed as a probability of measure defined over all paths . Thus we can adopt the results from Vapnik-Chernovenkis (VC) learning theory to analyze the proposed sampling-based algorithm. Theoretically, we obtain that the sample size,
,
only depends on the path length T of each random walk, for a given error-bound
and a confidence level
.
To capture the information of structural patterns, we extend the proposed algorithm by augmenting each vertex with a vector of structure-based features. Specifically, for vertex vi in the network, we first calculate the similarity between vi and all the other vertices using Panther. Then we construct a feature vector for vi by taking the largest D similarity scores as feature values. The resultant algorithm is referred to as Panther++. Panther++ is not only able to estimate similarity between vertices in a connected network, but also capable of estimating similarity betwen vertices from disconnected networks. Figure 2 shows an example of top-k similarity search across two disconnected networks, where v4, v6 and v5 are top-3 similar vertices to v0.

Figure2. Top-k similarity search
Details of the algorithm are presented in Algorithm 1,2 & 3.


Efficiency Performance
Table 2 lists statistics of the different Tencent sub-networks (355,591,065 users and 5,958,853,072 “following” relationships in total. ) and the efficiency performance of the comparison methods. Clearly, our methods (both Panther and Panther++) are much faster than the comparison methods. For example, on the Tencent6 sub-network, which consists of 443,070 vertices and 5,000,000 edges, Panther achieves a 390 speed-up , compared to the fastest (ReFeX) of all the comparative methods. we can also see that RWR, TopSim and RoleSim cannot complete top-k similarity search for all vertices within a reasonable time when the number of edges increases to 500,000. ReFeX can deal with larger networks, but also fails when the edge number increases to 10,000,000. Our methods can scale up to handle very large networks with more than 10,000,000 edges. On average, Panther only needs 0.0001 second to perform top-k similarity search for each vertex in a large network.
Baselines:
Random walk with restart (RWR):J.-Y. Pan, H.-J. Yang, C. Faloutsos, and P. Duygulu. Automatic multimedia cross-modal correlation discovery. In SIGKDD’04, pages 653–658, 2004.
TopSim:P. Lee, L. V. Lakshmanan, and J. X. Yu. On top-k structural similarity search. In ICDE’12, pages 774–785, 2012.
RoleSim: R. Jin, V. E. Lee, and H. Hong. Axiomatic ranking of network role similarity. In KDD’11, pages 922–930, 2011.
ReFeX:K. Henderson, B. Gallagher, L. Li, L. Akoglu, T. Eliassi-Rad, H. Tong, and C. Faloutsos. It’s who you know: graph mining using recursive structural features. In KDD’11, pages 663–671, 2011.

Qualitative Case study
“Who is similar to Barabási?” Albert-László Barabási is a famous Hungarian-American physicist, who proposed the Barabási–Albert (BA) model for generating random scale-free networks using a preferential attachment mechanism. We apply Panther++ to a scientific network to find researchers who have similar structural positions to that of Dr. Barabási. Figure 10 shows the results. It is interesting that different researchers play different roles in the network. Mark Newman and Vito Latora have similar structural patterns to that of Dr. Barabási. Some other researchers like Robert form a tightknit group with him. Panther++ successfully recognizes those researchers with similar structural positions.


Code and Open Data set
Panther Open Data set
We execute our code on Aminer coauthor network, and open the dataset and the execuated results here.
1. Data Overview
|
FileName
|
Desciprtion
|
Statistics
|
Size
|
|
Author profile with Index as unique identification
|
1,712,433 authors
|
167 MB
|
|
|
Coauthor network (ID to ID)
|
1,560,640 authors, 4,258,946 relationships
|
29 MB
|
|
|
Mapping from Index to ID
|
1,560,640 authors
|
9.2 MB
|
|
| AMinerCoauthor.panther.zip | Top-50 similar authors for any author by Panther | 1,560,640 authors | 189M |
| AMinerCoauthor.panther++.zip | Top-50 similar authors for any author by Panther++ | 1,560,640 authors | 362M |
| Tencent Networks | 9 randomly sampled Tencent "following" networks | 2.3G |
2. Data Description
The dataset is from Tencent Weibo, a popular Twitter-like microblogging service in China, and consists of over 355,591,065 users and 5,958,853,072 “following” relationships. The weight associated with each edge is set as 1.0 uniformly. This is the largest network in our experiments. We mainly use it to evaluate the efficiency performance of our methods.
Dataset size: #nodes / #edges
Tencent0: 6,523/10,000
Tencent1: 25,844/50,000
Tencent2: 48,837/100,000
Tencent3: 169,209/500,000
Tencent4: 230,103/1,000,000
Tencent5: 443,070/5,000,000
Tencent6: 702,049/10,000,000
Tencent7: 2,767,344/50,000,000
Tencent8: 5,355,507/100,000,000
We upload our source code here. We can find the main file under the folder src, and other source files under the folders of lib and include. In addition, we also include the source code of building kd-tree and querying kd-tree under the folder res/ann. Ann is the open source code by David M. Mount and Sunil Arya. You can find the detail from here.
How to run?
Note: please download and unzip the source code and run the code under linux OS.
Assume the root path is YOUR_PATH. Later we use $YOUR_PATH to represent the root path.
1) Enter $YOUR_PATH/rdsextr/ann and execute "make clean realclean"
2) Execute "make linux-g++" or "make macosx-g++" according to your platform.
3) Execute "mkdir $YOUR_PATH/run"
4) Exectue "cmake $YOUR_PATH/rdsextr" under the folder $YOUR_PATH/run
5) Execute "make" under the folder $YOUR_PATH/run
6) Execute "ln -sf $YOUR_PATH/rdsextr/ann/bin/ann_sample $YOUR_PATH/run"
7) Execute "mkdir $YOUR_PATH/run/data"
8) Download the data files AMiner-Author.zip, AMinerCoauthor.graph.zip, and AMinerCoauthor.dict.zip, unzip and copy them into the folder $YOUR_PATH/run/data
9) Copy run.sh into the folder $YOUR_PATH/run and execute it.
10) The results will be saved under the folder of result.
11) Copy findtwosimilar.py into the folder $YOUR_PATH/run and run it to check top 50 similar authors for any given authorIndex based on your ourput results.
Note: we can directly copy the results AMinerCoauthor.panther.zip and AMinerCoauthor.panther++.zip into the folder $YOUR_PATH/run/result, unzip them, and run findtwosimilar.py to check the results.
4. Example Results
Copy the results AMinerCoauthor.panther.zip and AMinerCoauthor.panther++.zip into the folder $YOUR_PATH/run/result, and run findtwosimilar.py with parameters "AMinerCoauthor_5_50_484 973203" to get the following example in AMiner coauthor network:
The format is Index:Name, Index:Name, ....
973203 : Michael I. Jordan , 1449150 : Tommi S. Jaakkola , 251055 : Lawrence K. Saul , 550923 : Zoubin Ghahramani , 747456 : Francis R. Bach , 919673 : David A. Patterson , 390372 : Armando Fox , 875089 : Andrew Y. Ng , 224633 : Robert A. Jacobs , 1509711 : David M. Blei , 509109 : Zhihua Zhang , 100313 : Marina MeilÄÂÂÂÂÂÂÂÂÂ�? , 864212 : Laurent El Ghaoui , 587351 : G. Lanckriet , 1699437 : Guang Dai , 16780 : David A. Rosenbaum , 521844 : Daniel V. Klein , 516753 : Chiranjib Bhattacharyya , 1222775 : Peter Bodik , 1139828 : Martin J. Wainwright , 1428458 : Richard M. Karp , 913466 : Tao Li , 24505 : S. J. Russell , 1614699 : Philip N. Sabes , 1701301 : Junming Yin , 129332 : Percy Liang , 834716 : L. Xu , 1087416 : Chris Bishop , 985597 : Eric Xing , 114681 : Simon Lacoste-Julien , 1340466 : Ling Huang , 1631519 : Alice X. Zheng , 872236 : I. Saira Mian , 830260 : Benjamin Taskar , 730910 : Nello Cristianini , 1436638 : Jens Nilsson , 668855 : XuanLong Nguyen , 634341 : Neil Lawrence , 287397 : Guillaume Obozinski , 1016632 : David Cohn , 1004012 : Wey Fun , 431393 : A. Rizki , 1603434 : P. Smyth , 308498 : D. Radisky , 1163105 : S. Sankararaman , 1284539 : Ben Liblit , 816298 : Jon McAuliffe , 777569 : Dan Klein , 1369029 : I. Stoica , 1171121 : Chris H. Q. Ding