This paper develops the exact linear relationship between the leading eigenvector of the unnormalized modularity matrix and the eigenvectors of the adjacency matrix. We propose a method for approximating the leading eigenvector of the modularity matrix, and we derive the error of the approximation. There is also a complete proof of the equivalence between normalized adjacency clustering and normalized modularity clustering. Numerical experiments show that normalized adjacency clustering can be as twice efficient as normalized modulairty clustering.
Although there has been a substantial amount of research conducted that concerns the matrix B and g-inverses, B~, of B, we feel that there are still many facts about the structure of g-inverses for B which have not yet been discovered. For example, a great deal of attention has been given to the problem of ob taining a particular form for an entire B~ but very little attention has been devoted to the structure of the various submatrices that might appear in a B~. One of the purposes of this paper is to show that the submatrices of B~ are entirely independent of each other and then to completely characterize the various classes of matrices which are blocks in a g-inverse for B, and thus characterize all g-inverses for B. As pointed out by Rao (1971), there is a need for efficient algorithms for computing g-inverses of B. As a consequence of our work, it will follow that the submatrices of any g-inverse for B may be computed separately and independently so that the sizes of the matrices involved in a computational scheme for a B~ may be greatly reduced, and thus opening the door to faster and more efficient algorithms. The equations on which such computations must be based are given in Theorem 4.2. It is hoped that our work may be useful in the future development of computational methods for g-inverses of B.
DAYAR, TUGRUL. Stability and Conditioning Issues on the Numerical Solution of Markov Chains. (Under the direction of William James Stewart.) Markovian modeling and analysis is extensively used in many disciplines in evaluating the performance of existing systems and in analyzing and designing systems to be developed. In most cases the systems under consideration are large, and therefore require the employment of numerical solution techniques. With the changing face of computing environments, there are several issues that need to be addressed in the numerical solution of Markov chains. Firstly, stability and conditioning issues in the solution of such large systems has to be of concern, and one should be in the lookout for algorithms that are more efficient and that produce more accurate results. This may be achieved by improving existing algorithms or developing new ones that take advantage of the increased computer power. In this thesis, the effects of using a modified version of Gaussian elimination in the iterative aggregation-disaggregation technique, which is especially suited to the solution of ill-conditioned nearly completely decomposable Markov chains, has been investigated. A second solution technique geared towards multivector computers has been extended to include the systems of interest. A third solution technique, which is useful in handling vast state spaces, is shown to be quite effective in computing performance measures for systems having highly unbalanced stationary probabilities. Experiments on real-life applications (involving mostly large systems) have been conducted in each case. Finally, the possibility of reducing the state space by combining states in nearly completely decomposable chains is discussed.
Updating a given a matrix A(mxn) by a rank-one matrix B = cd(T), where c and d are appropriately sized column vectors, is a common practice throughout all applied areas of mathematics, science, and engineering. Because rank is often tied to the number of degrees of freedom or the level of independence in underlying models or data, it can be imperative to know exactly how the update term affects rank. While it is well known that a rank-one update can only increase or decrease rank by at most one, there is not a widely known formula for exactly how this occurs. This note presents an expression in simply stated terms for the exact rank of a rank-one updated matrix.
In this paper the exact linear relation between the leading eigenvectors of the modularity matrix and the singular vectors of an uncentered data matrix is developed.Based on this analysis the concept of a modularity component is defined, and its properties are developed.It is shown that modularity component analysis can be used to cluster data similar to how traditional principal component analysis is used except that modularity component analysis does not require data centering.
That the Perron root of a square nonnegative matrix varies continuously with the entries in is a corollary of theorems regarding continuity of eigenvalues or roots of polynomial equations, the proofs of which necessarily involve complex numbers. But since continuity of the Perron root is a question that is entirely in the field of real numbers, it seems reasonable that there should exist a development involving only real analysis. This article presents a simple and completely self-contained development that depends only on real numbers and first principles.
We use a cluster ensemble to determine the number of clusters, k, in a group of data. A consensus similarity matrix is formed from the ensemble using multiple algorithms and several values for k. A random walk is induced on the graph defined by the consensus matrix and the eigenvalues of the associated transition probability matrix are used to determine the number of clusters. For noisy or high-dimensional data, an iterative technique is presented to refine this consensus matrix in way that encourages a block-diagonal form. It is shown that the resulting consensus matrix is generally superior to existing similarity matrices for this type of spectral analysis.
It is well known that good initializations can improve the speed and accuracy of the solutions of many nonnegative matrix factorization (NMF) algorithms. Many NMF algorithms are sensitive with respect to the initialization of W or H or both. This is especially true of algorithms of the alternating least squares (ALS) type, including the two new ALS algorithms that we present in this paper. We compare the results of six initialization procedures (two standard and four new) on our ALS algorithms. Lastly, we discuss the practical issue of choosing an appropriate convergence criterion.
The world's largest matrix computation. (This chapter is out of date and needs a major overhaul.) One of the reasons why Google TM is such an effective search engine is the PageRank TM algorithm developed by Google's founders, Larry Page and Sergey Brin, when they were graduate students at Stanford University. PageRank is determined entirely by the link structure of the World Wide Web. It is recomputed about once a month and does not involve the actual content of any Web pages or individual queries. Then, for any particular query, Google finds the pages on the Web that match that query and lists those pages in the order of their PageRank. Imagine surfing the Web, going from page to page by randomly choosing an outgoing link from one page to get to the next. This can lead to dead ends at pages with no outgoing links, or cycles around cliques of interconnected pages. So, a certain fraction of the time, simply choose a random page from the Web. This theoretical random walk is known as a Markov chain or Markov process. The limiting probability that an infinitely dedicated random surfer visits any particular page is its PageRank. A page has high rank if other pages with high rank link to it. Let W be the set of Web pages that can be reached by following a chain of hyperlinks starting at some root page, and let n be the number of pages in W. For Google, the set W actually varies with time, but by June 2004, n was over 4 billion. Let G be the n-by-n connectivity matrix of a portion of the Web, that is, g ij = 1 if there is a hyperlink to page i from page j and g ij = 0 otherwise. The matrix G can be huge, but it is very sparse. Its jth column shows the links on the jth page. The number of nonzeros in G is the total number of hyperlinks in W .
A novel framework for consensus clustering is presented which has the ability to determine both the number of clusters and a final solution using multiple algorithms. A consensus similarity matrix is formed from an ensemble using multiple algorithms and several values for k. A variety of dimension reduction techniques and clustering algorithms are considered for analysis. For noisy or high-dimensional data, an iterative technique is presented to refine this consensus matrix in way that encourages algorithms to agree upon a common solution. We utilize the theory of nearly uncoupled Markov chains to determine the number, k , of clusters in a dataset by considering a random walk on the graph defined by the consensus matrix. The eigenvalues of the associated transition probability matrix are used to determine the number of clusters. This method succeeds at determining the number of clusters in many datasets where previous methods fail. On every considered dataset, our consensus method provides a final result with accuracy well above the average of the individual algorithms.
Cluster analysis is a field of data analysis that extracts underlying patterns in data. One application of cluster analysis is in text-mining, the analysis of large collections of text to find similarities between documents. We used a collection of about 30,000 tweets extracted from Twitter just before the World Cup started. A common problem with real world text data is the presence of linguistic noise. In our case it would be extraneous tweets that are unrelated to dominant themes. To combat this problem, we created an algorithm that combined the DBSCAN algorithm and a consensus matrix. This way we are left with the tweets that are related to those dominant themes. We then used cluster analysis to find those topics that the tweets describe. We clustered the tweets using k-means, a commonly used clustering algorithm, and Non-Negative Matrix Factorization (NMF) and compared the results. The two algorithms gave similar results, but NMF proved to be faster and provided more easily interpreted results. We explored our results using two visualization tools, Gephi and Wordle.
A website's ranking on Google can spell the difference between success and failure for a new business. NCAA football ratings determine which schools get to play for the big money in postseason bowl games. Product ratings influence everything from the clothes we wear to the movies we select on Netflix. Ratings and rankings are everywhere, but how exactly do they work? Who's #1? offers an engaging and accessible account of how scientific rating and ranking methods are created and applied to a variety of uses.Amy Langville and Carl Meyer provide the first comprehensive overview of the mathematical algorithms and methods used to rate and rank sports teams, political candidates, products, Web pages, and more. In a series of interesting asides, Langville and Meyer provide fascinating insights into the ingenious contributions of many of the field's pioneers. They survey and compare the different methods employed today, showing why their strengths and weaknesses depend on the underlying goal, and explaining why and when a given method should be considered. Langville and Meyer also describe what can and can't be expected from the most widely used systems.The science of rating and ranking touches virtually every facet of our lives, and now you don't need to be an expert to understand how it really works. Who's #1? is the definitive introduction to the subject. It features easy-to-understand examples and interesting trivia and historical facts, and much of the required mathematics is included.
We explore the geometrical interpretation of the PCA based clustering algorithm Principal Direction Divisive Partitioning (PDDP). We give several examples where this algorithm breaks down, and suggest a new method, gap partitioning, which takes into account natural gaps in the data between clusters. Geometric features of the PCA space are derived and illustrated and experimental results are given which show our method is comparable on the datasets used in the original paper on PDDP.
A website's ranking on Google can spell the difference between success and failure for a new business. NCAA football ratings determine which schools get to play for the big money in postseason bowl games. Product ratings influence everything from the clothes we wear to the movies we select on Netflix. Ratings and rankings are everywhere, but how exactly do they work? Who's #1? offers an engaging and accessible account of how scientific rating and ranking methods are created and applied to a variety of uses. Amy Langville and Carl Meyer provide the first comprehensive overview of the mathematical algorithms and methods used to rate and rank sports teams, political candidates, products, Web pages, and more. In a series of interesting asides, Langville and Meyer provide fascinating insights into the ingenious contributions of many of the field's pioneers. They survey and compare the different methods employed today, showing why their strengths and weaknesses depend on the underlying goal, and explaining why and when a given method should be considered. Langville and Meyer also describe what can and can't be expected from the most widely used systems. The science of rating and ranking touches virtually every facet of our lives, and now you don't need to be an expert to understand how it really works. Who's #1? is the definitive introduction to the subject. It features easy-to-understand examples and interesting trivia and historical facts, and much of the required mathematics is included.
Cluster Analysis is a field of Data Mining used to extract underlying patterns in unclassified data. Many existing clustering algorithms are inadequate in that they require knowledge of how many clusters exist in the data, otherwise known as k, and that their underlying assumptions make them ineffective in certain situations. The method of Consensus Clustering seeks to rectify the latter problem by incorporating the results of multiple clustering algorithms to achieve one final grouping. We investigate a novel method of Iterative Consensus Clustering (ICC) which solves both issues. The iteration of the consensus clustering technique widens the eigengap associated with the Perron cluster, giving a more definitive, and more accurate, estimation of the number of clusters, k.