We propose a general framework for modelling network data that is designed to describe aspects of non-exchangeable networks. Conditional on latent (unobserved) variables, the edges of the network are generated by their finite growth history (with latent orders) while the marginal probabilities of the adjacency matrix are modeled by a generalization of a graph limit function (or a graphon). In particular, we study the estimation, clustering and degree behavior of the network in our setting. We determine (i) the minimax estimator of a composite graphon with respect to squared error loss; (ii) that spectral clustering is able to consistently detect the latent membership when the block-wise constant composite graphon is considered under additional conditions; and (iii) we are able to construct models with heavy-tailed empirical degrees under specific scenarios and parameter choices. This explores why and under which general conditions non-exchangeable network data can be described by a stochastic block model. The new modelling framework is able to capture empirically important characteristics of network data such as sparsity combined with heavy tailed degree distribution, and add understanding as to what generative mechanisms will make them arise. Keywords: statistical network analysis, exchangeable arrays, stochastic block model, nonlinear stochastic processes.
Summary We consider local linear estimation of the graphon function, which determines probabilities of pairwise edges between nodes in an unlabelled network. Real-world networks are typically characterized by node heterogeneity, with different nodes exhibiting different degrees of interaction. Existing approaches to graphon estimation are limited to local constant approximations, and are not designed to estimate heterogeneity across the full network. In this paper, we show how continuous node covariates can be employed to estimate heterogeneity in the network via a local linear graphon estimator. We derive the bias and variance of an oracle-based local linear graphon estimator, and thus obtain the mean integrated squared error optimal bandwidth rule. We also provide a plug-in bandwidth selection procedure that makes local linear estimation for unlabelled networks practically feasible. The finite-sample performance of our approach is investigated in a simulation study, and the method is applied to a school friendship network and an email network to illustrate its advantages over existing methods.
Data structures known as $k$-d trees have numerous applications in scientific computing, particularly in areas of modern statistics and data science such as range search in decision trees, clustering, nearest neighbors search, local regression, and so forth. In this article we present a scalable mechanism to construct $k$-d trees for distributed data, based on approximating medians for each recursive subdivision of the data. We provide theoretical guarantees of the quality of approximation using this approach, along with a simulation study quantifying the accuracy and scalability of our proposed approach in practice.
Designing scalable estimation algorithms is a core challenge in modern statistics. Here we introduce a framework to address this challenge based on parallel approximants, which yields estimators with provable properties that operate on the entirety of very large, distributed data sets. We first formalize the class of statistics which admit straightforward calculation in distributed environments through independent parallelization. We then show how to use such statistics to approximate arbitrary functional operators in appropriate spaces, yielding a general estimation framework that does not require data to reside entirely in memory. We characterize the $L^2$ approximation properties of our approach and provide fully implemented examples of sample quantile calculation and local polynomial regression in a distributed computing environment. A variety of avenues and extensions remain open for future work.
Abstract This article introduces a new class of models for multiple networks. The core idea is to parameterize a distribution on labeled graphs in terms of a Fréchet mean graph (which depends on a user-specified choice of metric or graph distance) and a parameter that controls the concentration of this distribution about its mean. Entropy is the natural parameter for such control, varying from a point mass concentrated on the Fréchet mean itself to a uniform distribution over all graphs on a given vertex set. We provide a hierarchical Bayesian approach for exploiting this construction, along with straightforward strategies for sampling from the resultant posterior distribution. We conclude by demonstrating the efficacy of our approach via simulation studies and two multiple-network data analysis examples: one drawn from systems biology and the other from neuroscience. This article has online supplementary materials.
We adopt the statistical framework on robustness proposed by Watson and Holmes in 2016 and then tackle the practical challenges that hinder its applicability to network models. The goal is to evaluate how the quality of an inference for a network feature degrades when the assumed model is misspecified. Decision theory methods aimed to identify model missespecification are applied in the context of network data with the goal of investigating the stability of optimal actions to perturbations to the assumed model. Here the modified versions of the model are contained within a well defined neighborhood of model space. Our main challenge is to combine stochastic optimization and graph limits tools to explore the model space. As a result, a method for robustness on exchangeable random networks is developed. Our approach is inspired by recent developments in the context of robustness and recent works in the robust control, macroeconomics and financial mathematics literature and more specifically and is based on the concept of graphon approximation through its empirical graphon.
We consider that a network is an observation, and a collection of observed networks forms a sample. In this setting, we provide methods to test whether all observations in a network sample are drawn from a specified model. We achieve this by deriving the joint asymptotic properties of average subgraph counts as the number of observed networks increases but the number of nodes in each network remains finite. In doing so, we do not require that each observed network contains the same number of nodes, or is drawn from the same distribution. Our results yield joint confidence regions for subgraph counts, and therefore methods for testing whether the observations in a network sample are drawn from: a specified distribution, a specified model, or from the same model as another network sample. We present simulation experiments and an illustrative example on a sample of brain networks where we find that highly creative individuals' brains present significantly more short cycles than found in less creative people. for this article are available online.
Abstract Algorithms are tools for decision-making. England's A-level results fiasco shows what happens when our tools are ill-suited or ill-designed and their workings poorly explained. By Sofia Olhede and Patrick J. Wolfe
We characterize the large-sample properties of network modularity in the presence of covariates, under a natural and flexible nonparametric null model. This provides for the first time an objective measure of whether or not a particular value of modularity is meaningful. In particular, our results quantify the strength of the relation between observed community structure and the interactions in a network. Our technical contribution is to provide limit theorems for modularity when a community assignment is given by nodal features or covariates. These theorems hold for a broad class of network models over a range of sparsity regimes, as well as weighted, multi-edge, and power-law networks. This allows us to assign $p$-values to observed community structure, which we validate using several benchmark examples in the literature. We conclude by applying this methodology to investigate a multi-edge network of corporate email interactions.
The study of complex relationships among the elements of a large collection of random variables lead to the development of a number of areas in probability and statistics such as probabilistic network analysis or random matrix theory. The aim of the workshop was to address the challenge to develop a coherent mathematical framework within which these areas can be integrated, for a successful analysis of massive and complicated data sets.
AbstractSofia Olhede and Patrick Wolfe discuss the current state of data-driven automation and its implications for jobs
The ubiquity of sensing devices, the low cost of data storage, and the commoditization of computing have together led to a big data revolution. We discuss the implication of this revolution for statistics, focusing on how our discipline can best contribute to the emerging field of data science.
The growing ubiquity of algorithms in society raises a number of fundamental questions concerning governance of data, transparency of algorithms, legal and ethical frameworks for automated algorithmic decision-making and the societal impacts of algorithmic automation itself. This article, an introduction to the discussion meeting issue of the same title, gives an overview of current challenges and opportunities in these areas, through which accelerated technological progress leads to rapid and often unforeseen practical consequences. These consequences-ranging from the potential benefits to human health to unexpected impacts on civil society-are summarized here, and discussed in depth by other contributors to the discussion meeting issue.This article is part of a discussion meeting issue 'The growing ubiquity of algorithms in society: implications, impacts and innovations'.
As nations race for dominance in the field of artificial intelligence, Sofia Olhede and Patrick Wolfe consider the implications for statistics and statisticians
Overcomplete representations such as wavelets and windowed Fourier expansions have become mainstays of modern statistical data analysis. In the present work, in the context of general finite frames, we derive an oracle expression for the mean quadratic risk of a linear diagonal de-noising procedure which immediately yields the optimal linear diagonal estimator. Moreover, we obtain an expression for an unbiased estimator of the risk of any smooth shrinkage rule. This last result motivates a set of practical estimation procedures for general finite frames that can be viewed as the generalization of the classical procedures for orthonormal bases. A simulation study verifies the effectiveness of the proposed procedures with respect to the classical ones and confirms that the correlations induced by frame structure should be explicitly treated to yield an improvement in estimation precision.
We prove that counting copies of any graph $F$ in another graph $G$ can be achieved using basic matrix operations on the adjacency matrix of $G$. Moreover, the resulting algorithm is competitive for medium-sized $F$: our algorithm recovers the best known complexity for rooted 6-clique counting and improves on the best known for 9-cycle counting. Underpinning our proofs is the new result that, for a general class of graph operators, matrix operations are homomorphisms for operations on rooted graphs.
We consider that a network is an observation, and a collection of observed networks forms a sample. In this setting, we provide methods to test whether all observations in a network sample are drawn from a specified model. We achieve this by deriving, under the null of the graphon model, the joint asymptotic properties of average subgraph counts as the number of observed networks increases but the number of nodes in each network remains finite. In doing so, we do not require that each observed network contains the same number of nodes, or is drawn from the same distribution. Our results yield joint confidence regions for subgraph counts, and therefore methods for testing whether the observations in a network sample are drawn from: a specified distribution, a specified model, or from the same model as another network sample. We present simulation experiments and an illustrative example on a sample of brain networks where we find that highly creative individuals' brains present significantly more short cycles.
As automated decisions affect more and different areas of our lives, we are faced with ethical and legal questions that are likely to change the way we think about algorithms and the law. By Sofia Olhede and Patrick Wolfe