We obtain local weak limits in probability for Collapsed Branching Processes (CBP), which are directed random networks obtained by collapsing random-sized families of individuals in a general continuous-time branching process. The local weak limit of a given CBP, as the network grows, is shown to be a related continuous-time branching process stopped at an independent exponential time. This is done through an explicit coupling of the in-components of vertices with the limiting object. We also show that the in-components of a finite collection of uniformly chosen vertices locally weakly converge (in probability) to i.i.d. copies of the above limit, reminiscent of propagation of chaos in interacting particle systems. We obtain as special cases novel descriptions of the local weak limits of directed preferential and uniform attachment models. We also outline some applications of our results for analyzing the limiting in-degree and PageRank distributions.
We study the evolution of opinions on a directed network with community structure. Individuals update their opinions synchronously based on a weighted average of their neighbors' opinions, their own previous opinions, and external media signals. Our model is akin to the popular Friedkin-Johnsen model, and is able to incorporate factors such as stubbornness, confirmation bias, selective exposure, and multiple topics, which are believed to play an important role in the formation of opinions. Our main result shows that, in the large graph limit, the opinion process concentrates around its mean-field approximation for any level of edge density, provided the average degree grows to infinity. Moreover, we show that the opinion process exhibits propagation of chaos. We also give results for the trajectories of individual vertices and the stationary version of the opinion process, and prove that the limits in time and in the size of the network commute. The mean-field approximation is explicit and can be used to quantify consensus and polarization.
We propose and analyze a mathematical model for the evolution of opinions on directed complex networks. Our model generalizes the popular DeGroot and Friedkin-Johnsen models by allowing vertices to have attributes that may influence the opinion dynamics. We start by establishing sufficient conditions for the existence of a stationary opinion distribution on any fixed graph, and then provide an increasingly detailed characterization of its behavior by considering a sequence of directed random graphs having a local weak limit. Our most explicit results are obtained for graph sequences whose local weak limit is a marked Galton-Watson tree, in which case our model can be used to explain a variety of phenomena, for example, conditions under which consensus can be achieved, mechanisms in which opinions can become polarized, and the effect of disruptive stubborn agents on the formation of opinions. Funding: This work was supported by the National Science Foundation [Grants NSF-DMS-1929298 and CMMI-2243261].
We present new results on community recovery based on the PageRank Nibble algorithm on a sparse directed stochastic block model (dSBM). Our results are based on a characterization of the local weak limit of the dSBM and the limiting PageRank distribution. This characterization allows us to estimate the probability of misclassification for any given connection kernel and any given number of seeds (vertices whose community label is known). The fact that PageRank is a local algorithm that can be efficiently computed in both a distributed and asynchronous fashion, makes it an appealing method for identifying members of a given community in very large networks where the identity of some vertices is known.
For a vertex-weighted directed graph G(Vn,En;An) on the vertices Vn={1,2,…,n}, we study the distribution of a Markov chain {R(k):k≥0} on Rn such that the ith component of R(k), denoted Ri(k), corresponds to the value of the process on vertex i at time k. We focus on processes {R(k):k≥0} where the value of Ri(k+1) depends only on the values {Rj(k):j→i} of its inbound neighbors, and possibly on vertex attributes. We then show that, provided G(Vn,En;An) converges in the local weak sense to a marked Galton–Watson process, the dynamics of the process for a uniformly chosen vertex in Vn can be coupled, for any fixed k, to a process {R0̸(r):0≤r≤k} constructed on the limiting marked Galton–Watson tree. Moreover, we derive sufficient conditions under which R0̸(k) converges, as k→∞, to a random variable R∗ that can be characterized in terms of the attracting endogenous solution to a branching distributional fixed-point equation. Our framework can also be applied to processes {R(k):k≥0} whose only source of randomness comes from the realization of the graph G(Vn,En;An).
We study the all-time supremum of the perturbed branching random walk, known to be the endogenous solution to the high-order Lindley equation: W=DmaxY,max1≤i≤N(Wi+Xi),where the {Wi} are independent copies of W, independent of the random vector (Y,N,{Xi}) taking values in R×N×R∞. Under Kesten assumptions, this solution satisfies P(W>t)∼He−αt,t→∞,where α>0 solves the Cramér–Lundberg equation E∑i=1NeαXi=1. This paper establishes the tail asymptotics of W by using the forward iterations of the map defining the fixed-point equation combined with a change of measure along a randomly chosen path. This new approach provides an explicit representation of the constant H and gives rise to unbiased and strongly efficient estimators for the rare event probabilities P(W>t).
The goal of this paper is to provide a general purpose result for the coupling of exploration processes of random graphs, both undirected and directed, with their local weak limits when this limit is a marked Galton-Watson process. This class includes in particular the configuration model and the family of inhomogeneous random graphs with rank-1 kernel. Vertices in the graph are allowed to have attributes on a general separable metric space and can potentially influence the construction of the graph itself. The coupling holds for any fixed depth of a breadth-first exploration process.
We present a hybrid importance sampling estimator that is strongly efficient for tail probabilities of the all-time maximum of a branching random walk, where the increments satisfy a Cramer-Lundberg condition. The estimator uses conditional Monte Carlo in combination with the population dynamics algorithm to compute an expression for the tail of the distribution obtained from a spine change of measure. It has computational complexity (measured by the number of input random vectors required) that is independent of the offspring distribution, allowing for fast computation even when the mean number of offspring is very large. We remark on consistency of this estimator and give numerical examples.
We characterize the tail behavior of the distribution of the PageRank of a uniformly chosen vertex in a directed preferential attachment graph and show that it decays as a power law with an explicit exponent that is described in terms of the model parameters. Interestingly, this power law is heavier than the tail of the limiting in-degree distribution, which goes against the commonly accepted power law hypothesis. This deviation from the power law hypothesis points at the structural differences between the inbound neighborhoods of typical vertices in a preferential attachment graph versus those in static random graph models where the power law hypothesis has been proven to hold (e.g., directed configuration models and inhomogeneous random digraphs). In addition to characterizing the PageRank distribution of a typical vertex, we also characterize the explicit growth rate of the PageRank of the oldest vertex as the network size grows.
The focus of this work is the asymptotic analysis of the tail distribution of Google's PageRank algorithm on large scale-free directed networks. In particular, the main theorem provides the convergence, in the Kantorovich-Rubinstein metric, of the rank of a randomly chosen vertex in graphs generated via either a directed configuration model or an inhomogeneous random digraph. The theorem fully characterizes the limiting distribution by expressing it as a random sum of i.i.d. copies of the attracting endogenous solution to a branching distributional fixed-point equation. In addition, we provide the asymptotic tail behavior of the limit and use it to explain the effect that in-degree/out-degree correlations in the underlying graph can have on the qualitative performance of PageRank.
We study a family of directed random graphs whose arcs are sampled independently of each other, and are present in the graph with a probability that depends on the attributes of the vertices involved. In particular, this family of models includes as special cases the directed versions of the Erdős‐Rényi model, graphs with given expected degrees, the generalized random graph, and the Poissonian random graph. We establish a phase transition for the existence of a giant strongly connected component and provide some other basic properties, including the limiting joint distribution of the degrees and the mean number of arcs. In particular, we show that by choosing the joint distribution of the vertex attributes according to a multivariate regularly varying distribution, one can obtain scale‐free graphs with arbitrary in‐degree/out‐degree dependence.
We study the typical behavior of Google's PageRank algorithm on inhomogeneous random digraphs, including directed versions of the Erdos-Renyi model, the Chung-Lu model, the Poissonian random graph and the generalized random graph. Specifically, we show that the rank of a randomly chosen vertex converges weakly to the attracting endogenous solution to the stochastic fixed-point equation R (D) double under bar Sigma(N)(i=1) CiRi + Q, where (N, Q, {C-i}(i >= 1)) is a real-valued vector with N epsilon N, and the {R-i} are i.i.d. copies of R, independent of (N, Q, {C-i}(i >= 1)); R (D) double under bar denotes equality in distribution. This result provides further evidence of the power-law behavior of PageRank on graphs whose in-degree distribution follows a power law. (C) 2019 Elsevier B.V. All rights reserved.
We propose a model for optimizing the last-mile delivery of n packages from a distribution center to their final recipients, using a strategy that combines the use of ride-sharing platforms (e.g., Uber or Lyft) with traditional in-house van delivery systems. The main objective is to compute the optimal reward offered to private drivers for each of the n packages such that the total expected cost of delivering all packages is minimized. Our technical approach is based on the formulation of a discrete sequential packing problem, in which bundles of packages are picked up from the warehouse at random times during the interval [Formula: see text]. Our theoretical results include both exact and asymptotic (as [Formula: see text]) expressions for the expected number of packages that are picked up by time T. They are closely related to the classical Rényi’s parking/packing problem. Our proposed framework is scalable with the number of packages.
We propose a model for optimizing the last-mile delivery of n packages from a distribution center to their final recipients, using a strategy that combines the use of ride-sharing platforms (e.g., Uber or Lyft) with traditional in-house van delivery systems. The main objective is to compute the optimal reward offered to private drivers for each of the n packages such that the total expected cost of delivering all packages is minimized. Our technical approach is based on the formulation of a discrete sequential packing problem, in which bundles of packages are picked up from the warehouse at random times during the interval [0,T]. Our theoretical results include both exact and asymptotic (as n→∞) expressions for the expected number of packages that are picked up by time T . They are closely related to the classical Rényi’s parking/packing problem. Our proposed framework is scalable with the number of packages.
We consider the distributional fixed-point equation: R 𝒟= Q ∨( ⋁_i=1^N C_i R_i ), where the {R_i} are i.i.d. copies of R, independent of the vector (Q, N, {C_i}), where N ∈ℕ, Q, {C_i}≥ 0 and P(Q > 0) > 0. By setting W = log R, X_i = log C_i, Y = log Q it is equivalent to the high-order Lindley equation W 𝒟=max{ Y, max_1 ≤ i ≤ N (X_i + W_i) }. It is known that under Kesten assumptions, P(W > t) ∼ H e^-α t, t →∞, where α>0 solves the Cramér-Lundberg equation E [ ∑_j=1^N C_i ^α] = E[ ∑_i=1^N e^α X_i] = 1. The main goal of this paper is to provide an explicit representation for P(W > t), which can be directly connected to the underlying weighted branching process where W is constructed and that can be used to construct unbiased and strongly efficient estimators for all t. Furthermore, we show how this new representation can be directly analyzed using Alsmeyer's Markov renewal theorem, yielding an alternative representation for the constant H. We provide numerical examples illustrating the use of this new algorithm.
Motivated by database locking problems in today’s massive computing systems, we analyze a queueing network with many servers in parallel (files) to which jobs (writing access requests) arrive according to a Poisson process. Each job requests simultaneous access to a random number of files in the database and will lock them for a random period of time. Alternatively, one can think of a queueing system where jobs are split into several fragments that are then randomly routed to specific servers in the network to be served in a synchronized fashion. We assume that the system operates on a first-come, first-served basis. The synchronization and service discipline create blocking and idleness among the servers, which leads to a strict stability condition compared with other distributed queueing models. We analyze the stationary waiting time distribution of jobs under a many-server limit and provide exact tail asymptotics. These asymptotics generalize the celebrated Cramér–Lundberg approximation for the single-server queue.
We study the convergence of the population dynamics algorithm, which produces sample pools of random variables having a distribution that closely approximates that of the {\em special endogenous solution} to a stochastic fixed-point equation of the form: $$R\stackrel{\mathcal D}{=} \Phi( Q, N, \{ C_i \}, \{R_i\}),$$ where $(Q, N, \{C_i\})$ is a real-valued random vector with $N \in \mathbb{N}$, and $\{R_i\}_{i \in \mathbb{N}}$ is a sequence of i.i.d. copies of $R$, independent of $(Q, N, \{C_i\})$; the symbol $\stackrel{\mathcal{D}}{=}$ denotes equality in distribution. Specifically, we show its convergence in the Wasserstein metric of order $p$ ($p \geq 1$) and prove the consistency of estimators based on the sample pool produced by the algorithm.
We consider a discrete-time Markov chain $\boldsymbol{\Phi}$ on a general state-space ${\sf X}$, whose transition probabilities are parameterized by a real-valued vector $\boldsymbol{\theta}$. Under the assumption that $\boldsymbol{\Phi}$ is geometrically ergodic with corresponding stationary distribution $\pi(\boldsymbol{\theta})$, we are interested in estimating the gradient $\nabla \alpha(\boldsymbol{\theta})$ of the steady-state expectation $$\alpha(\boldsymbol{\theta}) = \pi( \boldsymbol{\theta}) f.$$ To this end, we first give sufficient conditions for the differentiability of $\alpha(\boldsymbol{\theta})$ and for the calculation of its gradient via a sequence of finite horizon expectations. We then propose two different likelihood ratio estimators and analyze their limiting behavior.
We analyze the distribution of the distance between two nodes, sampled uniformly at random, in digraphs generated via the directed configuration model, in the supercritical regime. Under the assumption that the covariance between the in-degree and out-degree is finite, we show that the distance grows logarithmically in the size of the graph. In contrast with the undirected case, this can happen even when the variance of the degrees is infinite. The main tool in the analysis is a new coupling between a breadth-first graph exploration process and a suitable branching process based on the Kantorovich-Rubinstein metric. This coupling holds uniformly for a much larger number of steps in the exploration process than existing ones, and is therefore of independent interest.
We study the typical behavior of a generalized version of Google’s PageRank algorithm on a large family of inhomogeneous random digraphs. This family includes as special cases directed versions of classical models such as the Erdős-Rényi model, the Chung-Lu model, the Poissonian random graph and the generalized random graph, and is suitable for modeling scale-free directed complex networks where the number of neighbors a vertex has is related to its attributes. In particular, we show that the rank of a randomly chosen node in a graph from this family converges weakly to the attracting endogenous solution to the stochastic fixed-point equation R D = N ∑ i=1 CiRi +Q, where (N ,Q, {Ci}i≥1) is a real-valued vector with N ∈ {0, 1, 2, ...}, the {Ri} are i.i.d. copies of R, independent of (N ,Q, {Ci}i≥1), with {Ci} i.i.d. and independent of (N ,Q); D = denotes equality in distribution. This result can then be used to provide further evidence of the powerlaw behavior of PageRank on scale-free graphs.
N. Litvak合作论文数Faculty of Electrical Engineering, Mathematics and Computer Science
University of Twente2