The β-model is popular for characterizing the commonly observed degree heterogeneity phenomenon in real-world networks. In this study, we develop a cycle counting approach to estimate n node-specific parameters in the β-model for moderate or extremely sparse networks. Our proposed estimators, called Cycle Counting Ratio (CCR) Estimator, are based on the log-ratios of two network cycle counting statistics with explicit expressions and therefore easy to compute. We focus on conditions to guarantee statistical properties of the single estimator for each node. Under the very weak conditions that max_t θ_t → 0 and θ_t θ_1 →∞, we show that the CCR estimator is consistent and achieves the minimax rate in terms of the mean squared error, which is the squared signal-to-noise ratio for _t up to a constant factor. Here, _t is the CCR estimator of the node-specific parameter β_t, θ_t = exp(β_t) and θ=(θ_1, …, θ_n). Even if the whole network density is close to the Erdős-Rényi lower bound log n/n, the CCR estimator for the single parameter β_t is still consistent as long as θ_t θ_1 →∞. To the best of our knowledge, this is the first time to derive the minimax rate and consistency result under such weak conditions. Under a slight stronger condition, we further establish its uniform consistency and asymptotic normality, whose asymptotic variance is θ_t θ_1. Numerical studies and an application to a sparse network data set demonstrate our theoretical findings.
Evaluating diagnostic accuracy for multi-category outcomes remains a significant challenge, primarily due to computational limitations in existing performance metrics. The Polytomous Discrimination Index (PDI) has emerged as an order-agnostic solution suitable for nominal classifications. However, its broader adoption has been constrained by the lack of efficient implementation, especially for its variance estimation, which typically would require computationally intensive bootstrapping procedures. In this work, we address this limitation by proposing a novel asymptotic variance estimator for the PDI. Our method integrates classical U-statistic theory with recent advances in combinatorics, offering a scalable and theoretically grounded alternative. To assess the performance of the proposed approach, we conduct extensive simulation studies and observe remarkable gain in computing time. We further apply our method to a real-world brain image analysis where deep neural networks are used as a diagnostic tool. We can efficiently report the accuracy of the neural networks with different depth specifications.
For a fixed integer j≥1 and 0<p<1, we study the probability generating function (pgf) equation (1+2p) g(x)=2p x^j+g(x-px+px^2), 0≤ x≤1 , which governs the limiting degree distribution {p_k} of a family of evolving network models. The cases j=1 and j=2 are the treelike fast-growth model of Feng and Hu and the homogeneous evolving network of Feng, Li and Hu. We prove that for every j the equation has a unique pgf solution, of mean 2j, and we determine its coefficient tail exactly: p_k=k^-1-ρ Ψ_j(log_λk)+o(k^-1-ρ), where λ=1+p, ρ=log(1+2p)/log(1+p) is independent of j, and Ψ_j is continuous, strictly positive and 1-periodic, with explicit Fourier coefficients. This resolves two conjectures of Feng and coauthors: (1) the power-law order p_k=Θ(k^-1-ρ) and (2) its refinement to the multiplicatively periodic form p_k_j(log_λk) k^-1-ρ. The periodic factor is genuinely non-constant for p near 1, and, for the two network models, for all p outside a discrete set. Consequently, p_k is asymptotic to no constant multiple of k^-1-ρ. Our method is a self-contained local analysis of the supercritical Galton-Watson process with offspring law 1+Bernoulli(p), inspected at an independent geometric time. This time-changed process solves the equation observed by Feng and coauthors. The main results of this paper were obtained by the multi-agent system Eureka and have subsequently been verified by the authors.
The Zagreb index, which is defined as the sum of squares of degrees of the nodes of a tree, was studied in previous works by martingale techniques for random non-plane recursive trees and classes of random trees which are close to random plane recursive trees. These techniques are not easily amended to the generalized Zagreb index, which is defined similarly but with squares replaced by higher powers. We use the moment-transfer approach to (i) obtain the first-order asymptotics of moments and (ii) prove limit laws for the (suitably normalized) generalized Zagreb index for random non-plane and plane recursive trees. For the former, we show that for all higher powers the limit law is normal; for the latter, we show for cubes and fourth powers that it is a non-normal law.
We study the number of triangles $T_n$ in the sparse $\beta$ -model on n vertices, a random graph model that captures degree heterogeneity in real-world networks. Using the norms of the heterogeneity parameter vector, we first determine the asymptotic mean and variance of $T_n$ . Next, by applying the Malliavin-Stein method, we derive a non-asymptotic upper bound on the Kolmogorov distance between the normalized $T_n$ and the standard normal distribution. Under an additional assumption on degree heterogeneity, we further prove the asymptotic normality for $T_n$ as $n\to\infty$ .
This paper investigates the Fréchet mean of the Erdős-Rényi random graph G_n,p with respect to the Frobenius distance on graph Laplacians, a metric that captures global structural information beyond local edge flips. We first characterize the Fréchet mean set as consisting of quasi-regular graphs (i.e., graphs where all vertex degrees differ by at most one). We then analyze the asymptotic behavior of the Frobenius distance F_n=d_F(G_n,p,R) as n→∞, where R is any Fréchet mean. Closed-form expressions for the mean and variance of F_n^2 are derived, which are invariant to the choice of R. Leveraging these results, we establish several weak convergence laws for the Frobenius distance over all regimes of p ∈ (0,1) as n →∞. Finally, under the scaling condition n^2 p(1-p) →∞ we prove the asymptotic normality of this distance, which exhibits a phase transition governed by the growth rate of np(1-p). Our results reveal how metric selection fundamentally shapes Fréchet mean geometry in random graphs.
The degree of balance, a simple topological index of structural balance, in a signed Erd & odblac;s-R & eacute;nyi random graph model is investigated in this paper. This index in such a graph model is shown to have the asymptotic normality as the graph size tends to infinity and the edge probability tends to zero. Our main result is derived through the application of the delta method, and is based on the joint asymptotic normality of the numbers of balanced and unbalanced triangles, where a dependency graph approach is also used.
The Youden index is a widely used metric for assessing diagnostic accuracy in two-class classification problems, particularly for determining the optimal decision cutoff point. In this article, we extend this concept to a multiple-category classification framework by introducing semi-parametric estimators for the generalized Youden index and its associated optimal threshold, under the Lehmann assumption. Our proposed estimators are much easier to implement than the traditional nonparametric estimators. We further establish the theoretical properties of these estimators, ensuring their consistency and asymptotic normality. To evaluate the effectiveness of the proposed methods, we conduct extensive simulation studies and apply them to a real-world liver cancer dataset, demonstrating their practical applicability in medical diagnostics.
Assortativity measures the tendency of a vertex to bond with another based on their structural or functional features. The assortativity coefficient was originally proposed to specify the node degree–degree correlation for unweighted, undirected networks. This paper proposes a class of rank-based assortativity measures for weighted, directed networks, building upon and extending the work of Litvak and van der Hofstad. These new measures offer improved robustness to variations in edge weights, particularly in networks with extreme-valued edges, by reducing sensitivity to outliers in node connectivity patterns. Extensive simulation studies, based on the Erdös–Rényi random network model and preferential attachment random networks, are employed to compare the robustness of the proposed measures with existing assortativity measures. Finally, the proposed assortativity measures are applied to the networks generated from the World Input–Output Database, yielding insights that differ significantly from those obtained through existing methods.
In this paper, we study the asymptotic behavior of the generalized Zagreb indices of the classical Erdős–Rényi (ER) random graph G ( n , p ), as $n\to\infty$ . For any integer $k\ge1$ , we first give an expression for the k th-order generalized Zagreb index in terms of the number of star graphs of various sizes in any simple graph. The explicit formulas for the first two moments of the generalized Zagreb indices of an ER random graph are then obtained from this expression. Based on the asymptotic normality of the numbers of star graphs of various sizes, several joint limit laws are established for a finite number of generalized Zagreb indices with a phase transition for p in different regimes. Finally, we provide a necessary and sufficient condition for any single generalized Zagreb index of G ( n , p ) to be asymptotic normal.
The elephant random walk (ERW) is a discrete-time random walk on Z with a memory about the whole past. It has been shown that the asymptotic behavior of the ERW depends heavily on a memory parameter 0 <= p <= 1. In this work, we establish the strong invariance principle for the ERW in the diffusive regime 0 <= p < 3/4 and the critical regime p = 3/4, which enhances the known results of Coletti, Gava and Schutz.
The asymptotic behavior of the Jaccard index in $G(n,p)$, the classical Erdös-Rényi random graphs model, is studied in this paper, as $n$ goes to infinity. We first derive the asymptotic distribution of the Jaccard index of any pair of distinct vertices, as well as the first two moments of this index. Then the average of the Jaccard indices over all vertex pairs in $G(n,p)$ is shown to be asymptotically normal under an additional mild condition that $np\to\infty$ and $n^2(1-p)\to\infty$.
In this work, we consider a randomized version of the deterministic pseudofractal graph model, which was first proposed in the physical literature. Our model is a sequence of dynamic networks, in which the increment of network size at each time depends on the current size and a sequence of evolving parameters. We first show that the network size is closely related to a branching process in the varying environment. Under a mild assumption, we then prove that the asymptotic degree distribution in our model obeys the power law with an exponent three. Based on this asymptotic degree distribution, we show that the clustering coefficient of our network model converges to a given constant in probability.
In this paper we propose an evolving network model, which is a randomized version of the pseudofractal graphs by introducing an evolutionary parameter 0 < p < 1. Our network model grows exponentially over time, and can be generated in an iterative manner: at each time step, with probability p each existing edge recruits independently a new node and connects to it with both endpoints. We first briefly discuss the network size, which can correspond to a supercritical branching process. Then, it shows that the asymptotic degree distribution in our network model can be uniquely determined by a functional equation of its probability generating function.(c) 2022 Elsevier B.V. All rights reserved.
Distances between nodes are one of the most essential subjects in the study of complex networks. In this paper, we investigate the asymptotic behaviors of two types of distances in a model of geographic attachment networks (GANs): the typical distance and the flooding time. By generating an auxiliary tree and using a continuous-time branching process, we demonstrate that in this model the typical distance is asymptotically normal, and the flooding time converges to a given constant in probability as well.
Computation of hypervolume under ROC manifold (HUM) is necessary to evaluate biomarkers for their capability to discriminate among multiple disease types or diagnostic groups. However the original definition of HUM involves multiple integration and thus a medical investigation for multi‐class receiver operating characteristic (ROC) analysis could suffer from huge computational cost when the formula is implemented naively. We introduce a novel graph‐based approach to compute HUM efficiently in this article. The computational method avoids the time‐consuming multiple summation when sample size or the number of categories is large. We conduct extensive simulation studies to demonstrate the improvement of our method over existing R packages. We apply our method to two real biomedical data sets to illustrate its application.
In this paper, we analyze the time series data of the case and death counts of COVID-19 that broke out in China in December, 2019. The study period is during the lockdown of Wuhan. We exploit functional data analysis methods to analyze the collected time series data. The analysis is divided into three parts. First, the functional principal component analysis is conducted to investigate the modes of variation. Second, we carry out the functional canonical correlation analysis to explore the relationship between confirmed and death cases. Finally, we utilize a clustering method based on the Expectation-Maximization (EM) algorithm to run the cluster analysis on the counts of confirmed cases, where the number of clusters is determined via a cross-validation approach. Besides, we compare the clustering results with some migration data available to the public.
Medical multi-category diagnostic problems may involve discrete biomarkers. Many traditional accuracy measures are based on the assumption that all biomarkers follow continuous distributions and consequently may underestimate the true discrimination ability of the discrete markers. In particular, we focus on Hypervolume Under ROC Manifold (HUM) in this paper and propose an extension of the familiar continuous version of HUM to incorporate discrete biomarkers with ties. Statistical estimation and inference procedures are proposed along with asymptotic properties. We carry out simulation studies to examine the performance of our proposed estimators for the new HUM measure. A real medical example is analysed to illustrate our methodology.
The random variable Zn is investigated,the maximal node degree in a random k-tree at step n for k≥2.It is shown that as n→∞,Zn/n(k-1)/k has an almost sure limit,which is a positive random variable.The result is also extended to the random k-Apollonian networks model for k≥3.
Abstract The accessibility percolation model is investigated on random rooted labeled trees. More precisely, the number of accessible leaves (i.e. increasing paths) Zn and the number of accessible vertices Cn in a random rooted labeled tree of size n are jointly considered in this work. As n → ∞, we prove that (Zn, Cn) converges in distribution to a random vector whose probability generating function is given in an explicit form. In particular, we obtain that the asymptotic distributions of Zn + 1 and Cn are geometric distributions with parameters e/(1 + e) and 1/e, respectively. Much of our analysis is performed in the context of local weak convergence of random rooted labeled trees.
Hosam Mahmoud合作论文数Department of Statistics2