
A B-tree is a type of search tree where every node (except possibly for the root) contains between m and 2m keys for some positive integer m, and all leaves have the same distance to the root. We study sequences of B-trees that can arise from successively inserting keys, and in particular present a bijection between such sequences (which we call histories) and a special type of increasing trees. We describe the set of permutations for the keys that belong to a given history, and also show how to use this bijection to analyse statistics associated with B-trees.
We consider the number of occurrences of subwords (non-consecutive sub-sequences) in a given word. We first define the notion of subword entropy of a given word that measures the maximal number of occurrences among all possible subwords. We then give upper and lower bounds of minimal subword entropy for words of fixed length in a fixed alphabet, and also showing that minimal subword entropy per letter has a limit value. A better upper bound of minimal subword entropy for a binary alphabet is then given by looking at certain families of periodic words. We also give some conjectures based on experimental observations
We study the size of the connected component of early typical vertices in a subcritical inhomogeneous random graph with a kernel of preferential attachment type. The principal tools in our analysis are, first, a coupling of the neighbourhood of a typical vertex in the graph to a killed branching random walk and, second, an asymptotic result for the number of particles absorbed at the killing barrier in this branching random walk.
We introduce a model of tree-rooted planar maps weighted by their number of 2-connected blocks. We study its enumerative properties and prove that it undergoes a phase transition. We give the distribution of the size of the largest 2-connected blocks in the three regimes (subcritical, critical and supercritical) and further establish that the scaling limit is the Brownian Continuum Random Tree in the critical and supercritical regimes, with respective rescalings root n/log(n) and root n.
It is celebrated that a simple random walk on Z and Z(2) returns to the initial vertex v infinitely many times during infinitely many transitions, which is said recurrent, while it returns to v only finite times on Z(d) for d >= 3, which is said transient. It is also known that a simple random walk on a growing region on Z(d) can be recurrent depending on growing speed for any fixed d. This paper shows that a simple random walk on {0, 1,..., N}(n) with an increasing n and a fixed N can be recurrent depending on the increasing speed of n. Precisely, we are concerned with a specific model of a random walk on a growing graph (RW(o)GG) and show a phase transition between the recurrence and transience of the random walk regarding the growth speed of the graph. For the proof, we develop a pausing coupling argument introducing the notion of weakly less homesick as graph growing (weakly LHaGG).
Making use of a newly developed package in the computer algebra system SageMath, we show how to perform a full asymptotic analysis by means of the Mellin transform with explicit error bounds. As an application of the method, we answer a question of Bona and DeJonge on 132-avoiding permutations with a unique longest increasing subsequence that can be translated into an inequality for a certain binomial sum.
In the hardcore model, certain vertices in a graph are active: the active vertices must form an independent set. We extend this to a multicoloured version: instead of simply being active or not, the active vertices are assigned a colour; active vertices of the same colour must not be adjacent. This models a scenario in which two neighbouring resources may interfere when active - eg, short-range radio communication. However, there are multiple channels (colours) available; they only interfere if both use the same channel. Other applications include routing in fibreoptic networks. We analyse Glauber dynamics. Vertices update their status at random times, at which a uniform colour is proposed: the vertex is assigned that colour if it is available; otherwise, it is set inactive. We find conditions for fast mixing of these dynamics. We also use them to model a queueing system: vertices only serve customers in their queue whilst active. The mixing estimates are applied to establish positive recurrence of the queue lengths, and bound their expectation in equilibrium.
The Twelvefold Way represents Rota's classification, addressing the most fundamental enumeration problems and their associated combinatorial counting formulas. These distinct problems are connected to enumerating functions defined from a set of elements denoted by N into another one K. The counting solutions for the twelve problems are well known. We are interested in unranking algorithms. Such an algorithm is based on an underlying total order on the set of structures we aim at constructing. By taking the rank of an object, i.e. its number according to the total order, the algorithm outputs the structure itself after having built it. One famous total order is the lexicographic order: it is probably the one that is the most used by people when one wants to order things. While the counting solutions for Rota's classification have been known for years it is interesting to note that three among the problems have yet no lexicographic unranking algorithm. In this paper we aim at providing algorithms for the last three cases that remain without such algorithms. After presenting in detail the solution for set partitions associated with the famous Stirling numbers of the second kind, we explicitly explain how to adapt the algorithm for the two remaining cases. Additionally, we propose a detailed and fine-grained complexity analysis based on the number of bitwise arithmetic operations.
We present an analysis of the depth-first search algorithm in a random digraph model with independent outdegrees having an arbitrary distribution with finite variance. The results include asymptotics for the distribution of the stack index and depths of the search. The search yields a series of trees of finite size before and after the exploration of a giant tree. Our analysis mainly concerns the giant tree. Most results are first order. This analysis proposed by Donald Knuth in his next to appear volume of The Art of Computer Programming gives interesting insight in one of the most elegant and efficient algorithm for graph analysis due to Tarjan.
It is a classic result in spectral theory that the limit distribution of the spectral measure of random graphs G(n, p) converges to the semicircle law in case np tends to infinity with n. The spectral measure for random graphs G(n, c/n) however is less understood. In this work, we combine and extend two combinatorial approaches by Bauer and Golinelli (2001) and Enriquez and Menard (2016) and approximate the moments of the spectral measure by counting walks that span trees.
We propose the class of galled tree-child networks which is obtained as intersection of the classes of galled networks and tree-child networks. For the latter two classes, (asymptotic) counting results and stochastic results have been proved with very different methods. We show that a counting result for the class of galled tree-child networks follows with similar tools as used for galled networks, however, the result has a similar pattern as the one for tree-child networks. In addition, we also consider the (suitably scaled) numbers of reticulation nodes of random galled tree-child networks and show that they are asymptotically normal distributed. This is in contrast to the limit laws of the corresponding quantities for galled networks and tree-child networks which have been both shown to be discrete.
A fringe subtree of a rooted tree is a subtree that consists of a vertex and all its descendants. The number of distinct fringe subtrees in random trees has been studied by several authors, notably because of its connection to tree compaction algorithms. Here, we obtain a very precise result for binary search trees: it is shown that the number of distinct fringe subtrees in a binary search tree with n leaves is asymptotically equal to c(1)n/log n for a constant c(1) approximate to 2.4071298335, both in expectation and with high probability. This was previously shown to be a lower bound, our main contribution is to prove a matching upper bound. The method is quite general and can also be applied to similar problems for other tree models.
For d >= 2 and i.i.d. d-dimensional observations X-(1),X-(2),& mldr; with independent Exponential(1) coordinates, we revisit the study by Fill and Naiman (Electron. J. Probab., 25:Paper No. 92, 24 pp., 2020) of the boundary (relative to the closed positive orthant), or "frontier", Fn of the closed Pareto record-setting (RS) region RSn & ratio;={ 0 <= x is an element of & Ropf; (d )& ratio; x not less than X ( i ) for all 1 <= i <= n } at epoch n, where 0 <= x means that 0 <= x j for 1 <= j <= d and x < y means that x (j) < y (j) for 1 <= j <= d. With x( + )& ratio; = & sum; (d)( j )= 1 x j = & Vert; x & Vert; (& lscr; 1) ,let F - (n )& ratio; = min { x (+) & ratio; x is an element of F (n) } and F + (n) & ratio; = max { x (+) & ratio; x is an element of F (n)}. Almost surely ,there are for each n unique vectors lambda (n) is an element of F (n) and tau (n) is an element of F- n such that F (+)( n )= ( lambda (n) ) + and F- n =(tau n)+; we refer to lambda n and tau n as the leading and trailing points, respectively, of the frontier . Fill and Naiman provided rather sharp information about the typical and almost surebehaviour of (F-n(+)) , but somewhat crude information about (F-n(-)), namely, that for any epsilon > 0 and c (n) -> infinity we have & Popf;(F-n(-)- lnn is an element of (-(2 +epsilon) ln ln lnn, cn))-> 1 (describing typical behavior)and almost surely lim sup(n ->infinity)F-n- lnn/ln lnn <= 0 and lim inf n -> infinity F-n( -) - ln n / ln ln ln n is an element of [-2,-1]. In this paper we use the theory of generators (minima of F n ; specifically, we use a characterization of these objects and analysis of their expected count) together with the first - and second-moment methods to improve considerably the trailing-point location results to F-n(-)- (lnn- ln ln lnn)P ->- ln(d- 1) (describing typical behavior) and, for d >= 3, almost surely lim sup(n ->infinity)[F-n(-)- (ln n- ln ln ln n)] <= - l n ( d- 2) + ln2 and lim inf (n ->infinity)[F-n(-)- (ln n- ln ln ln n)] >=- ln d- ln 2. Our probabilistic upper bounds on F- n follow immediately from probabilistic upper bounds we derive on the(provably) slightly larger quantity F-n(-)& ratio;= (minimum coordinate - sum of any remaining record at epoch n) which is of interest in its own right and can be bounded more easily than F- n(-) using the second-moment method. This paper is a full-length version of the extended abstract.
In the asymptotic analysis of regular sequences as defined by Allouche and Shallit, it is usually advisable to study their summatory function because the original sequence has a too fluctuating behaviour. It might be that the process of taking the summatory function has to be repeated if the sequence is fluctuating too much. In this paper we show that for all regular sequences except for some degenerate cases, repeating this process finitely many times leads to a “nice” asymptotic expansion containing periodic fluctuations whose Fourier coefficients can be computed using the results on the asymptotics of the summatory function of regular sequences by the first two authors of this paper. In a recent paper, Hwang, Janson, and Tsai perform a thorough investigation of divide-and-conquer recurrences. These can be seen as 2-regular sequences. By considering them as the summatory function of their forward difference, the results on the asymptotics of the summatory function of regular sequences become applicable. We thoroughly investigate the case of a polynomial toll function.
Galled trees appear in problems concerning admixture, horizontal gene transfer, hybridization, and recombination. Building on a recursive enumerative construction, we study the asymptotic behavior of the number of rooted binary unlabeled (normal) galled trees as the number of leaves n increases, maintaining a fixed number of galls g. We find that the exponential growth with n of the number of rooted binary unlabeled normal galled trees with g galls has the same value irrespective of the value of g >= 0. The subexponential growth, however, depends on g; it follows c(g)n(2g-3/2), where c(g) is a constant dependent on g. Although for each g, the exponential growth is approximately 2.4833(n), summing across all g, the exponential growth is instead approximated by the much larger 4.8230(n).
We use a novel decomposition to create succinct data structures -- supporting a wide range of operations on static trees in constant time -- for a variety tree classes, extending results of Munro, Nicholson, Benkner, and Wild. Motivated by the class of AVL trees, we further derive asymptotics for the information-theoretic lower bound on the number of bits needed to store tree classes whose generating functions satisfy certain functional equations. In particular, we prove that AVL trees require approximately $0.938$ bits per node to encode.
In this paper we obtain some new results on the enumeration of parking functions and labeled forests. We introduce new statistics both for parking functions and for labeled forests that are connected to each other by means of a bijection. We determine the joint distribution of two statistics on parking functions and their counterparts on labeled forests. Our results on labeled forests also serve to explain the mysterious equidistribution between two seemingly unrelated statistics in parking functions recently identified by Stanley and Yin and give an explicit bijection between the two statistics.
Consider a tree T = (V, E) with root. and an edge length function l : E -> R+. The phylogenetic covariance matrix of T is the matrix C with rows and columns indexed by L, the leaf set of T, with entries C(i, j) := Sigma(e is an element of[i Lambda j,o]) l(e), for each i, j is an element of L. Recent work [Gorman & Lladser 2023] has shown that the phylogenetic covariance matrix of a large but random binary tree T is significantly sparsified, with overwhelmingly high probability, under a change-of-basis to the so-called Haar-like wavelets of T. Notably, this finding enables manipulating the spectrum of covariance matrices of large binary trees without the necessity to store them in computer memory but instead performing two post-order traversals of the tree [Gorman & Lladser 2023]. Building on the methods of the aforesaid paper, this manuscript further advances their sparsification result to encompass the broader class of k-regular trees, for any given k >= 2. This extension is achieved by refining existing asymptotic formulas for the mean and variance of the internal path length of random k-regular trees, utilizing hypergeometric function properties and identities.
The height of a random PATRICIA tree built from independent, identically distributed infinite binary strings with arbitrary diffuse probability distribution mu on {0, 1}(N) is studied. We show that the expected height grows asymptotically sublinearly in the number of leaves for any such mu, but can be made to exceed any specific sublinear growth rate by choosing mu appropriately.
We consider a synchronous process of particles moving on the vertices of a graph G, introduced by Cooper, McDowell, Radzik, Rivera and Shiraga (2018). Initially, M particles are placed on a vertex of G. In subsequent time steps, all particles that are located on a vertex inhabited by at least two particles jump independently to a neighbour chosen uniformly at random. The process ends at the first step when no vertex is inhabited by more than one particle; we call this (random) time step the dispersion time. In this work we study the case where G is the complete graph on n vertices and the number of particles is M = n/2 + alpha n(1/2) + o(n(1/2)), alpha is an element of R. This choice of M corresponds to the critical window of the process, with respect to the dispersion time. We show that the dispersion time, if rescaled by n(-1/2), converges in p-th mean, as n -> infinity and for any p is an element of R, to a continuous and almost surely positive random variable T-alpha. We find that T-alpha is the absorption time of a standard logistic branching process, thoroughly investigated by Lambert (2005), and we determine its expectation. In particular, in the middle of the critical window we show that E[T-0] = pi(3/2)/root 7, and furthermore we formulate explicit asymptotics when vertical bar alpha vertical bar gets large that quantify the transition into and out of the critical window. We also study the random variable counting the total number of jumps that are performed by the particles until the dispersion time is reached and prove that, if rescaled by n ln n, it converges to 2/7 in probability.