The maximum likelihood threshold of a statistical model is the minimum number of datapoints required to fit the model via maximum likelihood estimation. In this paper we determine the maximum likelihood thresholds of generic linear concentration models. This turns out to be the number one would expect from a naive dimension count, which is surprising and nontrivial to prove given that the maximum likelihood threshold is a semi-algebraic concept. We also describe geometrically how a linear concentration model can fail to exhibit this generic behavior and briefly discuss connections to rigidity theory.
In this expository article, we summarize what is known about maximum likelihood thresholds of Gaussian models, paying special attention to connections with rigidity theory.
We provide a new axiom system for flag matroids, characterize representability of uniform flag matroids, and give forbidden minor characterizations of full flag matroids that are representable over 𝔽_2 and 𝔽_3 along with regular full flag matroids. We also provide different equivalent characterizations for regular full flag matroids.
Associated to each graph G is a Gaussian graphical model. Such models are often used in high-dimensional settings, i.e. where there are relatively few data points compared to the number of variables. The maximum likelihood threshold of a graph is the minimum number of data points required to fit the corresponding graphical model using maximum likelihood estimation. Graphical lasso is a method for selecting and fitting a graphical model. In this project, we ask: when graphical lasso is used to select and fit a graphical model on n data points, how likely is it that n is greater than or equal to the maximum likelihood threshold of the corresponding graph? Our results are a series of computational experiments.
The maximum likelihood threshold (MLT) of a graph $G$ is the minimum number of samples to almost surely guarantee existence of the maximum likelihood estimate in the corresponding Gaussian graphical model. We give a new characterization of the MLT in terms of rigidity-theoretic properties of $G$ and use this characterization to give new combinatorial lower bounds on the MLT of any graph. We use the new lower bounds to give high-probability guarantees on the maximum likelihood thresholds of sparse Erd{\"o}s-R\'enyi random graphs in terms of their average density. These examples show that the new lower bounds are within a polylog factor of tight, where, on the same graph families, all known lower bounds are trivial. Based on computational experiments made possible by our methods, we conjecture that the MLT of an Erd{\"o}s-R\'enyi random graph is equal to its generic completion rank with high probability. Using structural results on rigid graphs in low dimension, we can prove the conjecture for graphs with MLT at most $4$ and describe the threshold probability for the MLT to switch from $3$ to $4$. We also give a geometric characterization of the MLT of a graph in terms of a new "lifting" problem for frameworks that is interesting in its own right. The lifting perspective yields a new connection between the weak MLT (where the maximum likelihood estimate exists only with positive probability) and the classical Hadwiger-Nelson problem.
A 1965 result of Crapo shows that every elementary lift of a matroid $M$ can be constructed from a linear class of circuits of $M$. In a recent paper, Walsh generalized this construction by defining a rank-$k$ lift of a matroid $M$ given a rank-$k$ matroid $N$ on the set of circuits of $M$, and conjectured that all matroid lifts can be obtained in this way. In this sequel paper we simplify Walsh's construction and show that this conjecture is true for representable matroids but is false in general. This gives a new way to certify that a particular matroid is non-representable, which we use to construct new classes of non-representable matroids. Walsh also applied the new matroid lift construction to gain graphs over the additive group of a non-prime finite field, generalizing a construction of Zaslavsky for these special groups. He conjectured that this construction is possible on three or more vertices only for the additive group of a non-prime finite field. We show that this conjecture holds for four or more vertices, but fails for exactly three.
The maximum likelihood threshold (MLT) of a graph $G$ is the minimum number of samples to almost surely guarantee existence of the maximum likelihood estimate in the corresponding Gaussian graphical model. Recently a new characterization of the MLT in terms of rigidity-theoretic properties of $G$ was proved \cite{Betal}. This characterization was then used to give new combinatorial lower bounds on the MLT of any graph. We continue this line of research by exploiting combinatorial rigidity results to compute the MLT precisely for several families of graphs. These include graphs with at most $9$ vertices, graphs with at most 24 edges, every graph sufficiently close to a complete graph and graphs with bounded degrees.
A graph G is fully reconstructible in ℂ d if the graph is determined from its d-dimensional measurement variety. The full reconstructibility problem has been solved for d = 1 and d = 2. For d = 3, some necessary and some sufficient conditions are known and K 5 , 5 falls squarely within the gap in the theory. In this paper, we show that K 5 , 5 is fully reconstructible in ℂ 3.
We characterize the combinatorial types of symmetric frameworks in the plane that are minimally generically symmetry-forced infinitesimally rigid when the symmetry group consists of rotations and translations. Along the way, we use tropical geometry to show how a construction of Edmonds that associates a matroid to a submodular function can be used to give a description of the algebraic matroid of a Hadamard product of two linear spaces in terms of the matroids of each linear space. This leads to new, short, proofs of Laman's theorem, and a theorem of Jord{á}n, Kaszanitzky, and Tanigawa, and Malestein and Theran characterizing the minimally generically symmetry-forced rigid graphs in the plane when the symmetry group contains only rotations.
A graph G is fully reconstructible in ℂd if the graph is determined from its d-dimensional measurement variety. The full reconstructibility problem has been solved for d=1 and d=2. For d=3, some necessary and some sufficient conditions are known and K5,5 falls squarely within the gap in the theory. In this paper, we show that K5,5 is fully reconstructible in ℂ3.
We study the problem of low-rank matrix completion for symmetric matrices. The minimum rank of a completion of a generic partially specified symmetric matrix depends only on the location of the specified entries, and not their values, if complex entries are allowed. When the entries are required to be real, this is no longer the case and the possible minimum ranks are called typical ranks. We give a combinatorial description of the patterns of specified entires of $n\times n$ symmetric matrices that have $n$ as a typical rank. Moreover, we describe exactly when such a generic partial matrix is minimally completable to rank $n$. We also characterize the typical ranks for patterns of entries with low maximal typical rank.
Many questions in applied algebraic geometry boil down to asking which polynomial functions within some family are generically finite-to-one. When the family consists of the coordinate projections of an irreducible variety, matroids arise. Two particular applications that have driven much research include matrix completion and rigidity theory. The ground sets of the matroids that appear therein are the edge sets of graphs.
A graph $G$ is fully reconstructible in $\mathbb{C}^d$ if the graph is determined from its $d$-dimensional measurement variety. The full reconstructibility problem has been solved for $d=1$ and $d=2$. For $d=3$, some necessary and some sufficient conditions are known and $K_{5,5}$ falls squarely within the gap in the theory. In this paper, we show that $K_{5,5}$ is fully reconstructible in $\mathbb{C}^3$.
We study the properties of alignment, a form of implicit regularization, in linear neural networks under gradient descent. We define alignment for fully connected networks with multidimensional outputs and show that it is a natural extension of alignment in networks with 1-dimensional outputs as defined by Ji and Telgarsky, 2018. While in fully connected networks, there always exists a global minimum corresponding to an aligned solution, we analyze alignment as it relates to the training process. Namely, we characterize when alignment is an invariant of training under gradient descent by providing necessary and sufficient conditions for this invariant to hold. In such settings, the dynamics of gradient descent simplify, thereby allowing us to provide an explicit learning rate under which the network converges linearly to a global minimum. We then analyze networks with layer constraints such as convolutional networks. In this setting, we prove that gradient descent is equivalent to projected gradient descent, and that alignment is impossible with sufficiently large datasets.
We study the properties of alignment, a form of implicit regularization, in linear neural networks under gradient descent. We define alignment for fully connected networks with multidimensional outputs and show that it is a natural extension of alignment in networks with 1-dimensional outputs as defined by Ji and Telgarsky, 2018. While in fully connected networks, there always exists a global minimum corresponding to an aligned solution, we analyze alignment as it relates to the training process. Namely, we characterize when alignment is an invariant of training under gradient descent by providing necessary and sufficient conditions for this invariant to hold. In such settings, the dynamics of gradient descent simplify, thereby allowing us to provide an explicit learning rate under which the network converges linearly to a global minimum. We then analyze networks with layer constraints such as convolutional networks. In this setting, we prove that gradient descent is equivalent to projected gradient descent, and that alignment is impossible with sufficiently large datasets.
We consider the task of learning a causal graph in the presence of latent confounders given i.i.d.~samples from the model. While current algorithms for causal structure discovery in the presence of latent confounders are constraint-based, we here propose a score-based approach. We prove that under assumptions weaker than faithfulness, any sparsest independence map (IMAP) of the distribution belongs to the Markov equivalence class of the true model. This motivates the \emph{Sparsest Poset} formulation - that posets can be mapped to minimal IMAPs of the true model such that the sparsest of these IMAPs is Markov equivalent to the true model. Motivated by this result, we propose a greedy algorithm over the space of posets for causal structure discovery in the presence of latent confounders and compare its performance to the current state-of-the-art algorithms FCI and FCI+ on synthetic data.
A finite unit norm tight frame is a collection of r vectors in Rn that generalizes the notion of orthonormal bases. The affine finite unit norm tight frame variety is the Zariski closure of the set of finite unit norm tight frames. Determining the fiber of a projection of this variety onto a set of coordinates is called the algebraic finite unit norm tight frame completion problem. Our techniques involve the algebraic matroid of an algebraic variety, which encodes the dimensions of fibers of coordinate projections. This work characterizes the bases of the algebraic matroid underlying the variety of finite unit norm tight frames in R3. Partial results towards similar characterizations for finite unit norm tight frames in Rn with n≥4 are also given. We provide a method to bound the degree of the projections based off of combinatorial data.
Given a dissimilarity map $\delta$ on finite set $X$, the set of ultrametrics (equidistant tree metrics) which are $l^\infty$-nearest to $\delta$ is a tropical polytope. We give an internal description of this tropical polytope which we use to derive a polynomial-time checkable test for the condition that all ultrametrics $l^\infty$-nearest to $\delta$ have the same tree structure. It was shown by Ardila and Klivans \cite{ardila-klivans2006} that the set of all ultrametrics on a finite set of size $n$ is the Bergman fan associated to the matroid underlying the complete graph on $n$ vertices. Therefore, we derive our results in the more general context of Bergman fans of matroids. This added generality allows our results to be used on dissimilarity maps where only a subset of the entries are known.
We consider the problem of exact low-rank matrix completion from a geometric viewpoint: given a partially filled matrix M, we keep the positions of specified and unspecified entries fixed, and study how the minimal completion rank depends on the values of the known entries. If the entries of the matrix are complex numbers, then for a fixed pattern of locations of specified and unspecified entries there is a unique completion rank which occurs with positive probability. We call this rank the generic completion rank. Over the real numbers there can be multiple ranks that occur with positive probability; we call them typical completion ranks. We introduce these notions formally, and provide a number of inequalities and exact results on typical and generic ranks for different families of patterns of known and unknown entries.
The Cayley-Menger variety is the Zariski closure of the set of vectors specifying the pairwise squared distances between $n$ points in $\mathbb{R}^d$. This variety is fundamental to algebraic approaches in rigidity theory. We study the tropicalization of the Cayley-Menger variety. In particular, when $d = 2$, we show that it is the Minkowski sum of the set of ultrametrics on $n$ leaves with itself, and we describe its polyhedral structure. We then give a new, tropical, proof of Laman's theorem.
Katherine St. John合作论文数City University of New York;The Graduate Center of the;Department of Computer Science1
M. A. Steel合作论文数Biomathematics Research Centre
Department of Mathematics and Statistics
University of Canterbury
1