We study when local reduced density operators, viewed as quantum marginals, can be assembled into a global quantum state with a prescribed Markov structure. The starting point is a canonical logarithmic construction T(ℛ), the noncommutative analogue of the junction-tree formula for decomposable graphical models. Unlike in the classical case, this formal construction may fail: noncommutativity can prevent it from being a normalized state with the prescribed marginals. We prove that this obstruction is captured exactly by a trace condition. For two overlapping marginals, and for clique marginals on a chordal graph, the condition Tr(T(ℛ))=1 is equivalent to the existence of a quantum Markov completion. When it exists, the completion is unique, equal to T(ℛ), and selected by the maximum entropy principle. In the two-clique case, we also give an equivalent conditional reconstruction characterization: the two natural one-sided sandwich reconstructions agree if and only if the trace condition holds. We introduce the global quantum information g I(𝒢)_ρ associated with a chordal graph 𝒢 and show that it is a relative-entropy discrepancy from ρ to the logarithmic candidate, with a trace correction when the candidate is not normalized. We also prove an intersection property for strictly positive quantum conditional independence. Three-qubit Pauli examples illustrate how the quantum obstructions are real: local consistency, feasibility, Markov feasibility, and maximum entropy can all separate.
Abstract This gives an overview of the basic theory of exponential families, including definitions, analytic properties, behaviour under conditioning and marginalization, basic asymptotic results for curved exponential families, standard conjugate families and their properties, and iterative methods for maximizing exponential likelihood functions.
We study quantum analogues of Bayesian networks on a directed acyclic graph (DAG), distinguishing two constructions for positive definite density operators on finite-dimensional tensor-product Hilbert spaces. The intrinsic construction starts from a joint state and its conditional-independence properties. The extrinsic construction assembles a state sequentially from prescribed local quantum kernels, following an ordering compatible with the arrows of the DAG. For the intrinsic construction, we prove the equivalence of the ordered, local, and global directed Markov properties, together with entropy, recursive-factorization, and logarithmic characterizations. The extrinsic construction always gives a normalized state and recovers each kernel as a conditional on all preceding systems. The same kernel, however, need not be recovered from the marginal on the vertex and its parents; a three-qubit example exhibits this obstruction. We prove that independence of the chosen topological ordering is sufficient exactly for transitive DAGs: every order-invariant kernel family then yields an intrinsically directed Markov state. Finally, we associate a logarithmic candidate with every positive definite state and DAG, prove that it is subnormalized, and show that the trace-one candidate is a directed Markov state. Both the candidate and the excess global information are invariant under DAG Markov equivalence.
This note establishes that if a sequence Pn,n=1,… of probability measures converges in total variation to the limiting probability measure P, and σ-algebras A and B are conditionally independent given H with respect to Pn for all n, then they are also conditionally independent with respect to the limiting measure P. As a corollary, this also extends to pointwise convergence of densities to a density.
In Gaussian graphical models, the likelihood equations must typically be solved iteratively. We investigate two algorithms: A version of iterative proportional scaling which avoids inversion of large matrices, and an algorithm based on convex duality and operating on the covariance matrix by neighbourhood coordinate descent, corresponding to the graphical lasso with zero penalty. For large, sparse graphs, the iterative proportional scaling algorithm appears feasible and has simple convergence properties. The algorithm based on neighbourhood coordinate descent is extremely fast and less dependent on sparsity, but needs a positive definite starting value to converge. We give an algorithm for finding such a starting value for graphs with low colouring number. As a consequence, we also obtain a simplified proof for existence of the maximum likelihood estimator in such cases.
The notion of multivariate total positivity has proved to be useful in finance and psychology but may be too restrictive in other applications. In this paper we propose a concept of local association, where highly connected components in a graphical model are positively associated and study its properties. Our main motivation comes from gene expression data, where graphical models have become a popular exploratory tool. The models are instances of what we term mixed convex exponential families and we show that a mixed dual likelihood estimator has simple exact properties for such families as well as asymptotic properties similar to the maximum likelihood estimator. We further relax the positivity assumption by penalizing negative partial correlations in what we term the positive graphical lasso. Finally, we develop a GOLAZO algorithm based on block-coordinate descent that applies to a number of optimization procedures that arise in the context of graphical models, including the estimation problems described above. We derive results on existence of the optimum for such problems.
Motivated by extreme value theory, max-linear Bayesian networks have been recently introduced and studied as an alternative to linear structural equation models. However, for max-linear systems the classical independence results for Bayesian networks are far from exhausting valid conditional independence statements. We use tropical linear algebra to derive a compact representation of the conditional distribution given a partial observation, and exploit this to obtain a complete description of all conditional independence relations. In the context-specific case, where conditional independence is queried relative to a specific value of the conditioning variables, we introduce the notion of a source DAG to disclose the valid conditional independence relations. In the context-free case we characterize conditional independence through a modified separation concept, $\ast$-separation, combined with a tropical eigenvalue condition. We also introduce the notion of an impact graph which describes how extreme events spread deterministically through the network and we give a complete characterization of such impact graphs. Our analysis opens up several interesting questions concerning conditional independence and tropical geometry.
We address the identifiablity and estimation of recursive max-linear structural equation models represented by an edge weighted directed acyclic graph (DAG). Such models are generally unidentifiable and we identify the whole class of DAGs and edge weights corresponding to a given observational distribution. For estimation, standard likelihood theory cannot be applied because the corresponding families of distributions are not dominated. Given the underlying DAG, we present an estimator for the class of edge weights and show that it can be considered a generalized maximum likelihood estimator. In addition, we develop a simple method for identifying the structures of the DAGs. With probability tending to one at an exponential rate with the number of observations, this method correctly identifies the class of DAGs and, similarly, exactly identifies the possible edge weights.
We study exponential families of distributions that are multivariate totally positive of order 2 (MTP2), show that these are convex exponential families, and derive conditions for existence of the MLE. Quadratic exponential familes of MTP2 distributions contain attractive Gaussian graphical models and ferromagnetic Ising models as special examples. We show that these are defined by intersecting the space of canonical parameters with a polyhedral cone whose faces correspond to conditional independence relations. Hence MTP2 serves as an implicit regularizer for quadratic exponential families and leads to sparsity in the estimated graphical model. We prove that the maximum likelihood estimator (MLE) in an MTP2 binary exponential family exists if and only if both of the sign patterns $(1,-1)$ and $(-1,1)$ are represented in the sample for every pair of variables; in particular, this implies that the MLE may exist with $n=d$ observations, in stark contrast to unrestricted binary exponential families where $2^d$ observations are required. Finally, we provide a novel and globally convergent algorithm for computing the MLE for MTP2 Ising models similar to iterative proportional scaling and apply it to the analysis of data from two psychological disorders.
This note attempts to understand graph limits as defined by Lovasz and Szegedy in terms of harmonic analysis on semigroups. This is done by representing probability distributions of random exchangeable graphs as mixtures of characters on the semigroup of unlabeled graphs with node-disjoint union, thereby providing an alternative derivation of de Finetti's theorem for random exchangeable graphs.
We analyze the problem of maximum likelihood estimation for Gaussian distributions that are multivariate totally positive of order two (MTP2). By exploiting connections to phylogenetics and single-linkage clustering, we give a simple proof that the maximum likelihood estimator (MLE) for such distributions exists based on at least 2 observations, irrespective of the underlying dimension. Slawski and Hein, who first proved this result, also provided empirical evidence showing that the MTP2 constraint serves as an implicit regularizer and leads to sparsity in the estimated inverse covariance matrix, determining what we name the ML graph. We show that we can find an upper bound for the ML graph by adding edges corresponding to correlations in excess of those explained by the maximum weight spanning forest of the correlation matrix. Moreover, we provide globally convergent coordinate descent algorithms for calculating the MLE under the MTP2 constraint which are structurally similar to iterative proportional scaling. We conclude the paper with a discussion of signed MTP2 distributions.
Graphical models (Lauritzen, 1996; Maathuis et al., 2019) are probabilistic or statistical models which describe complex relationships between systems of variables in a modular fashion, exploiting conditional independence relations encoded with mathematical graphs. The modularity enables simple specification, interpretation, communication, and computation associated with application of the models. The subject has reached maturity as a research area and graphical models are now applied in an abundance of contexts; for example in digital communication, machine learning, causal inference, genetics, decision support systems, social sciences, and forensic science. Modern applications of graphical models, as exemplified above, have disclosed a variety of theoretical and methodological challenges that must be addressed to take full advantage of their versatility and expressiveness. I am working on a number of different topics within this general area, including the development of graphical models for extremes, graphical models for social networks, but this little note will focus on my study of total positivity, mostly done in collaboration with Caroline Uhler (MIT) and Piotr Zwiernik (UPF Barcelona).
We derive representation theorems for exchangeable distributions on finite and infinite graphs using elementary arguments based on geometric and graph-theoretic concepts. Our results elucidate some of the key differences, and their implications, between statistical network models that are finitely exchangeable and models that define a consistent sequence of probability distributions on graphs of increasing size.
We study Bayesian networks based on max-linear structural equations as introduced in Gissibl and Kl\"uppelberg [16] and provide a summary of their independence properties. In particular we emphasize that distributions for such networks are generally not faithful to the independence model determined by their associated directed acyclic graph. In addition, we consider some of the basic issues of estimation and discuss generalized maximum likelihood estimation of the coefficients, using the concept of a generalized likelihood ratio for non-dominated families as introduced by Kiefer and Wolfowitz [21]. Finally we argue that the structure of a minimal network asymptotically can be identified completely from observational data.
Finn V. Jensen合作论文数Department of Computer Science, Aalborg University3
Bo Thiesson合作论文数Microsoft Research2