We introduce a new approach to goodness-of-fit testing in the high dimensional, sparse extended multinomial context. The paper takes a computational information geometric approach, extending classical higher order asymptotic theory. We show why the Wald – equivalently, the Pearson χ and score statistics – are unworkable in this context, but that the deviance has a simple, accurate and tractable sampling distribution even for moderate sample sizes. Issues of uniformity of asymptotic approximations across model space are discussed. A variety of important applications and extensions are noted.
This paper applies the tools of computation information geometry [3] – in particular, high dimensional extended multinomial families as proxies for the ‘space of all distributions’ – in the inferentially demanding area of statistical mixture modelling. A range of resultant benefits are noted.
This book focuses on the application and development of information geometric methods in the analysis, classification and retrieval of images and signals. It provides introductory chapters to help tho
We show how information geometry throws new light on the interplay between goodness-of-fit and estimation, a fundamental issue in statistical inference. A geometric analysis of simple, yet representative, models involving the same population parameter compellingly establishes the main theme of the paper: namely, that goodness-of-fit is necessary but not sufficient for model selection. Visual examples vividly communicate this. Specifically, for a given estimation problem, we define a class of least-informative models, linking these to both nonparametric and maximum entropy methods. Any other model is then seen to involve an informative rotation, often embodying extra-data considerations. We also look at the way that translation of models generates a form of bias-variance trade-off. Overall, our approach is a global extension of pioneering local work by Copas and Eguchi which, we note, was also geometrically inspired.
A novel approach is introduced to a very widely occurring problem, providing a complete, explicit resolution of it: minimisation of a convex quadratic under a general quadratic, equality or inequality, constraint. Completeness comes via identification of a set of mutually exclusive and exhaustive special cases. Explicitness, via algebraic expressions for each solution set. Throughout, underlying geometry illuminates and informs algebraic development. In particular, centrally to this new approach, affine equivalence is exploited to re-express the same problem in simpler coordinate systems. Overall, the analysis presented provides insight into the diverse forms taken both by the problem itself and its solution set, showing how each may be intrinsically unstable. Comparisons of this global, analytic approach with the, intrinsically complementary, local, computational approach of (generalised) trust region methods point to potential synergies between them. Points of contact with simultaneous diagonalisation results are noted.
The Information Geometry of extended exponential families has received much recent attention in a variety of important applications, notably categorical data analysis, graphical modelling and, more specifically, log-linear modelling. The essential geometry here comes from the closure of an exponential family in a high-dimensional simplex. In parallel, there has been a great deal of interest in the purely Fisher Riemannian structure of (extended) exponential families, most especially in the Markov chain Monte Carlo literature. These parallel developments raise challenges, addressed here, at a variety of levels: both theoretical and practical—relatedly, conceptual and methodological. Centrally to this endeavour, this paper makes explicit the underlying geometry of these two areas via an analysis of the limiting behaviour of the fundamental geodesics of Information Geometry, these being Amari’s (+1) and (0)-geodesics, respectively. Overall, a substantially more complete account of the Information Geometry of extended exponential families is provided than has hitherto been the case. We illustrate the importance and benefits of this novel formulation through applications.
In statistical practice model building, sensitivity and uncertainty are major concerns of the analyst. This paper looks at these issues from an information geometric point of view. Here, we define sensitivity to mean understanding how inference about a problem of interest changes with perturbations of the model. In particular it is an example of what we call computational information geometry. The embedding of simple models in much larger information geometric spaces is shown to illuminate these critically important issues.
This book focuses on the application and development of information geometric methods in the analysis, classification and retrieval of images and signals. It provides introductory chapters to help those new to information geometry and applies the theory to several applications. This area has developed rapidly over recent years, propelled by the major theoretical developments in information geometry, efficient data and image acquisition and the desire to process and interpret large databases of digital information. The book addresses both the transfer of methodology to practitioners involved in database analysis and in its efficient computational implementation.
Al-talib, M. See under Bhattacharaya, B. Atchadé, Y. See under Roy, S. Bandyopadhyay, S. and Subba Rao, S. A test for stationarity for irregularly spaced spatial data 95 Barber, R. F. See under Janson, L. Barber, R. F. and Ramdas, A. The p-fi lter: multilayer false discovery rate control for grouped hypotheses 1247 Bastide, P., Mariadassou, M. and Robin, S. Detection of adaptive shifts on phylogenies by using shifted stochastic processes on a tree 1067 Belloni, A., Rosenbaum, M. and Tsybakov, A. B. Linear and conic programming estimators in high dimensional errors-in-variables models 939 Bhattacharaya, B. and Al-talib, M. A minimum relative entropy based correlation model between the response and covariates 1095 Birr, S., Volgushev, S., Kley, T., Dette, H. and Hallin, M. Quantile spectral analysis for locally stationary time series 1619 Botev, Z. I. The normal law under linear restrictions: simulation and estimation via minimax tilting 125 Bradley, J. R., Wikle, C. K. and Holan, S. H. Regionalization of multiscale spatial processes by using a criterion for spatial aggregation error 815 Braekers, R. See under Prenen, L. Brockwell, P. J. and Matsuda, Y. Continuous auto-regressive moving average random fi elds on n 833 Brunner, E., Konietscke, F., Pauly, M. and Puri, M. L. Rank-based procedures in factorial designs: hypotheses about non-parametric treatment effects 1463 Cai, T. T. and Sun, W. Optimal screening and discovery of sparse signals with applications to multistage high throughput studies 197 Campi, M. C. See under Carè, A. Candès, E. See under Janson, L. Cannings, T. I. and Samworth, R. J. Random projection ensemble classifi cation (with discussion) 959 Contributors to the discussion: A. Ahmad, 1004; L. Anderlucci, 1000; B. Barney, 1007; W. Bergsma, 1004; X. Bing, 1006; R. Blaser, 1007; T. I. Cannings, 1027; M. de Carvalho, 1007; R. Casarin, 1008; Y. Chen, 1003; C. Cheng, 1016; P. Contreras, 1020; F. Critchley, 1001; L. Dalla Valle, 1023; E. Demirkaya, 1008; J. Derenski, 1009; R. J. Durrant, 1010; J. Fan, 1010; Y. Fan, 1009; Y. Feng, 1011; F. Fortunato, 1000, 1001; L. Frattarolo, 1008; P. Fryzlewicz, 1007; M. P. B. Gallaugher, 1011; M. Gataric, 1012; T. Gneiting, 1013; D. Hand, 997; C. Hennig, 996; G. M. James, 1009; H. Jamil, 1004; L. Janson, 1013; J. T. Kent, 1002; D. Kong, 1014; C. Leng, 1026; S. Lerch, 1013; B. Li, 1014; J. J. Li, 1025; Y. Ling, 1015; M. Liu, 1016; W. Lu, 1021; X. Lu, 1019; J. Lv, 1008; P. Marriott, 1020; J. Mateu, 1019; P. D. McNicholas, 1011; A. Montanari, 1000; F. Murtagh, 1020; G. L. Page, 1007; L. Rossini, 1008; R. Sabolová, 1020; R. J. Samworth, 1027; R. D. Shah, 1003; C. Shi, 1021; S. J. Shin, 1021; R. Song, 1021; J. Stander, 1023; M. Stehlík, 1023; L. Střelec, 1023; P. Switzer, 1024; M. Thulin, 1024; J. H. Tomal, 1024; H. Tong, 1025; X. Tong, 1025; C. Viroli, 996; X. Wang, 1026; M. Wegkamp, 1006; W. J. Welch, 1024; Y. Wu, 1021; J.-H. Xue, 1015, 1019; X. Yang, 1015; Y. G. Yatracos, 1026; K. Yu, 1014; R. H. Zamar, 1024; W. Zhang, 1000; C. Zheng, 1021; Z. Zhu, 1010 Carè, A., Garatti, S. and Campi, M. C. A coverage theory for least squares 1367 Caron, F. and Fox, E. B. Sparse graphs using exchangeable random measures (with discussion) 1295 Contributors to the discussion: J. Arbel, 1342; S. Banerjee, 1343; M. Battiston, 1343; K. Bharath, 1340; G. Bianconi, 1338; B. BloemReddy, 1341; C. Borgs, 1344; A. Bouchard-Côté, 1344; C. Briercliffe, 1344; T. Broderick, 1345; T. Campbell, 1345; F. Caron, 1360; R. Casarin, 1345; I. Castillo, 1347; S. Chakraborty, 1348; J. T. Chayes, 1344; H. Crane, 1349; D. Durante, 1350; M. Eckardt, 1353; S. Favaro, 1343; D. Firth, 1354; E. B. Fox, 1360; C. Gao, 1350; S. Ghosal, 1343; J. E. Griffi n, 1351; K. Heaukulani, 1351; M. Iacopini, 1345; L. F. James, 1351; S. Janson, 1352; W. S. Kendall, 1341; K. Kumar, 1352; F. Leisen, Index of authors, volume 79, 2017 J. R. Statist. Soc. B (2017) 79, Part 5, pp. 1667–1671
We give a personal view of what Information Geometry is, and what it is becoming, by exploring a number of key topics: dual affine families, boundaries, divergences, tensorial structures, and dimensionality. For each, we start with a graphical illustrative example (Sect. 1.1), give an overview of the relevant theory and key references (Sect. 1.2), and finish with a number of applications of the theory (Sect. 1.3). We treat ‘Information Geometry’ as an evolutionary term, deliberately not attempting a comprehensive definition. Rather, we illustrate how both the geometries used and application areas are rapidly developing.
Recent progress using geometry in the design of efficient Markov chain Monte Carlo (MCMC) algorithms have shown the effectiveness of the Fisher Riemannian structure. Furthermore, the theory of the underlying geometry of spaces of statistical models has made an important breakthrough by extending the classical theory on exponential families to their closures, the so-called extended exponential families. This paper looks at the underlying geometry of the Fisher information, in particular its limiting behaviour near boundaries, which illuminates the excellent behaviour of the corresponding geometric MCMC algorithms. Further, the paper shows how Fisher geodesics in extended exponential families smoothly attach the boundaries of extended exponential families to their relative interior. We conjecture that this behaviour could be exploited for trans-dimensional MCMC algorithms.
This paper applies the tools of computation information geometry [3] – in particular, high dimensional extended multinomial families as proxies for the ‘space of all distributions’ – in the inferentially demanding area of statistical mixture modelling. A range of resultant benefits are noted.
This paper lays the foundations for a new framework for numerically and computationally applying information geometric methods to statistical modelling.
We introduce a new approach to goodness-of-fit testing in the high dimensional, sparse extended multinomial context. The paper takes a computational information geometric approach, extending classical higher order asymptotic theory. We show why the Wald – equivalently, the Pearson χ ^2 and score statistics – are unworkable in this context, but that the deviance has a simple, accurate and tractable sampling distribution even for moderate sample sizes. Issues of uniformity of asymptotic approximations across model space are discussed. A variety of important applications and extensions are noted.
This paper takes an information-geometric approach to the challenging issue of goodness-of-fit testing in the high dimensional, low sample size context where-potentially-boundary effects dominate. The main contributions of this paper are threefold: first, we present and prove two new theorems on the behaviour of commonly used test statistics in this context; second, we investigate-in the novel environment of the extended multinomial model-the links between information geometry-based divergences and standard goodness-of-fit statistics, allowing us to formalise relationships which have been missing in the literature; finally, we use simulation studies to validate and illustrate our theoretical results and to explore currently open research questions about the way that discretisation effects can dominate sampling distributions near the boundary. Novelly accommodating these discretisation effects contrasts sharply with the essentially continuous approach of skewness and other corrections flowing from standard higher-order asymptotic analysis.
Khare, K., Oh, S.-Y. and Rajaratnam, B. A convex pseudolikelihood framework for high dimensional partial correlation estimation with convergence guarantees 803 Kidziński, Ł. See under Hörmann, S. Konietschke, F. See under Pauly, M. Kraus, D. Components and completion of partially observed functional data 777 Kreiss, J.-P. and Paparoditis, E. Bootstrapping locally stationary processes 267 Künsch, HR See under Sigrist, F.
An analytical solution is provided to the problem of minimising a convex quadratic under a general quadratic (inequality) constraint, additional affine constraints being easily accommodated. Such problems occur widely. The explicit, analytical solution provided offers insight into the diverse nature both of special cases of this problem and of their solution sets. This sheds light on algorithm performance and design, intrinsically unstable problems being a particular focus. Points of contact with simultaneous diagonalisation results are noted.
Computational Information Geometry... ...in mixture modelling Computational Information Geometry: mixture modelling Germain Van Bever1 , R. Sabolova1 , F. Critchley1 & P. Marriott2 . 1 The Open University (EPSRC grant EP/L010429/1), United Kingdom 2 University of Waterloo, USA GSI15, 28-30 October 2015, Paris Germain Van Bever CIG for mixtures 1/19 Computational Information Geometry... ...in mixture modelling Outline 1 Computational Information Geometry... Information Geometry CIG 2 ...in mixture modelling Introduction Lindsay’s convex geometry (C)IG for mixture distributions Germain Van Bever CIG for mixtures 2/19 Computational Information Geometry... ...in mixture modelling Information Geometry CIG Outline 1 Computational Information Geometry... Information Geometry CIG 2 ...in mixture modelling Introduction Lindsay’s convex geometry (C)IG for mixture distributions Germain Van Bever CIG for mixtures 3/19 Computational Information Geometry... ...in mixture modelling Information Geometry CIG Generalities The use of geometry in statistics gave birth to many different approaches. Traditionally, Information geometry refers to the application of differential geometry to statistica
A broad view of the nature and potential of computational information geometry in statistics is offered. This new area suitably extends the manifold-based approach of classical information geometry to a simplicial setting, in order to obtain an operational universal model space. Additional underlying theory and illustrative real examples are presented. In the infinite-dimensional case, challenges inherent in this ambitious overall agenda are highlighted and promising new methodologies indicated.