We discuss the evolution in the presentation of statistical science in English‐language textbooks, focusing on the period 1900–1970 as the field became increasingly influenced by research contributions of R. A. Fisher and Jerzy Neyman. George Udny Yule authored an early popular book that had 14 editions. Methods books authored by Fisher and George Snedecor guided scientists in implementing modern statistical methods. In the World War 2 era, textbooks authored by Maurice Kendall, Samuel Wilks, and Harald Cramér presented a dramatically different “mathematical statistics” portrayal that centered on theoretical foundations. The textbook emergence of the Bayesian approach occurred later, influenced by books by Harold Jeffreys and Leonard J. Savage. The quarter century after World War 2 saw an explosion of books in mathematical statistics and in particular topic areas. In addition to his highly cited research contributions, Sir David Cox was a prolific author of books on a great variety of topics. Most were published after the 1900–1970 period considered in this article, but we also summarize them as part of this special issue to honor his memory. We conclude by discussing the future of textbooks on the foundations of statistical science in the emerging, ever‐broader, era of data science.
Logistic regression is the closest model, given its sufficient statistics, to the model of constant success probability in terms of Kullback–Leibler information. A generalized binary model has this property for the more general ϕ-divergence. These results generalize to multinomial and other discrete data.
We describe two interesting and innovative strands of Murray Aitkin's research publications, dealing with mixture models and with Bayesian inference. Of his considerable publications on mixture models, we focus on a nonparametric random effects approach in generalized linear mixed modelling, which has proven useful in a wide variety of applications. As an early proponent of ways of implementing the Bayesian paradigm, Aitkin proposed an alternative Bayes factor based on a posterior mean likelihood. We discuss these innovative approaches and some research lines motivated by them and also suggest future related methodological implementations.
One of C. R. Rao’s many important contributions to statistical science was his introduction of the score test , based on the derivative of the log-likelihood function at the null hypothesis value of the parameter of interest. This article reviews methods for constructing score tests and score-test-based confidence intervals for analyzing parameters that arise in analyzing categorical data. A considerable literature indicates that score tests and their inversion for constructing confidence intervals perform well in a variety of settings and sometimes much better than Wald-test and likelihood-ratio test-based methods. We also discuss extensions of score-based inference and potential future research on generalizations for longitudinal data, complex sampling, and high-dimensional data.
Foundations of Statistics for Data Scientists: With R and Python is designed as a textbook for a one- or two-term introduction to mathematical statistics for students training to become data scientists. It is an in-depth presentation of the topics in statistical science with which any data scientist should be familiar, including probability distributions, descriptive and inferential statistical methods, and linear modeling. The book assumes knowledge of basic calculus, so the presentation can focus on "why it works" as well as "how to do it." Compared to traditional "mathematical statistics" textbooks, however, the book has less emphasis on probability theory and more emphasis on using software to implement statistical methods and to conduct simulations to illustrate key concepts. All statistical analyses in the book use R software, with an appendix showing the same analyses with Python. Key Features: Shows the elements of statistical science that are important for students who plan to become data scientists. Includes Bayesian and regularized fitting of models (e.g., showing an example using the lasso), classification and clustering, and implementing methods with modern software (R and Python). Contains nearly 500 exercises. The book also introduces modern topics that do not normally appear in mathematical statistics texts but are highly relevant for data scientists, such as Bayesian inference, generalized linear models for non-normal responses (e.g., logistic regression and Poisson loglinear models), and regularized model fitting. The nearly 500 exercises are grouped into "Data Analysis and Applications" and "Methods and Concepts." Appendices introduce R and Python and contain solutions for odd-numbered exercises. The book's website (http://stat4ds.rwth-aachen.de/) has expanded R, Python, and Matlab appendices and all data sets from the examples and exercises.
This chapter contains sections titled: Components of a Generalized Linear Model Generalized Linear Models for Binary Data Generalized Linear Models for Count Data Statistical Inference and Model Checking Fitting Generalized Linear Models Problems
We discuss how the foundations of statistical science have been presented historically in textbooks, with focus on the first half of the twentieth century after the field had become better defined by advances due to Francis Galton, Karl Pearson, and R. A. Fisher. Our main emphasis is on books that presented the theory underlying the subject, often identified as mathematical statistics, with primary focus on books authored by G. Udny Yule, Maurice Kendall, Samuel Wilks, and Harald Cramer. We also discuss influential books on statistical methods by R. A. Fisher and George Snedecor that showed scientists how to implement the theory. We then survey textbooks published in the quarter century after World War 2, as Statistics gathered more visibility as an academic subject and Departments of Statistics were formed at many universities. We also summarize how Bayesian presentations of Statistics emerged. In each section, we describe how the books were evaluated in reviews shortly after their publications. We conclude by briefly discussing the recent past, the present, and the future of textbooks on the foundations of statistical science and include comments about this by several notable statisticians.