Paired comparison models, such as Bradley-Terry and Thurstone-Mosteller, are commonly used to estimate relative strengths of pairwise compared items in tournament-style data. We discuss estimation of paired comparison models with a ridge penalty. A new approach is derived which combines empirical Bayes and composite likelihoods without any need to re-fit the model, as a convenient alternative to cross-validation of the ridge tuning parameter. Simulation studies demonstrate much better predictive accuracy of the new approach relative to ordinary maximum likelihood. A widely used alternative, the application of a standard bias-reducing penalty, is also found to improve appreciably the performance of maximum likelihood; but the ridge penalty, with tuning as developed here, yields greater accuracy still. The methodology is illustrated through application to 28 seasons of English Premier League football.
The rating of items based on pairwise comparisons has been a topic of statistical investigation for many decades. Numerous approaches have been proposed. One of the best known is the Bradley-Terry model. This paper seeks to assemble and explain a variety of motivations for its use. Some are based on principles or on maximising an objective function; others are derived from well-known statistical models, or stylised game scenarios. They include both examples well-known in the literature as well as what are believed to be novel presentations.
The Research Excellence Framework (REF) is a periodic UK-wide assessment of the quality of published research in universities. The most recent REF was in 2014, and the next will be in 2021. The published results of REF2014 include a categorical `quality profile' for each unit of assessment (typically a university department), reporting what percentage of the unit's REF-submitted research outputs were assessed as being at each of four quality levels (labelled 4*, 3*, 2* and 1*). Also in the public domain are the original submissions made to REF2014, which include -- for each unit of assessment -- publication details of the REF-submitted research outputs. In this work, we address the question: to what extent can a REF quality profile for research outputs be attributed to the journals in which (most of) those outputs were published? The data are the published submissions and results from REF2014. The main statistical challenge comes from the fact that REF quality profiles are available only at the aggregated level of whole units of assessment: the REF panel's assessment of each individual research output is not made public. Our research question is thus an `ecological inference' problem, which demands special care in model formulation and methodology. The analysis is based on logit models in which journal-specific parameters are regularized via prior `pseudo-data'. We develop a lack-of-fit measure for the extent to which REF scores appear to depend on publication venues rather than research quality or institution-level differences. Results are presented for several research fields.
In current applied research the most-used route to an analysis of composition is through log-ratios -- that is, contrasts among log-transformed measurements. Here we argue instead for a more direct approach, using a statistical model for the arithmetic mean on the original scale of measurement. Central to the approach is a general variance-covariance function, derived by assuming multiplicative measurement error. Quasi-likelihood analysis of logit models for composition is then a general alternative to the use of multivariate linear models for log-ratio transformed measurements, and it has important advantages. These include robustness to secondary aspects of model specification, stability when there are zero-valued or near-zero measurements in the data, and more direct interpretation. The usual efficiency property of quasi-likelihood estimation applies even when the error covariance matrix is unspecified. We also indicate how the derived variance-covariance function can be used, instead of the variance-covariance matrix of log-ratios, with more general multivariate methods for the analysis of composition. A specific feature is that the notion of `null correlation' -- for compositional measurements on their original scale -- emerges naturally.
We propose a hazard model for entry into marriage, based on a bell-shaped function to model the dependence on age. We demonstrate near-aliasing in an extension that estimates the support of the hazard and mitigate this via re-parameterization. Our proposed model parameterizes the maximum hazard and corresponding age, thereby facilitating more general models where these features depend on covariates. For data on women's marriages from the Living in Ireland Surveys 1994–2001, this approach captures a reduced propensity to marry over successive cohorts and an increasing delay in the timing of marriage with increasing education.
Frequently in sporting competitions it is desirable to compare teams based on records of varying schedule strength. Methods have been developed for sports where the result outcomes are win, draw, or loss. In this paper those ideas are extended to account for any finite multiple outcome result set. A principle-based motivation is supplied and an implementation presented for modern rugby union, where bonus points are awarded for losing within a certain score margin and for scoring a certain number of tries. A number of variants are discussed including the constraining assumptions that are implied by each. The model is applied to assess the current rules of the Daily Mail Trophy, a national schools tournament in England and Wales.
Penalization of the likelihood by Jeffreys' invariant prior, or by a positive power thereof, is shown to produce finite-valued maximum penalized likelihood estimates in a broad class of binomial generalized linear models. The class of models includes logistic regression, where the Jeffreys-prior penalty is known additionally to reduce the asymptotic bias of the maximum likelihood estimator; and also models with other commonly used link functions such as probit and log-log. Shrinkage towards equiprobability across observations, relative to the maximum likelihood estimator, is established theoretically and is studied through illustrative examples. Some implications of finiteness and shrinkage for inference are discussed, particularly when inference is based on Wald-type procedures. A widely applicable procedure is developed for computation of maximum penalized likelihood estimates, by using repeated maximum likelihood fits with iteratively adjusted binomial responses and totals. These theoretical results and methods underpin the increasingly widespread use of reduced-bias and similarly penalized binomial regression models in many applied fields.
Adaptive importance samplers are adaptive Monte Carlo algorithms to estimate expectations with respect to some target distribution which \textit{adapt} themselves to obtain better estimators over a sequence of iterations. Although it is straightforward to show that they have the same $\mathcal{O}(1/\sqrt{N})$ convergence rate as standard importance samplers, where $N$ is the number of Monte Carlo samples, the behaviour of adaptive importance samplers over the number of iterations has been left relatively unexplored. In this work, we investigate an adaptation strategy based on convex optimisation which leads to a class of adaptive importance samplers termed \textit{optimised adaptive importance samplers} (OAIS). These samplers rely on the iterative minimisation of the $\chi^2$-divergence between an exponential-family proposal and the target. The analysed algorithms are closely related to the class of adaptive importance samplers which minimise the variance of the weight function. We first prove non-asymptotic error bounds for the mean squared errors (MSEs) of these algorithms, which explicitly depend on the number of iterations and the number of samples together. The non-asymptotic bounds derived in this paper imply that when the target belongs to the exponential family, the $L_2$ errors of the optimised samplers converge to the optimal rate of $\mathcal{O}(1/\sqrt{N})$ and the rate of convergence in the number of iterations are explicitly provided. When the target does not belong to the exponential family, the rate of convergence is the same but the asymptotic $L_2$ error increases by a factor $\sqrt{\rho^\star} > 1$, where $\rho^\star - 1$ is the minimum $\chi^2$-divergence between the target and an exponential-family proposal.
This paper presents the R package PlackettLuce , which implements a generalization of the Plackett–Luce model for rankings data. The generalization accommodates both ties (of arbitrary order) and partial rankings (complete rankings of subsets of items). By default, the implementation adds a set of pseudo-comparisons with a hypothetical item, ensuring that the underlying network of wins and losses between items is always strongly connected. In this way, the worth of each item always has a finite maximum likelihood estimate, with finite standard error. The use of pseudo-comparisons also has a regularization effect, shrinking the estimated parameters towards equal item worth. In addition to standard methods for model summary, PlackettLuce provides a method to compute quasi standard errors for the item parameters. This provides the basis for comparison intervals that do not change with the choice of identifiability constraint placed on the item parameters. Finally, the package provides a method for model-based partitioning using covariates whose values vary between rankings, enabling the identification of subgroups of judges or settings with different item worths. The features of the package are demonstrated through application to classic and novel data sets.
This paper introduces a natural extension of the pair-comparison-with-ties model of Davidson (1970, J. Amer. Statist. Assoc.), to allow for ties when more than two items are compared. Properties of the new model are discussed. It is found that this "Davidson-Luce" model retains the many appealing features of Davidson's solution, while extending the scope of application substantially beyond the domain of pair-comparison data. The model introduced here already underpins the handling of tied rankings in the "PlackettLuce" R package.
Rankings of scholarly journals based on citation data are often met with scepticism by the scientific community. Part of the scepticism is due to disparity between the common perception of journals’ prestige and their ranking based on citation counts. A more serious concern is the inappropriate use of journal rankings to evaluate the scientific influence of researchers. The paper focuses on analysis of the table of cross-citations among a selection of statistics journals. Data are collected from the Web of Science database published by Thomson Reuters. Our results suggest that modelling the exchange of citations between journals is useful to highlight the most prestigious journals, but also that journal citation data are characterized by considerable heterogeneity, which needs to be properly summarized. Inferential conclusions require care to avoid potential overinterpretation of insignificant differences between journal ratings. Comparison with published ratings of institutions from the UK's research assessment exercise shows strong correlation at aggregate level between assessed research quality and journal citation ‘export scores’ within the discipline of statistics.
The term ‘quasi-likelihood’ has been used recently to describe a fairly wide variety of techniques for estimation and inference. This paper focuses on the method introduced by Wedderburn (1974), further developed by McCullagh (1983), but which, as pointed out by Crowder (1987), has roots at least as far back as Williams (1959, §4.5). A brief introduction and overview are given by McCullagh (1986), and more comprehensive treatments with examples of application may be found in McCullagh & Nelder (1989) and McCullagh (1991). The quasi-likelihood method of estimation is probably best viewed as a straightforward extension of generalized least squares. Suppose that y is a n × 1 response vector, assumed to be a realization of a random vector Y with
Rankings of scholarly journals based on citation data are often met with skepticism by the scientific community. Part of the skepticism is due to disparity between the common perception of journals' prestige and their ranking based on citation counts. A more serious concern is the inappropriate use of journal rankings to evaluate the scientific influence of authors. This paper focuses on analysis of the table of cross-citations among a selection of Statistics journals. Data are collected from the Web of Science database published by Thomson Reuters. Our results suggest that modelling the exchange of citations between journals is useful to highlight the most prestigious journals, but also that journal citation data are characterized by considerable heterogeneity, which needs to be properly summarized. Inferential conclusions require care in order to avoid potential over-interpretation of insignificant differences between journal ratings. Comparison with published ratings of institutions from the UK's Research Assessment Exercise shows strong correlation at aggregate level between assessed research quality and journal citation `export scores' within the discipline of Statistics.
Should the IS field relinquish its identity and surrender to the social, political and economic forces that are placing IS scholars outside of traditional IS departments and programs? Should IS scholars mesh into competing departments (e.g., marketing and accounting) and schools (e.g., information studies and engineering) instead of seeking placement in a traditional IS department? By touting IS as the blood that runs through all business functions, would we serve our stakeholders better by embracing our status as a field without a real home and becoming homeless? Is the key to the IS field's survival ripping IS departments apart, keeping them together or following a hybrid approach? This panel session brings together eminent IS scholars to argue both for and against the case that the IS field should surrender to the destructive forces that are pulling it in every direction.
Summary. In the course of national sports tournaments, usually lasting several months, it is expected that the abilities of teams taking part in the tournament will change over time. A dynamic extension of the Bradley–Terry model for paired comparison data is introduced to model the outcomes of sporting contests, allowing for time varying abilities. It is assumed that teams’ home and away abilities depend on past results through exponentially weighted moving average processes. The model proposed is applied to sports data with and without tied contests, namely the 2009–2010 regular season of the National Basketball Association tournament and the 2008–2009 Italian Serie A football season.
SummaryIn the course of national sports tournaments, usually lasting several months, it is expected that the abilities of teams taking part in the tournament will change over time. A dynamic extension of the Bradley–Terry model for paired comparison data is introduced to model the outcomes of sporting contests, allowing for time varying abilities. It is assumed that teams’ home and away abilities depend on past results through exponentially weighted moving average processes. The model proposed is applied to sports data with and without tied contests, namely the 2009–2010 regular season of the National Basketball Association tournament and the 2008–2009 Italian Serie A football season.
Composite likelihood methods are extensions of the Fisherian likelihood theory, one of the most influential approaches in statistics. Such extensions are generally motivated by the issue of computational feasibility arising in the application of the likelihood method in high-dimensional data analysis. Complex dependence presents substantial challenges in statistical modelling and methods and in substantive applications. The idea of projecting high-dimensional complicated likelihood functions to low-dimensional computationally feasible likelihood objects is methodologically appealing. Composite likelihood inherits many of the good properties of inference based on the full likelihood function, but is more easily implemented with high-dimensional data sets. This methodology is, to some extent, an alternative to the Markov Chain Monte Carlo method, and its impact is unbounded. The literature on both theoretical and practical issues for inference based on composite likelihood continues to expand quickly; the field of extremal processes for spatial data, of particular importance for climate modelling, is one of the most recent examples of an area where composite likelihood inference is both practical and efficient. The first international workshop on composite likelihood methods was held at the University of Warwick in April 2008. It attracted participants from all over the world and was widely viewed as very successful. Following the workshop, a special issue of the journal Statistica Sinica devoted to composite likelihood was announced; it was published in January 2011. This issue includes two long overview papers, one of which is devoted to applications in statistical genetics; several papers developing new theory for inference based on composite likelihood; new results in the application of composite likelihood to time series, spatial processes, longitudinal data and missing data. The methodology has drawn considerable attention in a broad range of applied disciplines in which complex data structures arise. Some notable application areas include, statistical genetics, genetic epidemiology, finance, panel surveys, computer experiments, geostatistics and biostatistics.