Abstract This paper considers the tactic of pressing in soccer. The investigation defines the conditions that constitute a traditional press, which then permits the identification of pressing situations using tracking data. The data consist of matches from the 2019 season of the Chinese Super League. Descriptive analyses suggest that traditional pressing is associated with more turnovers when the pressed player is given less time on the ball and as the intensity of the press increases (i.e. more pressers). It is also observed that counter-pressing is associated with better outcomes than a traditional press. Statistical modeling provides further insights with respect to pressing. For example, it is observed that pressing is associated with better outcomes when it is initiated deep and central in the opponent’s end of the field.
Event history data from sports competitions have recently drawn increasing attention in sports analytics to generate data-driven strategies. Such data often exhibit self-excitation in the event occurrence and dependence within event clusters. The conventional event models based on gap times may struggle to capture those features. In particular, while consecutive events may occur within a short timeframe, the self-excitation effect caused by previous events is often transient and continues for a period of uncertain time. This paper introduces an extended Hawkes process model with random self-excitation duration to formulate the dynamics of event occurrence. We present examples of the proposed model and procedures for estimating the associated model parameters. We employ the collection of the corner kicks in the games of the 2019 regular season of the Chinese Super League to motivate and illustrate the modeling and its usefulness. We also design algorithms for simulating the event process under proposed models. The proposed approach can be adapted with little modification in many other research fields such as Criminology and Infectious Disease.
Corner kicks are an important event in soccer because they are often the result of strong attacking play and can be of keen interest to sports fans and bettors. Peng, Hu, and Swartz (2024, Computational Statistics) frame the commonly available corner kick data as right-censored event times, formulate the mixture feature of corner kick times caused by previous corner kicks, and explore patterns of corner kicks associated with several factors. This paper extends their modeling to accommodate the potential correlations between corner kicks by the same teams within the same games. We consider a frailty model for event times and apply the Monte Carlo Expectation Maximization (MCEM) algorithm to obtain the maximum likelihood estimates for the model parameters. We compare the proposed model with the model in Peng, Hu, and Swartz (2024) using likelihood ratio tests. The 2019 Chinese Super League (CSL) data are employed throughout the paper for motivation and illustration.
Causal inference has become an accepted analytic framework in sports analytics, where experimentation is rarely feasible. A key consideration is the choice of estimand, specifically, whether to target the Average Treatment Effect (ATE), which reflects the effect of an action across the entire population, or the Average Treatment Effect on the Treated (ATT), which reflects the effect among those who actually took the action. Using data from nearly all 240 matches of the 2019 Chinese Super League season, we apply propensity score matching to estimate the causal effect of crossing on shot creation in soccer. The ATE and ATT are nearly identical (0.033 and 0.035 respectively), a result we attribute to substantial overlap in propensity score distributions between plays where a cross was and was not attempted. To illustrate when these estimands diverge, we construct two simulation scenarios with known ground truth: one reproducing the high-overlap structure of the real data, where ATE and ATT coincide, and one engineered to exhibit severe confounding and low overlap, where they diverge substantially. While empirical findings are specific to the 2019 Chinese Super League season, the case study and simulations provide a principled guide to estimand choice in causal analyses of sports data.
Evaluation of player value in sport can be measured in several ways. These measures, when captured over an entire career, provide insights concerning player contributions. Professional sports teams select young talent through a draft process with the goal of acquiring a player that will provide maximum value, but these expectations diminish as the pool of players grows smaller. In this paper, we develop valuation measures for draft picks in the National Hockey League (NHL) and analyze the value of each pick number with these measures. Specifically, we use different measures of player value to provide an expected value of that measure for each pick number in the draft. Our approach uses functional data analysis (FDA) to find a mean value curve from many observed functions in a nonparametric fashion. These functions are defined by each separate year of draft data. The resulting FDA model follows the assumption of monotonicity, ensuring that a smaller pick number always provides more expected value than any larger pick number. Based on a cross-validation approach, measuring value on annual salary provides the best predictive results. The proposed approach can be extended to sports in which an entry draft occurs and player career data are available.
This paper investigates unforced errors in tennis which is facilitated by a rich dataset. Descriptive analyses are carried out which studies the distribution of unforced error rates across professional tennis players, the identification of players with high and low rates, the relationship between rates and match winning percentage, the relationship between rates and aggressiveness (via hitting winners), the relationship between rates and touch number, and the study of rates versus time. Methods are then developed to assess the impact of unforced errors which are applicable to any racquet sport. We demonstrate the approach in the context of professional tennis with an investigation of the longstanding rivalry between Roger Federer and Rafael Nadal. The value of the approach is that we can provide estimates of the points lost, games lost, sets lost and matches lost due to unforced errors. The methods are based on a bootstrapping procedure which also yields standard errors for the estimates. The approach is valuable in terms of player evaluation, and can also be used for training purposes where it is possible to assess the quantification of improvement based on fewer unforced errors.
When considering future performance in sport, age is an important feature for prediction models. On average, players tend to improve from their rookie (earliest) season, plateau, and then decline in performance until they retire from the league. In this paper we apply Functional Principal Component Analysis to the careers of players from the National Hockey League in order to construct individual aging curves. The approach is nonparametric in the sense that a parametric structure is not imposed on the aging curves. A main aspect of our work is the consideration of selection bias whereby players who have long careers are not randomly sampled but tend to be exceptional players. Whereas the literature constructs aging curves that represent the average player, we produce aging curves for individual players; this is particularly useful in roster construction.
With the availability of tracking data, the determination of pitch control (field ownership) is an increasingly important topic in sports analytics. This paper reviews various approaches for the determination of pitch control and introduces a new field ownership metric that takes into account associated sporting dynamics. The methods that are proposed utilize the movement of the ball and players. Specifically, physical characteristics such as current velocity, acceleration and maximum velocity are considered. The determination of pitch control is based on the time that it takes the ball and the players to reach a given location. The main result of our investigation concerns the validation of the resultant pitch control diagram. Based on a sample of 5887 passes, the team identified as having pitch control was the observed recipient of the pass with 91
To understand the patterns of times to corner kicks in soccer and how they are associated with a few important factors, we analyze the corner kick records from the 2019 regular season of the Chinese Super League. This paper is particularly concerned with the elapsed time to a corner kick from a natural starting point. We overcome 2 challenges arising from such time-to-event analyses, which have not been discussed in the sports analytics literature. The first is that observations of times to corner kicks are subject to right-censoring. A given soccer starting point rarely ends with a corner kick but the occurrence of a different terminal event. The second issue is the mixture feature of short and typical gap times to the next corner kick from a particular one. There is often a subsequent corner kick quickly following a corner kick. The conventional event time models are thus inappropriate for formulating distributions of corner kick times. Our analysis reveals how the timing of corner kicks is associated with the factors of first versus second half of the game, home versus away team, score differential, betting odds prior to the game, and red card differential. We present applications of the developed statistical model for prediction to support tactics and sports betting.
The National Hockey League (NHL) Entry Draft has been an active area of research in hockey analytics over the past decade. Prior research has explored predictive modelling for draft results using player information and statistics as well as ranking data from draft experts. In this paper, we develop a new modelling framework for this problem using a Bayesian rank-ordered logit model based on draft ranking data from industry experts between 2019 and 2022. This model builds upon previous approaches by incorporating team tendencies, addressing within-ranking dependence between players, and solving various other challenges of working with rank-ordered outcomes, such as incorporating both unranked players and rankings that only consider a subset of the available pool of players (e.g., North American skaters, European goalies, etc.).
This paper attempts to identify football players who have a similar style to a player of interest. Playing style is not adequately quantified with traditional statistics, and therefore style statistics are created using tracking data. Tracking data allow us to monitor players throughout a match, and therefore include both "on-the-ball" and "off-the-ball" observations. Having developed style features, tractable discrepancy measures are introduced that are based on Kullback-Leibler divergence in the context of multivariate normal distributions. Examples are provided where a pool of players from the Chinese Super League are identified as having a playing style that is similar to players of interest.
The Banister impulse-response (IR) model was designed to predict an athlete's performance ability from their past training. Despite its long history, the model's usefulness remains limited due to difficulties in obtaining precise parameter estimates and performance predictions. To help address these challenges, we developed a Bayesian implementation of the IR model, which formalises the combined use of prior knowledge and data. We report the following methodological contributions: 1) we reformulated the model to facilitate the specification of informative priors, 2) we derived the IR model in Bayesian terms, and 3) we developed a method that enabled the JAGS software to be used while enforcing parameter constraints. To demonstrate proof-of-principle, we applied the model to the data of a national-class middle-distance runner. We specified the priors from published values of IR model parameters, followed by estimating the posterior distributions from the priors and the athlete's data. The Bayesian approach led to more precise and plausible parameter estimates than nonlinear least squares. We conclude that the Bayesian implementation of the IR model shows promise in addressing a primary challenge to its usefulness for athlete monitoring.
This article proposes increasingly complex models based on publicly available data involving rally length. The models provide insights regarding player characteristics involving the ability to extend rallies and relates these characteristics to performance measures. The analysis highlights some important features that make a difference between winning and losing, and therefore provides feedback on how players may improve.
This paper considers the impact of unforced errors in sport. Although the proposed methods are applicable to various sports, we demonstrate the approach in the context of professional tennis. The value of the approach is that we can provide estimates of the points lost, games lost, sets lost and matches lost due to unforced errors. The methods are based on a bootstrapping procedure which also yields standard errors for the estimates. The approach is valuable in terms of player evaluation, and can also be used for training purposes where it is possible to assess the quantification of improvement based on fewer unforced errors.
This paper considers how player acceleration changes in soccer relative to age. A plot of average maximum acceleration versus age is produced. The construction of the plot is based on methods from functional data analysis and the availability of tracking data from the 2019 season of the Chinese Super League. For an individual player, we calculate his maximum acceleration for each single match of the 2019 season. Since the players' maximum accelerations are observed only on a single season instead of their entire careers, we treat them as incomplete functional data, called functional snippets. The average maximum acceleration, i.e., the mean function of the functional snippets rather than full curves is estimated by a local linear smoothing method. The most important observation is that the shape of the acceleration curve closely resembles curves of soccer performance versus age. This observation has implications for predicting future performance since acceleration is more easily and more accurately measured than performance.
Accepted by: Phil ScarfThis paper investigates optimal target locations for throw-ins in soccer. The investigation is facilitated by the use of tracking data which provide the positioning of players measured at frequent intervals (i.e. 10 times per second). The methods for the investigation are necessarily causal since there are confounding variables that impact both the throw-in location and the result of the throw-in. A simple causal analysis indicates that on average, backwards throw-ins are beneficial and lead to an extra two shots per 100 throw-ins. We also observe that there is a benefit to long throw-ins where on average, they result in roughly four more shots per 100 throw-ins. These results are corroborated by a more complex causal analysis that relies on the spatial structure of throw-ins.
Anticipating an opponent’s serve is a salient skill in tennis: a skill that undoubtedly requires hours of deliberate study to properly hone. Awareness of one’s own serve tendencies is equally as important, and helps maintain unpredictable serve patterns that keep the returner unbalanced. This paper investigates intended serve direction with Bayesian hierarchical models applied on an extensive, and now publicly available data source of professional tennis players at Roland Garros. We find discernible differences between men’s and women’s tennis, and between individual players. General serve tendencies such as the preference of serving towards the Body on second serve and on high pressure points are revealed.
This short communication considers the calculation of player speed from tracking data. Whereas there are many player tracking systems, all rely on the collection of Cartesian coordinates corresponding to the players on the pitch. From these Cartesian coordinates, there are many ways that one could approximate player speed and acceleration. We introduce some simple principles from exploratory data analysis, which help yield more reliable speed calculations. The general principles are illustrated on various player tracking systems.
This paper explores defensive play in soccer. The analysis is predicated on the assumption that the area of the convex hull formed by the players on a team provides a proxy for defensive style where small areas coincide with a greater defensive focus. With the availability of tracking data, the massive dataset considered in this paper consists of areas of convex hulls, related covariates and shots taken during matches. Whereas the pre-processing of the data is an exercise in data science, the statistical analysis is carried out using linear models. The resultant messages are nuanced but the primary message suggests that an extreme defensive style (defined by a small convex hull) is negatively associated with generating shots.