
Abstract The maximum 147 break is snooker’s most celebrated achievement, yet its frequency has never been analysed statistically. We compile a season-level dataset of 44 seasons (1981–82 to 2024–25) and 217 maximum breaks. Using negative binomial regression with a log-frames offset, we find the 147 rate rising exponentially at 6.8 % per season (IRR = 1.068, 95 % CI 1.052–1.083). This outpaces the century rate (4.8 % per season), so the extreme tail of the break distribution is fattening. Because every 147 is also a century, we decompose the trend into three stages: the 147 rate, the century rate, and the conditional probability that a century reaches 147, rather than a single collinear model. All three rise; the conditional probability increases about 1.3 % per season, though only marginally significantly. Financial incentives, which ranged from £0 to £147,000, show no significant association in any specification (all p > 0.5). The mean age of 147-makers has risen 3.1 years per decade ( p < 0.001) as the age distribution widens, and the number of unique makers per ten-season window has increased more than fivefold (8–45). Maximum breaks are becoming more common as break-building skill improves, not as a response to incentives or expanding opportunity.
Abstract This paper considers the tactic of pressing in soccer. The investigation defines the conditions that constitute a traditional press, which then permits the identification of pressing situations using tracking data. The data consist of matches from the 2019 season of the Chinese Super League. Descriptive analyses suggest that traditional pressing is associated with more turnovers when the pressed player is given less time on the ball and as the intensity of the press increases (i.e. more pressers). It is also observed that counter-pressing is associated with better outcomes than a traditional press. Statistical modeling provides further insights with respect to pressing. For example, it is observed that pressing is associated with better outcomes when it is initiated deep and central in the opponent’s end of the field.
Abstract This paper introduces a new generalisation of the Skellam distribution based on independent mean-parameterised Conway-Maxwell-Poisson (MPCMP) marginals with flexible, unequal dispersion. The resulting MPCMP-Skellam model can capture both under- and overdispersion – including bidispersion – in the difference of two count processes, which arise frequently in sports applications. We demonstrate the utility of this model through simulation studies and three real sports datasets: goal differences from Italy’s Serie A and Tunisia’s Ligue 1, and football yellow card differences. The proposed model offers an improvement over the Skellam model in each case, is implemented in reproducible code and extends the toolkit available to sports analysts for modelling count differences subject to less variability than permiited via adoption of the standard Skellam model.
This paper proposes a binary time-series model which captures intransitivity in paired comparison by extending the classic Elo Rating System (ELO). Most existing rating systems assume that players' ratings are totally ordered and that transitivity holds, which precludes situations where a player performs exceptionally well against a specific opponent regardless of their overall ability. The proposed model, called the Pairwise-Elo (P-ELO) model, allows us to (i) construct a hypothesis test for the existence of pairwise advantages (i.e., intransitivity) across the players, and (ii) estimate winning probabilities for each pair while incorporating these pairwise effects. We compare P-ELO to ELO and mElo2k on Sumo Wrestling, Mixed Martial Arts, StarCraft II, and Large Language Models (LLMs) evaluations, in terms of match outcome prediction and probability estimation.
Understanding the key factors driving attacking success represents a critical challenge in soccer match analysis. A promising approach to address this issue is the application of expected possession value (EPV) models. Therefore, this paper aims to develop an EPV model with high explainability to provide detailed practical insights into the keys to attacking success. Tracking and event data of the Bundesliga season 2022/23 were analyzed (306 matches). From three main categories (match performance offense, defense, & match situation context), 21 features were carefully selected by professional match analysts. Afterward, machine learning classifiers were used (e.g. Random Forest, XGBoost) to predict the goal probability of possessions. The selected EPV model showed satisfactory prediction performance (xGBoost: Accuracy = 0.99, Recall = 0.06, F1-Score = 0.10, AUC = 0.85, logloss = 0.05, ECE = 0.01). The most important features in predicting attacking success were the distance (1st) and angle (2nd) to the goal, the offensive space control in the final third (3rd), and the relative pitch position (4th, defined by the number of outplayed opposing formation lines). By applying the presented EPV model and interpreting its most important features in individual match situations or over a whole match highly practice-relevant information about the tactical match performance of players can be gained.
Football clubs analyze large amounts of data in an attempt to improve their performance and gain a competitive advantage over rivals. Several attempts have been made in formulating, detecting, and measuring team-based indicators. One team-based indicator popular with football analysts and managers, is a so-called playing style. Analysts at all levels of the game regularly use the term playing style to better understand the complexity of football matches and team tactics. A formal definition, let alone a proper quantification, of a playing style is typically not provided. In this paper, we introduce a method for quantifying a team's playing style based on match event data. In particular, we define playing styles based on the location and patterns of a team's consecutive actions with the ball. Specifically, we apply Latent Dirichlet allocation to ball movement patterns to obtain distributions over such ball movement patterns with similar structure; that is, the playing styles. Using our method, a team's playing style is represented by a "style" vector, that summarizes the playing style in an interpretable way. We apply our methodology to a publicly available data set, and illustrate how the resulting playing styles can be used in practice.
This study develops a dynamic framework for evaluating professional football player contracts under institutional trading constraints. We incorporate transfer windows and release clauses, which restrict when transfers can occur and under what terms. Player performance is represented by a stochastic process based on a revenue-adjusted Opta metric, and the club's decision problem is formulated with time-dependent constraints. Numerical analysis reveals two main mechanisms. First, under realistic market conditions, voluntary sellouts are not optimal ("rational retention"). As contract expiration approaches, the declining remaining duration expands the set of situations in which accepting a transfer becomes optimal, generating a pronounced deadline effect. Second, release clauses reduce the club's discretion by making high-performing players harder to retain when interest from other clubs arrives. Sensitivity analysis shows that the resulting cost is particularly large for high-growth scenarios and in high-liquidity markets with frequent transfer offers, while it is less affected by downside risks such as injuries. An empirical application using proprietary data from a J.League (Japan Professional Football League) club shows that incorporating these institutional features changes contract evaluations materially, yields a lower bound tied to the release fee relative to a static DCF benchmark, and shows through counterfactual experiments that the deadline effect is driven by horizon contraction rather than short-term performance fluctuations.
Understanding offensive player profiles in basketball requires analyzing both the spatial distribution of shot attempts and the corresponding scoring effectiveness across court locations. Motivated by this, we propose a semiparametric spatial point process model that combines flexible spline-based intensity estimation with player-specific mixed effects to characterize shooting intensity as a function of distance and angle from the basket. The model balances flexibility and interpretability while accommodating the hierarchical structure of player-level data through random effects. We apply this framework to shooting data from the NBA 2024/2025 season, focusing on High-Volume Shooters who exhibit diverse spatial patterns. The model captures both population-level shooting tendencies and individual deviations, revealing substantial heterogeneity in offensive behavior across players. Extending the framework to marked point processes allows us to construct spatial scoring probability maps that quantify efficiency beyond shot frequency. Collectively, the proposed methodology offers a unified statistical approach for modeling spatial shooting behavior and provides actionable insights for player evaluation, tactical decision-making, and roster construction.
In basketball and other team sports, assessing players' and lineups' performances is crucial for team management. This assessment should address both individual players' contributions and the synergies among them. Here, we aim to analyze the importance of collaboration among NBA players from the 2002-2003 to the 2022-2023 season. To this end, we propose a modification of the extended Regularized Adjusted Plus-Minus model, where the efficacy of each lineup is determined by adding up the non-negative contributions of players on the court, adjusted for a positive or negative interaction term. The proposed specification improves the readability of the estimated values of players' and lineups' effects, which can be straightforwardly visualized and used more efficiently for actual decision-making. The main analyses based on this model specification revealed that, in the NBA, the importance of synergies among players decreases over the years, suggesting a shift from team to individual-oriented style of play. The only notable exception occurs during the COVID-19 lockdown period, when the relevance of lineups effects temporarily increases, indicating a change towards a team-oriented style of play. The proposed model can also be adopted to evaluate the difference between home and away-specific performances of players and lineups.
This study utilises the Markov Chain process to generate football odds, focusing on historic La Liga matches between FC Barcelona and Real Madrid. Using data from 93 encounters, it categorises matches based on home and away performances. Transition matrices derived from historical results calculate probabilities for match outcomes (win, draw, loss) and over/under goal markets, which are then converted into betting odds. Findings reveal Barcelona's strong home advantage, with a 55 % probability of winning (odds: 1.82), while Real Madrid has a 59 % probability at home (odds: 1.69). However, both teams experience a significant drop when playing away - Barcelona with 24 % (odds: 4.17) and Real Madrid with 23 % (odds: 4.35). The study also shows a higher likelihood of over 1.5 and 2.5 goals in their matches. This research highlights the Markov Chain process as an effective tool for improving football odds accuracy, offering a data-driven alternative to subjective betting models.
Recently, UEFA changed the group stage of its international soccer competitions to an incomplete round robin tournament. Previously, teams were divided into groups, each playing a double round robin tournament with a resulting ranking table. In contrast, the new format has all teams competing in one league, producing a single ranking. We investigate the effect of the new format on the probability of competitive matches in the UEFA Champions League. A match is non-competitive if the prize for at least one opponent does not depend on the match outcome, or if there exists an opportunity for both opponents to collude; otherwise, we call a match competitive. Using Monte Carlo simulations and integer programming, we show that the new format results in proportionally more competitive matches than the old format, although it also results in a noticeable increase in the risk of collusion.
This paper introduces a Bayesian Additive Regression Trees (BART) approach to correct sprint kayak and canoe race times for environmental conditions and predict future winning times using historical data. The model delivers refined estimates, enabling accurate forecasts and real-time adjustments under changing conditions. Trained on data from elite international and Olympic competitions over the past decade, BART achieved an out-of-sample mean relative error (MRE) of 3.8 % and an in-sample MRE of 1.4 % for 1,000 m, 500 m, and 200 m events at the 2022 Halifax Championships. Consistent with prior research, wind strength, air temperature, and water temperature were the most influential variables, while salinity, which is not generally used in the analysis, proved important, with event venues classified as freshwater or saltwater. A major strength of BART is its Bayesian framework, which yields predictive distributions that allow sport scientists to estimate probabilities such as achieving a target time. Model validation included meta-analysis of variable importance across 13 distance and boat-class sub-models, Gelman-Rubin MCMC convergence checks, and cross-validation for hyper-parameter tuning. Temporal hold-out validations for Halifax 2022 and the London 2012, Rio 2016, and Tokyo 2021 Olympics confirmed BART's superior performance, with out-of-sample MRE less than half that of Bayesian linear models.
The use of high-fidelity camera-vision and Doppler tracking systems in Major League Baseball (MLB) has created an influx of advanced analytics that have transformed game and personnel strategy. However, due to the cost and complexity of these systems, advanced statistics are severely limited in most amateur games. In this work, we develop a robust pitch reconstruction methodology using artificial neural networks (ANN) and a three-point reconstruction technique, requiring the knowledge of only three baseball spatiotemporal locations. The ANN models were trained to predict the initial baseball speed and spin rate, which are then used as initial conditions to integrate the full pitch trajectory. We use numerically simulated baseball pitches to train two ANN models, each with clean and noisy training inputs, respectively. We then performed a robustness analysis to test ANN performance with increasingly noisy data to simulate low-fidelity camera tracking systems. Probability distributions of predicted model output are calculated using the Monte Carlo method to quantify model uncertainty. We show that the ANN models accurately predict the true baseball trajectory from both high-fidelity and noise-injected testing data. The results demonstrate the effectiveness of ANN models for quick and robust pitch trajectory reconstruction using minimal input data, even in the presence of noise. The present methodology provides a first step towards enabling pitch tracking and advanced analytics in a wider variety of baseball games.
Because the decathlon tests many facets of athleticism, including sprinting, throwing, jumping, and endurance, many consider it to be the ultimate test of athletic ability. On this view, estimating the maximal decathlon score and understanding what it would take to achieve that score provides insight into the upper limits of human athletic potential. To this end, we develop a Bayesian composition model for forecasting how individual decathletes perform in each of the 10 decathlon events of time. Besides capturing potential non-linear temporal trends in performance, our model carefully captures the dependence between performance in an event and all preceding events. Using our model, we can simulate and evaluate the distribution of the maximal possible scores and identify profiles of decathletes who could realistically attain scores approaching this limit.
We propose a new probabilistic and interpretable approach to quantify competitive balance in sports leagues based on the precise moment when a tournament can no longer be perceived as perfectly balanced: the longer it takes, the more balanced it is. We analyzed 1,539 seasons from 175 sports leagues in basketball, soccer, handball, and volleyball and observed that only 5 % of the seasons could be seen as perfectly balanced throughout their entire duration. For the others, there is a turning point round after which the points distribution permanently diverges from a range of behaviors likely to occur in perfectly balanced tournaments. Our initial results agree with the general literature since soccer has a considerably more random behavior, i.e., its perceived balance is higher overall. Given the explicit temporal dependence, we also proposed a modification to remove the bias that the order of matches could impose, thereby enabling fair comparisons across different leagues and sports. Our modified coefficient highlights the substantial imbalance in heavily debated leagues, such as the Premier League and the NBA. Furthermore, combining both metrics enables the discovery of anomalous tournament schedules where even simple random changes to the match order could have improved the perceived competitive balance.
Following a penalty in rugby union, teams typically choose between attempting a shot at goal or kicking to touch to pursue a try. We develop an Expected Points (EP) framework that quantifies the value of each option as a function of both field location and game context. Using phase-level data from the 2018/19 Premiership Rugby season (35,199 phases across 132 matches) and an angle-distance model of penalty kick success estimated from international records, we construct two surfaces: (i) the expected points of a possession beginning with a lineout, and (ii) the expected points of a kick at goal, taking into account the in-game consequences of made and missed kicks. We then compare these surfaces to produce decision maps that indicate where kicking for goal or kicking to touch maximizes expected return, and we analyze how the boundary shifts with game context and the expected meters gained to touch. Our results provide a unified, data-driven method for evaluating penalty decisions and can be tailored to team-specific kickers and lineout units. This study offers, to our knowledge, the first comprehensive EP-based assessment of penalty strategy in rugby union and outlines extensions to win-probability analysis and richer tracking data.
This paper presents the quantile cube, a novel three-dimensional summary representation designed to analyze external load using GPS-derived movement data. While broadly applicable, we demonstrate its utility through an application to data from elite female soccer athletes across 23 matches. The quantile cube segments athlete movements into discrete quantiles of velocity, acceleration, and movement angle across match halves, providing a structured and interpretable framework to capture complex movement dynamics. Statistical analysis revealed significant differences in movement distributions between the first and second halves for individual athletes across the vast majority of matches (188/198). Principal component analysis identified matches with unique movement dynamics, particularly at the start and end of the season. Dirichlet-multinomial regression further explored how factors such as athlete position, playing time, and match characteristics influenced movement profiles. Our analysis reveals external load variations over time and provides insights into performance optimization. The integration of these statistical techniques demonstrates the potential of data-driven strategies to enhance athlete monitoring and workload management in women's soccer.
This study evaluates the effectiveness of the two-for-one strategy in basketball by applying a causal inference framework to play-by-play data from the 2018-19 and 2021-22 National Basketball Association regular seasons. Incorporating factors such as player lineup, betting odds, and player ratings, we compute the average treatment effect and find that the two-for-one strategy has a positive impact on game outcomes, suggesting it can benefit teams when employed effectively. Additionally, we investigate potential heterogeneity in the strategy's effectiveness using the causal forest framework, with tests indicating no significant variation across different contexts. These findings offer valuable insights into the tactical advantages of the two-for-one strategy in professional basketball.
Evaluating the value-added of coaches in the NBA is a challenging task as the coaches with the best win/loss records often have the best players. This prompts a question of attribution: if two coaches had the same roster, which one would win? This paper attempts to answer this question by introducing a method for quantifying coaching effect in the NBA. We propose a method for isolating the effect of a coach's in-game scheme on their team's probability of winning a game while controlling for other factors, namely the relative strength of the two competing teams. To control for team strength, player performance metrics are aggregated into "Team-Adjusted VORP Difference" or Delta tVORP, meant to account for the difference in quality of on-court product between both teams. We model each coach's win probability as a function of Delta tVORP using probit monotone Bayesian Additive Regression Trees. In comparing coaches' win probability curves, we find some of the winningest coaches are close to average in terms of scheme, while other coaches are found to be truly great contributors to their teams.