
To prevent tanking, sports leagues should require every team to trade away its own first-round draft pick—the pick for the draft following that season—before the season begins. After the season, draft order would still be determined in reverse standings order, but no team would own the pick generated by its own performance. Losing would no longer improve a team's draft position, eliminating in-season tanking incentives. More broadly, the Draft Pick Trade Deadline implements a novel design principle for sports drafts: compensation should be assigned before the season based on expected, rather than realized, performance. Expectations-based draft compensation eliminates tanking incentives while maintaining the draft's core purpose of promoting parity. I prove a theorem showing that the only way to eliminate in-season tanking incentives is to base draft compensation on information unaffected by teams’ in-season performance. I also show, using both a theoretical model and an empirical analysis of NBA data, that any expectations-based system achieves what no lottery reform can: a complete decoupling of in-season incentives from draft compensation. Teams expected to perform poorly still receive value, paid before the season rather than after. Expectations-based approaches eliminate tanking while preserving competitive balance and can be implemented immediately without transition costs.
The increasing use of synthetic data in sports science reflects persistent limitations in real-world sport datasets, including small sample sizes, class imbalance, privacy restrictions, and the absence of ground truth. While synthetic data generation has been applied across a wide range of sporting contexts, its use remains heterogeneous, with limited consistency in design, evaluation, and reporting practices. This study synthesises evidence from twelve sport studies to address a central gap in the literature: the lack of an established, domain-specific framework to guide the systematic generation, evaluation, and deployment of synthetic data in sport. Rather than proposing a new generative model, this work introduces a decision-oriented framework that structures synthetic data projects across six interrelated dimensions: objective of use, data structure, generation strategy, domain constraints, utility and fidelity evaluation, and deployment risk. Analysis of the reviewed studies shows that synthetic data are most effective when generation strategies are aligned with sport-specific domain knowledge and clearly defined operational goals. The framework highlights recurring challenges related to domain shift, constrained realism, and class imbalance, and demonstrates how these issues can be addressed through explicit design and evaluation choices.
The 2022 NFL rule change guaranteeing both teams possession in playoff overtime raised a strategic question: should the coin-toss winner receive or kick? We develop a Monte Carlo simulation framework focused on two asymmetric factors: sudden-death positioning (favoring receiving, since the receiver retains first possession if scores are matched) and information asymmetry (favoring kicking, since the kicker knows what is required before its possession). Across a wide range of offensive efficiencies (TD rates 15–35%, FG rates 7.5–27.5%), receiving is favored throughout; within the parameter space examined, no combination renders kicking optimal. At baseline parameters with optimal two-point conversion strategy, the receiving team wins 53.14% of simulated games (95% CI: 52.84%–53.44%), a receive advantage of +6.47 percentage points (95% CI: +5.87, +7.07). Among 16 decided games under current rules, receiving teams won 9 (56.3%; Wilson 95% CI: 33.2%–76.9%), a small sample compatible with simulation predictions but unable to establish them. Despite this simulated advantage, most coin-toss winners have elected to kick. We interpret the sudden-death positioning advantage as outweighing the information-asymmetry advantage, and suggest that the prevailing “information advantage” intuition may lead teams toward a strategy that underperforms in simulation.
In Major League Baseball, every ballpark is unique, with its own geometry and climate. Some ballparks may be more conducive to home runs than others. Quantifications of home run friendliness abound, but are often based on limited data and typically do not include uncertainty assessment. Further, personnel effects from individual players are rarely considered. We fit generalized linear models, taking as the observational unit the combination of game and handedness-matchup of the batter and pitcher, usually leading to four home run totals per game. The Poisson model provides a good fit for counts observed in the 2010-2024 season and generalizes well out-of-sample. We model personnel effects by constructing “elsewhere” measures of individual batter and pitcher home run tendency using data from parks other than the one in which the response is observed. All pairwise comparisons of ballparks are made, with multiplicity adjustment, using means adjusted to teams of average batters facing average pitchers in average handedness frequencies. Estimated standard errors for these means and differences are reported. We find that adjusted home run frequencies are substantially different from observed frequencies, leading to considerably different ballpark rankings than those based on unitless park factors appearing in baseball media.
Final time and rank do not fully characterize orienteering performance: athletes with similar outcomes can differ substantially in consistency and in the severity of major navigational mistakes. We propose a split-time framework that decomposes performance into three interpretable components: speed, instability, and mistake severity. The framework combines a descriptive layer based on leg-level relative log-times with an additive model for log split times that includes athlete effects (baseline pace), leg effects (common difficulty), and athlete-leg residuals. Under a complete split-time table, least-squares estimation has a transparent closed-form representation in terms of doubly centered log split times. Using elite split-time data from the 2025 World Orienteering Championships middle- and long-distance finals, we show that speed is most closely aligned with final outcome, while instability and mistake severity provide complementary execution-level information rather than replacements for official results. We also report bootstrap uncertainty intervals and a minimal missing-data sensitivity analysis. The proposed framework is simple to compute, interpretable in practice, and extensible for applied sports analytics.
Sports spectatorship is fundamentally affective, and the rise of fan-generated data has enabled large-scale computational analysis of fan sentiment and emotion. This paper synthesises 64 studies (2000–2025) that infer fan affect from (i) digital text and online discourse, (ii) broadcast-linked second-screen interaction, and (iii) in-venue or non-textual signals such as crowd audio, video, and wearable sensing. We organise the literature into three modality-based categories: text-centric discourse analytics, second-screen/social-TV behaviour, and in-venue or multimodal sensing. Across studies, a consistent empirical pattern is event-driven synchrony: aggregate affective signals shift rapidly around salient match events and controversy. However, three structural limitations constrain behavioural inference and generalisability: strong platform dependence (especially Twitter/X), overreliance on coarse polarity sentiment, and conceptual slippage between affective expression and behaviour. Research on harmful expressions (e.g., toxicity, hate speech) is expanding and behaviourally relevant, but introduces methodological and ethical challenges. Overall, the literature is effective at detecting affective expression but weaker in linking affect to explicit behavioural outcomes or integrating evidence across modalities. We highlight directions for behaviour-centred, multimodal fan analytics, including improved construct validity, clearer outcome definitions, cross-modality integration, and stronger cross-platform and ethical considerations.
With advancements in soccer analytics, considerable attention has been devoted to developing sophisticated measures for quality scoring opportunities - such as expected goals and expected threat. Far less effort, however, has gone toward adjusting these and other performance metrics for game context, including factors like score and red card differential. It is well known that certain match scenarios prompt systematic tactical adjustments - for example, teams often defend more conservatively when leading or playing shorthanded. In this study, we employ generalized additive mixed-effects models (GAMMs) on minute-by-minute match event data from Europe’s five major leagues to quantify how such contextual variables influence offensive production metrics like shots on target and expected goals. This approach allows us to estimate global effects of contextual variables on offensive production, while also allowing for league-specific distinctions to be captured via random effects. We then use these estimates to adjust observed team statistics, projecting each team’s performance onto a standardized baseline scenario - a tied home game played at even manpower. We believe the resulting adjusted measures to better isolate underlying team quality by removing distortions arising from transient tactical responses to game state, thereby providing improved inputs for downstream analytical and predictive tasks.
Understanding how humans adapt to streaky performance under uncertainty is central to research on belief updating, attention, and decision-making. Professional basketball provides a naturalistic setting in which these processes can be observed under real incentives and time pressure. Using NBA player-tracking data, we examine how defenders adjust their behavior in response to offensive shooting outcomes, shifting focus from shooter performance to defensive adaptation. We conduct three complementary studies relating multiple spatial defensive metrics to recent shooting history, including asymmetric responses to makes versus misses, cumulative success over short memory windows, and escalation during extended streaks. Across analyses, we estimate regression-based models with appropriate clustering and controls to isolate behavioral updating from contextual factors. The results show clear evidence of asymmetric updating: defenders tighten coverage more after made shots than they relax following misses, with effects strongest for the closest defender. Defensive adjustments are driven primarily by short recency windows of approximately three to five shots. Escalation during long hot streaks is limited, with defensive responses stabilizing rather than increasing indefinitely. Together, these findings suggest that defensive behavior reflects psychologically structured but bounded belief updating rather than full rational inference or indiscriminate overreaction to streaks.
The Model Context Protocol (MCP) enables large language models to query and synthesize sports data across distributed sources, yet it lacks built-in mechanisms for provenance, integrity, and athlete-controlled access. This study proposes a hybrid MCP–blockchain architecture for LLM4Sports , a domain-specific model fine-tuned on localized sports datasets. In this design, MCP orchestrates multi-database retrieval, while a blockchain layer immutably anchors query and data-bundle hashes. Smart contracts manage time-bound and granular consent through decentralized identifiers and verifiable credentials, and auditable logs ensure end-to-end traceability. We contribute: (i) a trust-enhanced reference architecture combining MCP and blockchain; (ii) reusable prompt and workflow templates that embed consent validation within natural-language tasks (e.g., “Summarize athlete X's performance with verified consent”); and (iii) prototype implementations for both team analytics and personalized coaching. Evaluations on synthetic but realistic workloads show high verification accuracy and minimal orchestration overhead, demonstrating the feasibility of real-time, consent-aware analytics. The proposed framework enhances interoperability, regulatory compliance (e.g., GDPR), and athlete autonomy, while bolstering practitioner trust. We conclude by discussing scalability, legacy integration, and privacy trade-offs, and by outlining next steps toward field pilots and multimodal extensions (e.g., video provenance) across professional and amateur sports ecosystems.
For a sport that is approaching a state of analytic saturation, major league baseball is devoid of one seemingly critical metric: a manager-value estimator. Whether managers have ever meaningfully influenced their teams’ performances and still do today are matters of significant dissensus. Without a manager-value estimator, that debate (critical, at a minimum, to informed front-office decisionmaking) cannot be convincingly resolved. Based on a sample of over 500 managers spanning the history of AL/NL seasons since 1901, this paper develops a manager-value estimator based on manager performance in relation to team records predicted by aggregate player WARs. Simulation and Bayesian methods are used to test for False Discovery Rates and to form posterior estimates in relation to Regions of Practical Equivalence. Results suggest that a substantial fraction of managers (including current and recently active ones) have over their careers influenced team “winning percentages” ≥ ± 0.012, the equivalent of ± 2 wins per 162 games. In addition to enabling historical and contemporary comparisons, the mWAR Estimator can also be calibrated to reflect the asymmetric-error costs and risk preferences that characterize the tournament structure of contemporary MLB economics.
This study investigated temporal changes in service returns in elite table tennis by analyzing men's and women's singles matches from the 2012, 2016, and 2021 Olympic Games. Quarterfinals and subsequent matches were analyzed, focusing on service placement and service return stroke types. Chi-square tests and effect sizes were employed to assess longitudinal and sex-related differences. The results revealed significant temporal changes in women's matches, characterized by a consistent increase in backhand-based service returns, particularly backhand topspin, accompanied by a shift in service placement toward the short-forehand area. Consequently, sex differences in service return stroke types for services directed toward the backhand and middle areas diminished over time. In contrast, men's matches exhibited no consistent temporal changes in the frequency of service return stroke types. However, observed shifts in service placement suggest changes in tactical patterns that were not fully captured by usage frequency alone. Despite an overall convergence in certain aspects, marked sex differences persisted in service return stroke types for services directed to the forehand side. These findings indicate that service returns in elite table tennis evolved differently for men and women during the examined period and have practical implications for coaching and training, particularly for female players.
Penalty shootouts in association football are sometimes criticized by fans and pundits as an imperfect tie-breaking procedure. In this study, we analyze through a Bayesian model if shootouts are governed more by skill or by chance. Using a representative dataset from twelve recent European seasons, we fit a hierarchical logistic model with appropriate random effects and a within–shootout latent autoregressive state to capture evolving pressure. The model is implemented through Hamiltonian Monte Carlo approach. Our proposed framework allows us to quantify the amount of skill involved in shootouts by the proportion of logit-scale variance attributable to persistent heterogeneity versus idiosyncratic and state noise. We also compare the full specification to nested alternatives via PSIS-LOO, stacking, and decision-oriented scores computed from leave-one-shootout-out predictive distributions. Empirically, persistent individual effects are found to be small: the posterior SkillShare is near zero in the full model, suggesting that shootouts are primarily chance dominated. As a by-product of our approach, we show how the proposed model can also be utilized to rank the shooters as well as to optimize penalty-taking orders. We also discuss a few alternative tie-breaking procedures as future recommendations which can be evaluated in a similar modeling framework.
The aim of this study was to analyse lap performance stability, stroke kinematic variables in the clean swimming phase and turn performance in the finals of the four 200 m short course (SCM) events (freestyle, backstroke, breaststroke, and butterfly). The sample included thirty-two elite female swimmers who competed in the 200 m finals at the 2019 LEN European Short Course Championships in Glasgow. Lap times, stroke kinematics, and turns were analysed using two approaches: (i) official lap time comparison (50 m), and; (ii) comparison of consecutive segments (25 m). Based on the official lap time comparison, performance significantly decreased over time (p < 0.001) with strong effect sizes in all strokes (η 2 between 0.89 and 0.96). Pairwise comparisons between consecutive 25 m segments were not significant. Turn time increased in all four strokes with strong effect sizes (p < 0.001; η 2 between 0.71 and 0.84). The start contributed ∼6% and turns ∼52% of the total race time. These findings highlight that traditional pacing analysis based on official splits may lead to misleading interpretations of performance in SCM events. Analysing 25 m segments provides a more accurate representation of race dynamics and contributes to a clearer understanding of pacing stability in short-course swimming.
Implementing state-transition models in team sports such as football (soccer), can be complex and time-consuming if observational methods need to be used. To help alleviate this, the aims of this study are two-fold. First, based on existing phases of play models, we introduce a hand-crafted expert state system that combines in- and out-of-possession phases of play to form a set of comprehensive tactical states. Second, we introduce a machine learning approach for automatically detecting the introduced tactical states based on positional data. This is implemented by capturing positional configurations of the players and the corresponding ball location. By clustering the player configurations for each ball zone and learning a “translation table” between the identified clusters and the expert state system, using three training games, the framework can be used to automatically detect states from the positional data of new games. The performance of our framework is evaluated by comparing the automatically detected states to the state label assigned by a human annotator using twelve test games from the 2021/2022 Bundesliga season. This showed good agreement, with F1-scores greater than 0.9 for most states. Hence, our framework may provide a method for time-efficient and comprehensive match analysis in football performance analysis.
This study analyzed post match data from the men's singles events at the 2025 Australian Open, French Open, and Wimbledon Championships to examine the relationship between match outcomes and technical performance across different court surfaces. Technical variables were compared between winners and losers using paired-samples t-tests, followed by exploratory binary logistic regression to identify key performance indicators. Surface-related differences in champions’ performance were examined descriptively across hard, clay, and grass courts. The results indicated that first serve points won and return points won were key predictors of match outcomes across all surfaces, with return points won showing predictive value specifically at the Australian Open and Wimbledon, and break point conversion rate being predictive at the French Open. Championship players recorded a higher number of winners at the French Open compared with the Australian Open and Wimbledon, suggesting that surface-specific bounce characteristics may facilitate more aggressive offensive play.
This paper evaluates the effectiveness of timeouts in the EuroLeague using Play-by-Play data to estimate short-run within-game effects and Box-Score data to assess season-level outcomes over the 2021–22 to 2023–24 regular seasons. Within games, timeout effectiveness is identified using an event-study framework comparing team performance in fixed possession windows immediately before and after each timeout, complemented by a difference-in-differences strategy exploiting situations in which teams are unable to call additional timeouts. Performance is measured by the point differential (points scored minus points conceded) over these windows. Under the identifying assumption that pre-timeout trends are comparable across treated and control situations, timeouts are associated with a statistically significant short-run improvement: the average point differential increases from − 3.74 before a timeout to − 1.20 afterward. Although economically meaningful in reducing opponent scoring runs by approximately 2.5 points on average, the post-timeout differential remains negative, indicating that timeouts primarily stabilize performance rather than reverse game momentum. To assess whether these short-run effects translate into sustained success, we extend the standard Four-Factor Model of Wins with a fifth factor capturing team-level timeout effectiveness. This additional factor provides no incremental explanatory power. From a coaching perspective, timeouts appear to function as short-term damage-control devices but do not systematically affect season-long competitive outcomes.
The intermittent work in soccer demands rapid changes in velocity and direction due to the unpredictable events that occur during match play. Hence, the players’ ability to accelerate and decelerate are important indicators for physical performance and load. It is common for elite teams to use inertial measurement units (IMU) to monitor the external training load, of which accelerations and decelerations are considered important variables. Accelerations and decelerations have traditionally been determined relative to a fixed threshold (e.g., >2.0 m·s −2 ). The aim of this study was to examine accelerations and decelerations in several zones (graded from low to high) in two different levels of sensitivity and in three different thresholds for determination of accelerations and decelerations. Thirty-five semi-professional soccer players were monitored with a foot-mounted IMU (Playermaker) during 60 official matches, where number- and meters of accelerations and decelerations were distributed from low to high in sensitivity level 1 (6 zones), sensitivity level 2 (3 zones), and three different thresholds for determination of accelerations and decelerations. Sensitivity and threshold are shown to influence positional differences when acceleration and deceleration values are relatively low but appear to be absent when the values are relatively high. High sensitivity leads to positional differences that are not present for low sensitivity, whereas low sensitivity polarizes differences between zones. For thresholds, the use of acceleration meters leads to greater positional differences compared to number of accelerations. These findings should be considered in the interpretation of the data output and in decisions regarding training load management.
In judged sports, such as rhythmic gymnastics, figure skating, and baton twirling, inter-judge variability in scoring—the degree to which judges differ in their scoring of the same athlete—is a common concern. This study quantitatively examined inter-judge variability using freestyle scores from the World Baton Twirling Championships held in 2018 and 2022. Data were collected from the preliminary rounds of senior and junior women's divisions, with scores assigned by seven judges. Welch's analyses of variance were performed to assess the effects of athlete ranking group (high, middle, low) on inter-judge variability across two scoring axes: technical merit (TM) and artistic expression (AE). These analyses were conducted separately for each competition year (2018 and 2022) and division (senior and junior). In the senior division in 2018, inter-judge variabilities for both TM and AE were significantly lower in the high-ranking group than in the other ranking groups. In the junior division in 2018, inter-judge variabilities for both TM and AE were significantly lower in the high-ranking group than in the low-ranking group. These findings are interpreted in terms of the interactions among the competitive structure of baton twirling and the cognitive processes involved in judging.
This work reconciles two perspectives on the Elo ranking that coexist in the literature: the practitioner's view as a heuristic feedback rule, and the statistician's view as online maximum likelihood estimation via stochastic gradient ascent. Both perspectives coincide exactly in the binary case (iff the expected score is the logistic function). However, estimation noise forces a principled decoupling between the model used for ranking and the model used for prediction: the effective scale and home-field advantage parameter must be adjusted to account for the noise. We provide both closed-form corrections and a data-driven identification procedure. For multilevel outcomes, an exact relationship exists when outcome scores are uniformly spaced, but approximations are preferred in general: they account for estimation noise and better fit the data. The decoupled approach substantially outperforms the conventional one that reuses the ranking model for prediction, and serves as a diagnostic of convergence status. Applied to six years of FIFA men's ranking, we find that the ranking had not converged for the vast majority of national teams. The paper is written in a semi-tutorial style accessible to practitioners, with all key results accompanied by closed-form expressions and numerical examples.
In sports, players transitioning between leagues often experience changes in performance statistics due to differences in competition level and player pools. For example, G League players (part of the NBA's minor league system) may see declines in performance metrics when called up to the NBA. Quantifying league translation factors, i.e., the expected difference in performance between leagues, is crucial for accurately contextualizing player performance and understanding key differences between leagues. We present a new method for constructing league translation factors using a matching method and difference-in-differences (DD) estimator, providing a causal estimate of how a player's existing statistics might have appeared in a different league. Unlike traditional approaches that rely on league-wide averages or Z-scores, our method constructs a comparable player pool and accounts for potential aging effects. We apply this approach to construct G League-to-NBA translation factors and compare it to the "same season" method, which examines players competing in both leagues within the same season. Our findings show that most performance statistics decline when players transition from the G League to the NBA. The DD approach produces translation factors that are directionally similar but generally smaller in magnitude than the same season approach, providing a more conservative and stable estimate.