The recent growth in data availability in football has increased the risk of incorrect use of data-driven models, making guidelines on their validation and application necessary. The Expected Threat (xT) model is an accessible option for football organizations that start building in-house methods, yet little is known about how to assess its quality. The aim of this study is twofold: to examine how the model error depends on the number of game states and the number of training points, and to translate these results into guidelines for constructing and applying the model. Using the Markov chain underlying the model, we perform theoretical analyses and simulations to study the model error. These show that the model error is approximately log-normally distributed for a specified number of training points and game states. Additionally, we combine the simulations with expert consultation to establish the model error beyond which player evaluations based on the Expected Threat model become unreliable for scouting applications. From this, we derive rules of thumb to ensure the quality of an Expected Threat model before application, and we illustrate through an example how a validated model can be applied in practice. Because the approach generalizes to Expected Possession Value models, this paper illustrates a framework to systematically quantify model quality, despite the ground truth being unobservable in football analytics.
We consider nonparametric estimation of the distribution function F of squared sphere radii in the classical Wicksell problem. Under smoothness conditions on Fin a neighborhood of x, in Gili et al. (2024) it is shown that the Isotonic Inverse Estimator (IIE) is asymptotically -efficient antt attains mic of convergente log constant or a interval containing x, the optimal rate of convergence increases to s and the IIE attains this rate adaptively, i.e. without explicitly using the knowledge of local constancy. However, in this case, the asymptotic distribution is not normal. In this paper, we introduce three informed projection-type estimators of F, which use knowledge on the interval of constancy and show these are all asymptotically equivalent and normal. Furthermore, we establish a local asymptotic minimax lower bound in this setting, proving that the three informed estimators are asymptotically efficient and a convolution result showing that the IIE is not efficient. We also derive the asymptotic distribution of the difference of the IIE with the efficient estimators, demonstrating that the IIE is not asymptotically equivalent the informed estimators. Through a simulation study, we provide evidence that the performance of the IIE closely resembles that of its competitors, supporting the use of the IIE as the standard choice when no information about F is available.
We address the problem of uncertainty quantification for the deconvolution model Z = X + Y, where X and Y are nonnegative random variables and the goal is to estimate the signal's distribution of X ∼ F_0 supported on [0,∞), from observations where the noise distribution is known. Existing frequentist methods often produce confidence intervals for F_0(x) that depend on unknown nuisance parameters, such as the density of X and its derivative, which are difficult to estimate in practice. This paper introduces a novel and computationally efficient nonparametric Bayesian approach, based on projecting the posterior, to overcome this limitation. Our method leverages the solution p to a specific Volterra integral equation as in , which relates the cumulative distribution function (CDF) of the signal, F_0, to the distribution of the observables. We place a Dirichlet Process prior directly on the distribution of the observed data Z, yielding a simple, conjugate posterior. To ensure the resulting estimates for F_0 are valid CDFs, we isotonize posterior draws taking the Greatest Convex Majorant of the primitive of the posterior draws and defining what we term the Isotonic Inverse Posterior. We show that this framework yields posterior credible sets for F_0 that are not only computationally fast to generate but also possess asymptotically correct frequentist coverage after a straightforward recalibration technique for the so-called Bayes Chernoff distribution introduced in . Our approach thus does not require the estimation of nuisance parameters to deliver uncertainty quantification for the parameter of interest F_0(x). The practical effectiveness and robustness of the method are demonstrated through a simulation study with various noise distributions for Y.
In this paper, we propose a novel Bayesian approach for nonparametric estimation in Wicksell's problem. This has important applications in astronomy for estimating the distribution of the positions of the stars in a galaxy given projected stellar positions and in materials science to determine the 3D microstructure of a material, using its 2D cross-sections. We deviate from the classical Bayesian nonparametric approach, which would place a Dirichlet Process (DP) prior on the distribution function of the unobservables, by directly placing a DP prior on the distribution function of the observables. Our method offers computational simplicity due to the conjugacy of the posterior and allows for asymptotically efficient estimation by projecting the posterior onto the L-2 subspace of increasing, right-continuous functions. Indeed, the resulting Isotonized Inverse Posterior (IIP) satisfies a Bernstein-von Mises (BvM) phenomenon with minimax asymptotic variance g0(x)/2 gamma, where gamma > 1/2 reflects the degree of H & ouml;lder continuity of the true cdf at x. Since the IIP gives automatic uncertainty quantification, it eliminates the need to estimate gamma. Our results provide the first semiparametric Bernstein-von Mises theorem for projection-based posteriors with a DP prior in inverse problems.
While it is well-known how to compute the cells of a Laguerre tessellation for a given set of weighted generator points, it is not obvious how to invert a Laguerre tessellation. That is, given that one observes a Laguerre tessellation, how can one retrieve the weighted generators corresponding to the observed cells. In this paper, we consider inversion of a class of random Laguerre tessellations known as Poisson-Laguerre tessellations. The weighted generators of observed cells of a Poisson-Laguerre tessellation are of interest because knowledge of these weighted generators is useful for statistical inference of Poisson-Laguerre tessellations. For general Laguerre tessellations we provide a characterization of all configurations of weighted generator points which yield the same Laguerre tessellation. For Poisson-Laguerre tessellations we propose a method for consistent inversion, meaning that as one observes the tessellation through increasing observation windows, a closer approximation of the original weighted generators can be obtained. In a simulation study we examine both performance of the inversion procedure, as well as the use of the obtained approximated weighted generators for nonparametrically estimating the weight distribution function corresponding to a Poisson-Laguerre tessellation.
Obtaining information about the 3D grain size distribution of metallic microstructures is crucial for understanding the mechanical behavior of metals. This paper addresses the problem of estimating the 3D grain size distribution from 2D cross sections. This is a well-known stereological problem and different estimators have been proposed in the literature. We propose a statistical estimation procedure that provides consistent estimates without relying on arbitrary binning choices. When applying this procedure to space filling structures, we investigate the impact of the choice of grain shape and propose a heuristic to choose the best grain shape. To validate our approach, we employ simulations using Laguerre-Voronoi diagrams and apply our methodology to a sample of Interstitial-Free steel, obtained via EBSD.
Battery lifetime prediction is crucial in industrial applications. However, the lack of diversity in training data often poses challenges regarding the robustness and generalization of lifetime predictions for batteries from different batches. Motivated by the early cycle data from lithium-ion batteries, this article proposes a robust transfer learning method by employing a model average framework, where the weights are determined based on the distance between the source domain and the target domain. Kernel regression is used to build the prediction of battery lifetime using early cycle data, and transfer component analysis is utilized to transfer knowledge between different domains. The case study on lithium-ion phosphate/graphite cells demonstrates that the proposed method can mitigate the impact of negative transfer and has superior performance compared to traditional methods.
With an average football (soccer) match recording over 3,000 on-ball events, effective use of this event data is essential for practitioners at football clubs to obtain meaningful insights. Models can extract more information from this data, and explainable methods can make them more accessible to practitioners. The Expected Threat model has been praised for its explainability and offers an accessible option. However, selecting the grid size is a challenging key design choice that has to be made when applying the Expected Threat model. Using a finer grid leads to a more flexible model that can better distinguish between different situations, but the accuracy of the estimates deteriorates with a more flexible model. Consequently, practitioners face challenges in balancing the trade-off between model flexibility and model accuracy. In this study, the Expected Threat model is analyzed from a theoretical perspective and simulations are performed based on the Markov chain of the model to examine its behavior in practice. Our theoretical results establish an upper bound on the error of the Expected Threat model for different flexibilities. Based on the simulations, a more accurate characterization of the model's error is provided, improving over the theoretical bound. Finally, these insights are converted into a practical rule of thumb to help practitioners choose the right balance between the model flexibility and the desired accuracy of the Expected Threat model.
In this paper, we consider statistical inference for Poisson-Laguerre tessellations in & Ropf;d$$ {\mathbb{R}}<^>d $$. The object of interest is a distribution function F$$ F $$ which describes the distribution of the arrival times of the generator points. The function F$$ F $$ uniquely determines the intensity measure of the underlying Poisson process. Two nonparametric estimators for F$$ F $$ are introduced, which depend only on the points of the Poisson process that generate non-empty cells and the actual cells corresponding to these points. The proposed estimators are proven to be strongly consistent as the observation window expands unboundedly to the whole space. We also consider a stereological setting, where one is interested in estimating the distribution function associated with the Poisson process of a higher-dimensional Poisson-Laguerre tessellation, given that a corresponding sectional Poisson-Laguerre tessellation is observed.
In the uniform deconvolution problem one is interested in estimating the distribution function F_0 of a nonnegative random variable, based on a sample with additive uniform noise. A peculiar and not well understood phenomenon of the nonparametric maximum likelihood estimator in this setting is the dichotomy between the situations where F_0(1)=1 and F_0(1)<1. If F_0(1)=1, the MLE can be computed in a straightforward way and its asymptotic pointwise behavior can be derived using the connection to the so-called current status problem. However, if F_0(1)<1, one needs an iterative procedure to compute it and the asymptotic pointwise behavior of the nonparametric maximum likelihood estimator is not known. In this paper we describe the problem, connect it to interval censoring problems and a more general model studied in Groeneboom (2024) to state two competing naturally occurring conjectures for the case F_0(1)<1. Asymptotic arguments related to smooth functional theory and extensive simulations lead us to to bet on one of these two conjectures.
We construct bootstrap confidence intervals for a monotone regression function. It has been shown that the ordinary nonparametric bootstrap, based on the nonparametric least squares estimator (LSE) , is inconsistent in this situation. We show that an ‐consistent bootstrap can be based on the smoothed , to be called the SLSE (Smoothed Least Squares Estimator). The asymptotic pointwise distribution of the SLSE is derived. The confidence intervals, based on the smoothed bootstrap, are compared to intervals based on the (not necessarily monotone) Nadaraya Watson estimator and the effect of Studentization is investigated. We also give a method for automatic bandwidth choice, correcting work in Sen and Xu (2015). Analogous methods for constructing confidence intervals in the current status model are discussed, improving on work in Groeneboom and Hendrickx (2018).
The prediction of remaining useful life (RUL) is a critical component of prognostic and health management for industrial systems. In recent decades, there has been a surge of interest in RUL prediction based on degradation data of a well-defined degradation index (DI). However, in many real-world applications, the DI may not be readily available and must be constructed from complex source data, rendering many existing methods inapplicable. Motivated by multivariate sensor data from industrial induction motors, this paper proposes a novel prognostic framework that develops a nonlinear DI, serving as an ensemble of representative features, and employs a similarity-based method for RUL prediction. The proposed framework enables online prediction of RUL by dynamically updating information from the in-service unit. Simulation studies and a case study on three-phase industrial induction motors demonstrate that the proposed framework can effectively extract reliability information from various channels and predict RUL with high accuracy.
Often the question arises whether can be predicted based on using a certain model. Especially for highly flexible models such as neural networks one may ask whether a seemingly good prediction is actually better than fitting pure noise or whether it has to be attributed to the flexibility of the model. This paper proposes a rigorous permutation test to assess whether the prediction is better than the prediction of pure noise. The test avoids any sample splitting and is based instead on generating new pairings of . It introduces a new formulation of the null hypothesis and rigorous justification for the test, which distinguishes it from the previous literature. The theoretical findings are applied both to simulated data and to sensor data of tennis serves in an experimental context. The simulation study underscores how the available information affects the test. It shows that the less informative the predictors, the lower the probability of rejecting the null hypothesis of fitting pure noise and emphasizes that detecting weaker dependence between variables requires a sufficient sample size.
We consider nonparametric estimation in Wicksell's problem which has relevant applications in astronomy for estimating the distribution of the positions of the stars in a galaxy given projected stellar positions and in material sciences to determine the 3D microstructure of a material, using its 2D cross sections. In the classical setting, we study the isotonized version of the plug-in estimator (IIE) for the underlying cdf $F$ of the spheres' squared radii. This estimator is fully automatic, in the sense that it does not rely on tuning parameters, and we show it is adaptive to local smoothness properties of the distribution function $F$ to be estimated. Moreover, we prove a local asymptotic minimax lower bound in this non-standard setting, with $\sqrt{\log{n}/n}$-asymptotics and where the functional $F$ to be estimated is not regular. Combined, our results prove that the isotonic estimator (IIE) is an adaptive, easy-to-compute, and efficient estimator for estimating the underlying distribution function $F$.
Consider an opaque medium which contains 3D particles. All particles are convex bodies of the same shape, but they vary in size. The particles are randomly positioned and oriented within the medium and cannot be observed directly. Taking a planar section of the medium we obtain a sample of observed 2D section profile areas of the intersected particles. In this paper the distribution of interest is the underlying 3D particle size distribution for which an identifiability result is obtained. Moreover, a nonparametric estimator is proposed for this size distribution. The estimator is proven to be consistent and its performance is assessed in a simulation study.
The past decade has seen an increased interest in human activity recognition based on sensor data. Most often, the sensor data come unannotated, creating the need for fast labelling methods. For assessing the quality of the labelling, an appropriate performance measure has to be chosen. Our main contribution is a novel post-processing method for activity recognition. It improves the accuracy of the classification methods by correcting for unrealistic short activities in the estimate. We also propose a new performance measure, the Locally Time-Shifted Measure (LTS measure), which addresses uncertainty in the times of state changes. The effectiveness of the post-processing method is evaluated, using the novel LTS measure, on the basis of a simulated dataset and a real application on sensor data from football. The simulation study is also used to discuss the choice of the parameters of the post-processing method and the LTS measure.
In the recent paper [5], a Bayesian approach for constructing confidence intervals in monotone regression problems is proposed, based on credible intervals. We view this method from a frequentist point of view, and show that it corresponds to a percentile bootstrap method of which we give two versions. It is shown that a (non-percentile) smoothed bootstrap method has better behavior and does not need correction for over- or undercoverage. The proofs use martingale methods.
In various stereological problems an $n$-dimensional convex body is intersected with an $(n-1)$-dimensional Isotropic Uniformly Random (IUR) hyperplane. In this paper the cumulative distribution function associated with the $(n-1)$-dimensional volume of such a random section is studied. This distribution is also known as chord length distribution and cross section area distribution in the planar and spatial case respectively. For various classes of convex bodies it is shown that these distribution functions are absolutely continuous with respect to Lebesgue measure. A Monte Carlo simulation scheme is proposed for approximating the corresponding probability density functions.
Coastal climate impact studies make increasing use of multi-source and multi-dimensional atmospheric and environmental datasets to investigate relationships between climate signals and the ecological response. The large quantity of numerically simulated data may, however, include redundancy, multi-colinearity and excess information not relevant to the studied processes. In such cases techniques for feature extraction and identification of latent processes prove useful. Using dimensionality reduction techniques this research provides a statistical underpinning of variable selection to study the impacts of atmospheric processes on coastal chlorophyll-a concentrations, taking the Dutch Wadden Sea as case study. Dimension reduction techniques are applied to environmental data simulated by the Delft3D coastal water quality model, the HIRLAM numerical weather prediction model and the Euro-CORDEX climate modelling experiment. The dimension reduction techniques were selected for their ability to incorporate (1) spatial correlation via multi-way methods (2), temporal correlation through Dynamic Factor Analysis, and (3) functional variability using Functional Data Analysis. The data reduction potential and explanatory value of these methods are showcased and important atmospheric variables affecting the chlorophyll-a concentration are identified. Our results indicate room for dimensionality reduction in the atmospheric variables (2 principle components can explain the majority of variance instead of 7 variables), in the chlorophyll-a time series at different locations (two characteristic patterns can describe the 10 locations), and in the climate projection scenarios of solar radiation and air temperature variables (a single principle component function explains 77% of the variation for solar radiation and 57% of the variation for air temperature). It was also found that solar radiation followed by air temperature are the most important atmospheric variables related to coastal chlorophyll-a concentration, noting that regional differences exist, for instance the importance of air temperature is greater in the Eastern Dutch Wadden Sea at Dantziggat than in the Western Dutch Wadden Sea at Marsdiep Noord. Common trends and different regional system characteristics have also been identified through dynamic factor analysis between the deeper channels and the shallower intertidal zones, where the onset of spring blooms occurs earlier. The functional analysis of climate data showed clusters of atmospheric variables with similar functional features. Moreover, functional components of Euro-CORDEX climate scenarios have been identified for radiation and temperature variables, which provide information on the dominant mode (pattern) of variation and its uncertainties. The findings suggest that radiation and temperature projections of different Euro-CORDEX scenarios share similar characteristics and mainly differ in their amplitudes and seasonal patterns, offering opportunities to construct statistical models that do not assume independence between climate scenarios but instead borrow information (“borrow strength”) from the larger pool of climate scenarios. The presented results were used in follow up studies to construct a Bayesian stochastic generator to complement existing Euro-CORDEX climate change scenarios and to quantify climate change induced trends and uncertainties in phytoplankton spring bloom dynamics in the Dutch Wadden Sea.
This study proposes a new approach to determine phenomenological or physical relations between microstructure features and the mechanical behavior of metals bridging advanced statistics and materials science in a study of the effect of hard precipitates on the hardening of metal alloys. Synthetic microstructures were created using multi-level Voronoi diagrams in order to control microstructure variability and then were used as samples for virtual tensile tests in a full-field crystal plasticity solver. A data-driven model based on Functional Principal Component Analysis (FPCA) was confronted with the classical Voce law for the description of uniaxial tensile curves of synthetic AISI 420 steel microstructures consisting of a ferritic matrix and increasing volume fractions of M23C6 carbides. The parameters of the two models were interpreted in terms of carbide volume fractions and texture using linear mixed-effects models.