Zador's celebrated theorem is a cornerstone of optimal quantisation, establishing both the weak limit of the empirical distribution of an n-point optimal quantiser in R^d and the decay rate of the associated L_s-mean quantisation error. However, for large dimensions d, observing this asymptotic behaviour demands an astronomically large sample size n, which grows super-exponentially with d. Through a detailed analysis of the quantisation problem for spherically symmetric distributions, we demonstrate that for moderate n random quantisers uniformly distributed on a sphere of suitable radius r achieve exceptional performance. The expected distortion, expressed as a triple integral, can be computed with arbitrary precision, and the optimal radius r can be efficiently determined numerically. Leveraging results from extreme-value theory, we derive approximations for r, particularly in scenarios where n scales with d. Depending on the growth rate of n, r may either converge to zero or approach a limiting value that is independent of s.
Real-time bidding has transformed the digital advertising landscape, allowing companies to buy website advertising space in a matter of milliseconds in the time it takes a webpage to load. Joint research between Cardiff University and Crimtan has employed statistical modelling in conjunction with machine-learning techniques on big data to develop computer algorithms that can select the most appropriate person to which an ad should be shown. These algorithms have been used to identify suitable bidding strategies for that particular advert in order to make the whole process as profitable as possible for businesses. Crimtan's use of the algorithms have enabled them to improve the service that they offer to clients, save money, make significant efficiency gains and attract new business. This has had a knock-on effect with the clients themselves, who have reported an increase in conversion rates as a result of more targeted, accurate and informed advertising. We have also used mixed Poisson processes for modelling for analysing repeat-buying behaviour of online customers. To make numerical comparisons, we use real data collected by Crimtan in the process of running several recent ad campaigns.
Singular spectrum analysis (SSA) is a technique of time series analysis and forecasting combining elements of classical time series analysis, multivariate statistics, multivariate geometry, dynamical
It has been shown by Pronzato et al. (Bernoulli 23(4A):2617–2642, 2017; J Multivar Anal 168:276–289, 2018) that simplicial volumes formed by independent copies of random variables can be used to extend the definition of generalised variances. It is shown in this paper that exterior algebra is a natural environment in which to study these constructions. This is used to extend the formulation to covariances and correlations. The theory leads naturally to dispersion ordering, that is partial orderings in which one random variable is more disperse than another if one squared simplicial volume stochastically dominates the other.
This chapter starts by considering, in Sect. 1.1, properties of high-dimensional cubes and balls; we will use many of these properties in other sections of this chapter and in Chap. 3 . In Sect. 1.2, we discuss various aspects of uniformity and space-filling and demonstrate that good uniformity of a set of points is by no means implying its good space-filling. In Sect. 1.3, we consider space-filling from the viewpoint of covering and pay much attention to the concept of weak covering, where only a large part of a cube (or other set) has to be covered by the balls with centres at given points, rather than the full cube as in the standard covering. In Sect. 1.4, we provide bibliographic notes and give additional references.
The main objective of this chapter is development of accurate approximations for the volume of intersection of a cube and a ball in $$\mathbb {R}^d$$ . In Chap. 2 , the approximations developed in this chapter will be the cornerstone of more evolved approximations and methods of construction of efficient exploration strategies in global optimization and space-filling designs in high-dimensional sets.
The main developments in this chapter are:
The main purpose of the paper is to uncover the connections between kriging, energy minimization and properties of the ordinary least squares and best linear unbiased estimators in the location model with correlated observations. We emphasize the special role of the constant function and illustrate our results by several examples.
We consider global optimization problems, where the feasible region 𝒳 is a compact subset of ℝ^d with d ≥ 10 . For these problems, we demonstrate that the actual convergence of global random search algorithms is much slower than that given by the classical estimates, based on the asymptotic properties of random points, and that the usually recommended space exploration schemes are inefficient in the non-asymptotic regime. Moreover, we show that uniform sampling on entire 𝒳 is much less efficient than uniform sampling on a suitable subset of 𝒳 , and that the effect of replacement of random points by low-discrepancy sequences can be felt in small dimensions only.
A design is a collection of distinct points in a given set $X$, which is assumed to be a compact subset of $R^d$, and the mesh-ratio of a design is the ratio of its fill distance to its separation radius. The uniformity constant of a sequence of nested designs is the smallest upper bound for the mesh-ratios of the designs. We derive a lower bound on this uniformity constant and show that a simple greedy construction achieves this lower bound. We then extend this scheme to allow more flexibility in the design construction.
We consider the problem of predicting values of a random process or field satisfying a linear model y(x)=θ ^⊤ f(x) + ε (x) , where errors ε (x) are correlated. This is a common problem in kriging, where the case of discrete observations is standard. By focussing on the case of continuous observations, we derive expressions for the best linear unbiased predictors and their mean squared error. Our results are also applicable in the case where the derivatives of the process y are available, and either a response or one of its derivatives need to be predicted. The theoretical results are illustrated by several examples in particular for the popular Matérn 3/2 kernel.
The inverse square root of a covariance matrix is often desirable for performing data whitening in the process of applying many common multivariate data analysis methods. Direct calculation of the inverse square root is not available when the covariance matrix is either singular or nearly singular, as often occurs in high dimensions. We develop new methods, which we broadly call polynomial whitening , to construct a low-degree polynomial in the empirical covariance matrix which has similar properties to the true inverse square root of the covariance matrix (should it exist). Our method does not suffer in singular or near-singular settings, and is computationally tractable in high dimensions. We demonstrate that our construction of low-degree polynomials provides a good substitute for high-dimensional inverse square root covariance matrices, in both d < N and d ≥ N cases. We offer examples on data whitening, outlier detection and principal component analysis to demonstrate the performance of the proposed method.
For large classes of group testing problems, we derive lower bounds for the probability that all significant items are uniquely identified using specially constructed random designs. These bounds allow us to optimize parameters of the randomization schemes. We also suggest and numerically justify a procedure of constructing designs with better separability properties than pure random designs. We illustrate theoretical considerations with a large simulation-based study. This study indicates, in particular, that in the case of the common binary group testing, the suggested families of designs have better separability than the popular designs constructed from disjunct matrices. We also derive several asymptotic expansions and discuss the situations when the resulting approximations achieve high accuracy.
In this paper, we study the behaviour of the so-called k -simplicial distances and k -minimal-variance distances between a point and a sample. The family of k -simplicial distances includes the Euclidean distance, the Mahalanobis distance, Oja’s simplex distance and many others. We give recommendations about the choice of parameters used to calculate the distances, including the size of the sub-sample of simplices used to improve computation time, if needed. We introduce a new family of distances which we call k -minimal-variance distances. Each of these distances is constructed using polynomials in the sample covariance matrix, with the aim of providing an alternative to the inverse covariance matrix, that is applicable when data is degenerate. We explore some applications of the considered distances, including outlier detection and clustering, and compare how the behaviour of the distances is affected for different parameter choices.
Let $${\mathbb {Z}}_n = \{Z_1, \ldots , Z_n\}$$ be a design; that is, a collection of n points $$Z_j \in [-1,1]^d$$ . We study the quality of quantisation of $$[-1,1]^d$$ by the points of $${\mathbb {Z}}_n$$ and the problem of quality of coverage of $$[-1,1]^d$$ by $${{{\mathcal {B}}}}_d({\mathbb {Z}}_n,r)$$ , the union of balls centred at $$Z_j \in {\mathbb {Z}}_n$$ . We concentrate on the cases where the dimension d is not small, $$d\ge 5$$ , and n is not too large, $$n\le 2^d$$ . We define the design $${{\mathbb {D}}_{n,\delta }}$$ as a $$2^{d-1}$$ design defined on vertices of the cube $$[-\delta ,\delta ]^d$$ , $$0\le \delta \le 1$$ . For this design, we derive a closed-form expression for the quantisation error and very accurate approximations for the coverage area $${\text {vol}}{([-1,1]^d \cap {{{\mathcal {B}}}}_d({\mathbb {Z}}_n,r))}$$ . We provide results of a large-scale numerical investigation confirming the accuracy of the developed approximations and the efficiency of the designs $${{\mathbb {D}}_{n,\delta }}$$ .