Solving large scale Optimal Transport (OT) in machine learning typically relies on sampling measures to obtain a tractable discrete problem. While the discrete solver's accuracy is controllable, the rate of convergence of the discretization error is governed by the intrinsic dimension of our data. Therefore, the true bottleneck is the knowledge and control of the sampling error. In this work, we tackle this issue by introducing novel estimators for both sampling error and intrinsic dimension. The key finding is a simple, tuning-free estimator of OT_c(ρ, ) that utilizes the semi-dual OT functional and, remarkably, requires no OT solver. Furthermore, we derive a fast intrinsic dimension estimator from the multi-scale decay of our sampling error estimator. This framework unlocks significant computational and statistical advantages in practice, enabling us to (i) quantify the convergence rate of the discretization error, (ii) calibrate the entropic regularization of Sinkhorn divergences to the data's intrinsic geometry, and (iii) introduce a novel, intrinsic-dimension-based Richardson extrapolation estimator that strongly debiases Wasserstein distance estimation. Numerical experiments demonstrate that our geometry-aware pipeline effectively mitigates the discretization error bottleneck while maintaining computational efficiency.
We consider a regularly varyingstationary sequenceof random variables (X-t) with tail index alpha < 2. For this sequence we study the joint convergenceof sums, l(p)-type moduli and maxima. We focus on ratio statistics, including the studentized sums and sums normalized by the corresponding maxima, and study the existence of moments for the limit ratios. We consider particular examples of processes (X-t) whose limit ratios possess all moments as in the iid setting. But, in contrast to the latter situation, there also exist dependent sequences (X-t) where certain moments of the limit ratio are infinite. This phenomenon results from extremal clusters in the sequence.
The Kalman Filter (KF) is used to estimate and predict state-space models. It frequently encounters uncertain covariates, such as those derived from noisy measurements of physical signals (e.g., sensor readings, weather data, etc.). Our aim is to account for their uncertainties and develop robust KF approaches when incorporating the uncertainties of the covariates into the model; that is, by treating them as stochastic variables with known variances. To enhance the robustness of the KF viewed as a Bayesian method that estimates the posterior distribution of the system state conditioned on past observations, we propose two methods. The linear Bayes estimator yields an explicit, unbiased forecast with minimal variance. The optimal risk estimator is nonlinear and remains robust even when the covariates are highly noisy. This second method is approximated using a sequential Monte Carlo technique.
We study predictive probability inference in classification tasks using random forests under class imbalance. We focus on two simplified variants of Breiman's algorithm, namely subsampling Infinite Random Forests (IRFs) and under-sampling IRFs, and establish their asymptotic normality. In the under-sampling setting, training data from both classes are re-sampled to achieve balance, which enhances minority class representation but introduces a biased model. To correct this, we propose a debiasing procedure based on Importance Sampling (IS) using odds ratios. We instantiate our results using 1-Nearest Neighbor (1-NN) classifiers as base learners in the IRFs and prove the near-minimax optimality of the approach for Lipschitz continuous objectives. We also show that the IS bagged 1-NN estimator matches the convergence rate of its subsampled counterpart while attaining lower asymptotic variance in most cases. Our theoretical findings are supported by simulation studies, highlighting the empirical benefits of the proposed approach.
Abstract Many numerical weather prediction models and their associated postprocessed models are available. Combining all of these predictions in an optimal way is, however, not straightforward. This can be achieved thanks to expert aggregation (EA) which has many advantages, such as being online, being adaptive to model changes, and having theoretical guarantees. In this paper, we propose a method for making deterministic temperature predictions with EA. We used exponentially weighted average, MLprod and MLpol, and Bernstein online aggregation. Hence, we combine and outperform the forecasts of the raw and postprocessed Integrated Forecasting System (IFS), forecasts of Applications of Research to Operations at Mesoscale (AROME), Action de Recherche Petite Echelle Grande Echelle (ARPEGE), and quantiles of the postprocessed Prévision d’Ensemble ARPEGE (PEARP). We also compare the different EA strategies in various settings and show that they outperform the National Blend of Models. Finally, we discuss certain limitations. Significance Statement There are many different weather forecasting models nowadays, so it can be difficult to choose a prediction. The aim of this study is to combine several temperature forecasts in an optimal way. To achieve this, we use a framework known as “expert aggregation” and compare various algorithms. Our results show that this framework can help to improve the temperature predictions on average. However, the challenge of efficiently utilizing experts that perform poorly on average but may be accurate at specific times remains.
Abstract In this paper, we improve on the temperature predictions made with online Expert Aggregation (EA) in a previous work. In particular, we make the aggregation more reactive, whilst improving the root mean squared error and reducing the number of large errors. We have achieved this by using the Sleeping Expert Framework (SEF), which allows a more efficient use of biased experts (bad on average but which may be good at some point). To deal with the fact that we do not know in advance when to use these biased experts, we resorted to Gradient Boosted Regression Trees (GBRT). We applied the SEF with GBRT in a fully online way on Bernstein Online Aggregation (BOA), an adaptive aggregation with second-order regret bounds.
We study nonparametric regression over Besov spaces from noisy observations under sub-exponential noise, aiming to achieve minimax-optimal guarantees on the integrated squared error that hold with high probability and adapt to the unknown noise level. To this end, we propose a wavelet-based online learning algorithm that dynamically adjusts to the observed gradient noise by adaptively clipping it at an appropriate level, eliminating the need to tune parameters such as the noise variance or gradient bounds. As a by-product of our analysis, we derive high-probability adaptive regret bounds that scale with the ℓ_1-norm of the competitor. Finally, in the batch statistical setting, we obtain adaptive and minimax-optimal estimation rates for Besov spaces via a refined online-to-batch conversion. This approach carefully exploits the structure of the squared loss in combination with self-normalized concentration inequalities.
We present the winning strategy for the EVA2025 Data Challenge, which aimed to estimate the probability of extreme precipitation events. These events occurred at most once in the dataset making the challenge fundamentally one of extrapolating extreme values. Given the scarcity of extreme events, we argue that a simple, robust modeling approach is essential. We adopt univariate models instead of multivariate ones and model Peaks Over Thresholds using Extreme Value Theory. Specifically, we fit an exponential distribution to model exceedances of the target variable above a high quantile (after seasonal adjustment). The novelty of our approach lies in using martingale testing to evaluate the extrapolation power of the procedure and to agnostically select the level of the high quantile. While this method has several limitations, we believe that framing extrapolation as a game opens the door to other agnostic approaches in Extreme Value Analysis.
We study online adversarial regression with convex losses against a rich class of continuous yet highly irregular competitor functions,% prediction rules, modeled by Besov spaces $B_{pq}^s$ with general parameters $1 \leq p,q \leq \infty$ and smoothness $s > \tfrac{d}{p}$. We introduce an adaptive wavelet-based algorithm that performs sequential prediction without prior knowledge of $(s,p,q)$, and establish minimax-optimal regret bounds against any comparator in $B_{pq}^s$. We further design a locally adaptive extension capable of sequentially adapting to spatially inhomogeneous smoothness. This adaptive mechanism adjusts the resolution of the predictions over both time and space, yielding refined regret bounds in terms of local regularity. Consequently, in heterogeneous environments, our adaptive guarantees can significantly surpass those obtained by standard global methods.
Extremes occur in stationary regularly varying time series as short periods with several large observations, known as extremal blocks. We study cluster statistics summarizing the behavior of functions acting on these extremal blocks. Examples of cluster statistics are the extremal index, cluster size probabilities, and other cluster indices. The purpose of our work is twofold. First, we state the asymptotic normality of block estimators for cluster inference based on consecutive observations with large ℓ ^p- norms, for p > 0 . The case p=α , where α > 0 is the tail index of the time series, has specific nice properties thus we analyze the asymptotics of block estimators when approximating α using the Hill estimator. Second, we verify the conditions we require on classical models such as linear models and solutions of stochastic recurrence equations. Regarding linear models, we prove that the asymptotic variance of classical index cluster-based estimators is null as first conjectured in Hsing (Probab. Theory Related Fields 95, 331–356 1993). We illustrate our findings on simulations.
Nonlinear and delayed effects of covariates often render time series forecasting challenging. To this end, we propose a novel forecasting framework based on ridge regression with signature features calculated on sliding windows. These features capture complex temporal dynamics without relying on learned or hand-crafted representations. Focusing on the discrete-time setting, we establish theoretical guarantees, namely universality of approximation and stationarity of signatures. We introduce an efficient sequential algorithm for computing signatures on sliding windows. The method is evaluated on both synthetic and real electricity demand data. Results show that signature features effectively encode temporal and nonlinear dependencies, yielding accurate forecasts competitive with those based on expert knowledge.
Many Numerical Weather Prediction models and their associated Post-Processed Models are available. Combining all of these predictions in an optimal way is however not straightforward. This can be achieved thanks to Expert Aggregation (EA) which has many advantages, such as being online, being adaptive to model changes and having theoretical guarantees. In this paper, we propose a method for making deterministic temperature predictions with EA. We used Exponentially Weighted Average, MLprod and MLpol and Bernstein Online Aggregation. Hence, we combine and outperform the forecasts of the raw and post-processed Integrated Forecasting System (IFS), forecasts of Application of Research to Operations at Mesoscale (AROME), Action de Recherche Petite Echelle Grande Echelle (ARPEGE) and quantiles of the post processed Prévision d'Ensemble ARPEGE (PEARP). We also compare the different EA strategies in various settings and show that they outperform the National Blend of Models. Finally, we discuss certain limitations.
A multitude of Numerical Weather Prediction (NWP) models, along with their associated Model Output Statistics (MOS), are readily available. Expert Aggregation (EA) algorithms combine them in an online and adaptive manner. While EA competes optimally against the best-fixed combination of experts (Wintenberger 2017), it falls short in handling rapid changes. We introduce the class of Markov-EA algorithms, extending the seminal work of Mourtada and Maillard (2017) on Exponentiated Weights to other EA algorithms such as BOA and ML-Poly. Understanding how and when to adjust the weights is crucial for obtaining optimal second-order regret bounds. Assuming a (non-homogeneous) Markovian dynamic, we enhance the EA predictions of short and poorly predicted events, such as the cold event in the Chamonix valley, using weight sharing and strategies involving sleeping experts. This work is done in collaboration with Leo Pfitzner and Olivier Mestre (Météo France).
We introduce an online mathematical framework for survival analysis, allowing real time adaptation to dynamic environments and censored data. This framework enables the estimation of event time distributions through an optimal second order online convex optimization algorithm-Online Newton Step (ONS). This approach, previously unexplored, presents substantial advantages, including explicit algorithms with non-asymptotic convergence guarantees. Moreover, we analyze the selection of ONS hyperparameters, which depends on the exp-concavity property and has a significant influence on the regret bound. We introduce an adaptive aggregation method that ensures robustness in hyperparameter selection while maintaining fast regret bounds. These findings can extend beyond the survival analysis field, and are relevant for any case characterized by poor exp-concavity and unstable ONS. Additionally, we propose a stochastic approach for ONS that guarantees logarithmic regret in the case of an exponential hazard model. Next, these assertions are illustrated by simulation experiments, followed by an application to a real dataset.
We investigate the semi-discrete Optimal Transport (OT) problem, where a continuous source measure $\mu$ is transported to a discrete target measure $\nu$, with particular attention to the OT map approximation. In this setting, Stochastic Gradient Descent (SGD) based solvers have demonstrated strong empirical performance in recent machine learning applications, yet their theoretical guarantee to approximate the OT map is an open question. In this work, we answer it positively by providing both computational and statistical convergence guarantees of SGD. Specifically, we show that SGD methods can estimate the OT map with a minimax convergence rate of $\mathcal{O}(1/\sqrt{n})$, where $n$ is the number of samples drawn from $\mu$. To establish this result, we study the averaged projected SGD algorithm, and identify a suitable projection set that contains a minimizer of the objective, even when the source measure is not compactly supported. Our analysis holds under mild assumptions on the source measure and applies to MTW cost functions,whic include $\|\cdot\|^p$ for $p \in (1, \infty)$. We finally provide numerical evidence for our theoretical results.
Many classification tasks involve imbalanced data, in which a class is largely underrepresented. Several techniques consists in creating a rebalanced dataset on which a classifier is trained. In this paper, we study theoretically such a procedure, when the classifier is a Centered Random Forests (CRF). We establish a Central Limit Theorem (CLT) on the infinite CRF with explicit rates and exact constant. We then prove that the CRF trained on the rebalanced dataset exhibits a bias, which can be removed with appropriate techniques. Based on an importance sampling (IS) approach, the resulting debiased estimator, called IS-ICRF, satisfies a CLT centered at the prediction function value. For high imbalance settings, we prove that the IS-ICRF estimator enjoys a variance reduction compared to the ICRF trained on the original data. Therefore, our theoretical analysis highlights the benefits of training random forests on a rebalanced dataset (followed by a debiasing procedure) compared to using the original data. Our theoretical results, especially the variance rates and the variance reduction, appear to be valid for Breiman's random forests in our experiments.
In this paper we improve on the temperature predictions made with (online) Expert Aggregation (EA) [Cesa-Bianchi and Lugosi, 2006] in Part I. In particular, we make the aggregation more reactive, whilst maintaining at least the same root mean squared error and reducing the number of large errors. We have achieved this by using the Sleeping Expert Framework (SEF) [Freund et al., 1997, Devaine et al., 2013], which allows the more efficient use of biased experts (bad on average but which may be good at some point). To deal with the fact that, unlike in Devaine et al. [2013], we do not know in advance when to use these biased experts, we resorted to gradient boosted regression trees [Chen and Guestrin, 2016] and provide regret bounds against sequences of experts [Mourtada and Maillard, 2017] which take into account this uncertainty. We applied this in a fully online way on BOA [Wintenberger, 2024], an adaptive aggregation with second order regret bounds, which had the best results in Part I. Finally, we made a meta-aggregation with the EA follow the leader. This chooses whether or not to use the SEF in order to limit the possible noise added by the SEF.
We study the joint limit behavior of sums, maxima and $\ell^p$-type moduli for samples taken from an $\mathbb{R}^d$-valued regularly varying stationary sequence with infinite variance. As a consequence, we can determine the distributional limits for ratios of sums and maxima, studentized sums, and other self-normalized quantities in terms of hybrid characteristic functions and Laplace transforms. These transforms enable one to calculate moments of the limits and to characterize the differences between the iid and stationary cases in terms of indices which describe effects of extremal clustering on functionals acting on the dependent sequence.
Adding entropic regularization to Optimal Transport (OT) problems has become a standard approach for designing efficient and scalable solvers. However, regularization introduces a bias from the true solution. To mitigate this bias while still benefiting from the acceleration provided by regularization, a natural solver would adaptively decrease the regularization as it approaches the solution. Although some algorithms heuristically implement this idea, their theoretical guarantees and the extent of their acceleration compared to using a fixed regularization remain largely open. In the setting of semi-discrete OT, where the source measure is continuous and the target is discrete, we prove that decreasing the regularization can indeed accelerate convergence. To this end, we introduce DRAG: Decreasing (entropic) Regularization Averaged Gradient, a stochastic gradient descent algorithm where the regularization decreases with the number of optimization steps. We provide a theoretical analysis showing that DRAG benefits from decreasing regularization compared to a fixed scheme, achieving an unbiased $\mathcal{O}(1/t)$ sample and iteration complexity for both the OT cost and the potential estimation, and a $\mathcal{O}(1/\sqrt{t})$ rate for the OT map. Our theoretical findings are supported by numerical experiments that validate the effectiveness of DRAG and highlight its practical advantages.
We establish sample-path large deviation principles for the centered cumulative functional of marked Poisson cluster processes in the Skorokhod space equipped with the M1 topology, under joint regular variation assumptions on the marks and the offspring distributions governing the propagation mechanism. These findings can also be interpreted as hidden regular variation of the cluster processes' functionals, extending the results in Dombry et al. (2022) to cluster processes with heavy-tailed characteristics, including mixed Binomial Poisson cluster processes and Hawkes processes. Notably, by restricting to the adequate subspace of measures on D([0, 1], R+), and applying the correct normalization and scaling to the paths of the centered cumulative functional, the limit measure concentrates on paths with multiple large jumps.