
We consider a one-dimensional, nearest-neighbour, recurrent random walk $(X_k)$, $k\leq n$, in a random i.i.d. environment (RWRE). Assuming that the environment has finite support, which is treated as a parameter, we establish the Local Asymptotic Mixed Normality (LAMN) property for this parameter. The convergence rate is $\sqrt n$, and the asymptotic Fisher information is random and expressed in terms of the invariant measure in the infinite valley, as introduced in Gantert et al., 2010. We further show that the Maximum Likelihood Estimator (MLE) of the support parameter converges to a mixture of normal distributions and is asymptotically efficient. We further show that the Maximum Likelihood Estimator (MLE) of the support parameter converges to a mixture of normal distributions and is asymptotically efficient. The proofs are based on a recent result by Comets et al.,2024, which extends to recurrent RWRE the method of the ``environment viewed from the particle'', originally introduced in \cite{KozMol} for the transient ballistic RWRE.
Calibration of sensors in partially observed and correlated environments raises fundamental challenges for variable selection and model interpretation. When observations are noisy and influenced by unmeasured factors, classical selection criteria based on regression coefficients, cross-validation errors, or sensitivity indices may fail to identify the variables that most effectively reduce uncertainty in the target quantity. This paper introduces a probabilistic framework for variable selection based on variance reduction and prediction stability. The proposed approach relies on the systematic evaluation of models built from different subsets of observed variables and on a criterion that minimizes conditional prediction variance under a parsimony constraint. This criterion naturally penalizes spurious correlations and distinguishes variables that contribute to uncertainty reduction from those that merely compensate for missing information. The framework is illustrated through analytical examples and numerical experiments on simulated data, which highlight its robustness to noise and unmeasured confounders. Although motivated by calibration problems in environmental sensing, the proposed methodology is general and applicable to a wide range of regression and inference tasks involving correlated inputs and unobserved factors.
Numerical simulation is widely used in many fields of engineering to study complex physical systems. The numerical models, designed to faithfully represent the underlying physical phenomena, are subject to uncertainties of different natures (either numerical, stochastic or epistemic) that degrade the accuracy of the simulated outputs. Part of epistemic uncertainty arises from limited knowledge regarding some input model parameters theta. This component can be reduced through Bayesian calibration of the model against experimental data. Before calibration itself, sensitivity analysis can be used to better understand how parameter uncertainties impact the model output, and this may help confine calibration to the most impactful parameters. In this work, we show that kernel methods, especially those based on the Hilbert-Schmidt independence criterion (HSIC), are effective tools in support of Bayesian calibration, both for a single model and for two chained models. In the latter case, our main contribution is a screening methodology for the parameters theta of the downstream model, which accounts for the posterior distribution of the upstream model parameters lambda. By taking the expectation of the HSIC over lambda, we define a new sensitivity measure that is able to incorporate the residual uncertainty due to the upstream model calibration. We show that the resulting sensitivity indices can be estimated from the same data used for conditional Bayesian calibration. We further demonstrate that the corresponding estimators are consistent and achieve convergence rates comparable to those of classical Monte Carlo estimators. Importantly, we construct two test procedures that enable rigorous decisions on which parameters among theta should be selected. Finally, we apply the proposed approach to nuclear fuel simulation to screen the calibration parameters theta of a fission gas behavior model which follows an upstream thermal model whose thermal conductivity lambda was calibrated in previous work.
Explaining the outcome of dynamical systems is non-trivial due to the temporal nature and correlation of the variables. In this work, we propose a novel framework of history-aware sensitivity analysis for stationary time-series to quantify different memory effects and clarify their roles. For this purpose, we decompose the output time series into non-correlated components, namely the instantaneous component and the memory components. The latter are sorted in decreasing order of variance to reflect the importance of the variables. We highlight the compensation phenomena between the resulting components and illustrate them in the case of independent variables in a linear setting. To enable history-aware explanations, variance-based sensitivity indices are derived from the obtained decomposition. We demonstrate the effectiveness of our methodology in providing insights to explain output time-series in both synthetic and real-world cases.
Nonparametric regression analysis has broad applications. In some cases, the regression function with jumps (i.e., the regression curve is discontinuous) seems to be more appropriate to describe the related phenomena. A number of methods exist for estimating discontinuous curve, most of which are based on complete data, which is unrealistic in many practical situations. In this paper, we consider estimating discontinuous nonparametric model with covariate with missing values. Based on inverse selection probability weighted and jump-preserving techniques, a jump-preserving estimation procedure is proposed. The proposed method is capable of automatically accommodating possible jumps in the nonparametric function, without the requirement of prior knowledge regarding the number and locations of jump points. The proposed estimator for the discontinuous regression function is shown to be oracally efficient in the sense that it is uniformly indistinguishable from that when the selection probabilities are known. Furthermore, it is proved that the fitted curve by this procedure is consistent in the entire design space. Numerical simulation also indicates the finite sample performance of this method is efficient and reliable.
This paper considers a continuous-time bidimensional risk model with a constant interest force, where there exist dependence structures among the claim sizes of two business lines and the inter-arrival times of the claim sizes. When the claim sizes have subexponential distributions, some uniform asymptotics for the finite-time absolute ruin probabilities of the bidimensional risk model are established.
A faithful bi-interweaving relation is a Markovian similarity-type relation between two Markov chains, strengthening the Markovian intertwining relation and introducing warming-up times after which the time-marginal distributions of the chains can be tightly compared (for any initial distributions). For irreducible transition kernels on the same finite state space, these relations are shown to be equivalent to the generalised isospectrality relation, but this is no longer true for nontransient transition kernels, contrary to the faithful bi-intertwining relations. Some bounds are deduced on corresponding warming-up times, when the eigenvalues are furthermore assumed to be real (but still allowing for Jordan blocks). When the eigenvalues are non-negative, the same approach enables us to construct strong stationary times for irreducible Markov chains through interweaving relations with model absorbed Markov chains, thus extending a result due to Matthews in the reversible situation.
We address the problem of estimating multiple modes of a multivariate density using persistent homology, a central tool in Topological Data Analysis. We introduce a method based on the preliminary estimation of the H0-persistence diagram to infer the number of modes, their locations, and the corresponding local maxima. For broad classes of piecewise-continuous functions with geometric control on discontinuity loci, we identify a critical separation threshold between modes, equiv-alently interpretable in our framework in terms of modes' prominence, below which modes inference is impossible and above which our procedure achieves minimax optimal rates.
This work addresses the interpolation of probability measures within a spatial statistics framework. We develop a Kriging approach in the Wasserstein space, leveraging the quantile function representation of the one-dimensional Wasserstein distance. To mitigate the inaccuracies in semivariogram estimation that arise from sparse datasets, we combine this formulation with cross-validation techniques. In particular, we introduce a variant of the virtual cross-validation formulas tailored to quantile functions. The effectiveness of the proposed method is demonstrated on a controlled toy problem as well as on a real-world application from nuclear safety.
We propose a new importance sampling framework for the estimation and analysis of Sobol' indices. We show that a Sobol' index defined under a reference input distribution can be consistently estimated from samples drawn from other sampling distributions by reweighting the estimator appropriately to account for the distribution change. We derive the optimal sampling distribution that minimizes the asymptotic variance and demonstrate its strong impact on estimation accuracy. Beyond variance reduction, the framework supports distributional sensitivity analysis via reverse importance sampling, enabling robust exploration of input distribution uncertainty with negligible additional computational cost.
We study a random walk on the subgroup of lower triangular matrices of SL_2, with i.i.d. increments. We prove that the process of the lower corner of the random walk satisfies a Rogers-Pitman criterion to be a Markov chain if and only if the increments are distributed according to a Generalized Inverse Gaussian (GIG) law on their diagonals. For this, we prove a new characterization of these laws. We prove a discrete-time version of the Dufresne identity. We show how to recover the Matsumoto-Yor theorem by taking the continuous limit of the random walk.
We introduce the concept of an imprecise Markov semigroup 𝐐. It is a tool that allows us to represent ambiguity around both the transition probabilities and the invariant measure of a continuous-time Markov process via a collection of Markov semigroups, each associated with a (possibly different) Markov process. We use techniques from topology, geometry, and probability to analyze ergodic limits under model uncertainty encoded by 𝐐. We establish long-term bounds that are uniform in the initial state and identify regimes in which the imprecision in these bounds collapses asymptotically. Our results are proved in progressively more general settings. We first assume that 𝐐 is compact and that the state space is Euclidean or a Riemannian manifold, working with a fixed bounded observable. We then allow the state space to be standard Borel, while keeping 𝐐 compact and the observable fixed. Finally, we drop compactness and work on Polish metric spaces of finite diameter, where we treat arbitrary bounded Lipschitz observables. The importance of our findings for the fields of artificial intelligence and computer vision is also discussed at a high level; In particular, in the study of how the probability of an output evolves over time as we perturb the input of a convolutional autoencoder.
We present an orthogonal expansion for real, function-regulated, second-order random measures over ℝd with measure covariance. Such an expansion, which can be seen as a Karhunen– Loève decomposition, consists in a series of deterministic real measures weighted by uncorrelated real random variables with the variances forming a convergent series. The convergence of the series is in a mean-square sense stochastically and against measurable and bounded test functions (with compact support if the random measure is not finite) in the measure sense, which implies set-wise convergence. This is proven taking advantage of the extra requirement of having a covariance measure over ℝd × ℝd describing the covariance structure of the random measure, for which we also provide a series expansion. These results cover for instance the cases of Gaussian White Noise, Poisson and Cox point processes, and can be used to obtain expansions for trawl processes.
Consider a random graph $G$ of size $N$ constructed according to a \textit{graphon} $w \, : \, [0,1]^{2} \mapsto [0,1]$ as follows. First embed $N$ vertices $V = \{v_1, v_2, \ldots, v_N\}$ into the interval $[0,1]$, then for each $i < j$ add an edge between $v_{i}, v_{j}$ with probability $w(v_{i}, v_{j})$. Given only the adjacency matrix of the graph, we might expect to be able to approximately reconstruct the permutation $\sigma$ for which $v_{\sigma(1)} < \ldots < v_{\sigma(N)}$ if $w$ satisfies the following \textit{linear embedding} property introduced in [Janssen 2019]: for each $x$, $w(x,y)$ decreases as $y$ moves away from $x$. For a large and non-parametric family of graphons, we show that (i) the popular spectral seriation algorithm [Atkins 1998] provides a consistent estimator $\hat{\sigma}$ of $\sigma$, and (ii) a small amount of post-processing results in an estimate $\tilde{\sigma}$ that converges to $\sigma$ at a nearly-optimal rate, both as $N \rightarrow \infty$.
Under various conditions, sewing lemmas provide convergence of the Riemann-type sum Sigma[s,t]is an element of pi Xi s,t for a given two-parametric map Xi as the mesh sizes of the considered partitions pi tend to zero. In this note, we prove a stochastic sewing lemma for two-parameter processes whose increments, when viewed as functions with values in Lm(Omega;V) for m >= 2 and a real separable Banach space V with a non-trivial martingale type, are of Besov regularity. The contribution is two-fold: First, we generalize the stochastic sewing lemma of L & ecirc; [Electron. J. Probab. 25 (2020) 1-55] for processes whose increments belong to a Besov and not necessarily H & ouml;lder space. Second, we show here that the assumptions of the Besov sewing lemma of Friz et al. [J. Differ. Equ. 339 (2022) 152-231] can be relaxed if stochastics is incorporated in the sewing from the beginning. As an application, Besov regularity of the Ito integral of Brownian functionals is obtained.
In this paper we develop a representation formula of Clark-Ocone type for any integrable Poisson functionals, which extends the Poisson imbedding for point processes. This repre- sentation formula differs from the classical Clark-Ocone formula on three accounts. First the representation holds with respect to the Poisson measure instead of the compensated one; second the representation holds true in L1 and not in L2; and finally contrary to the classical Clark-Ocone formula the integrand is defined as a pathwise operator and not as a L2-limiting object. We make use of Malliavin’s calculus and of a decomposition with uncompensated iterated integrals derived in [HR24] to establish this non-compensated Clark-Ocone representation formula and to characterize the integrand, which turns out to be a predictable integrable process.
The term noncentral moderate deviations is used in the literature to mean a class of large deviation principles that, in some sense, fills the gap between the convergence in probability to a constant (governed by a reference large deviation principle) and a weak convergence to a non-Gaussian (and non-degenerating) distribution. Some noncentral moderate deviation results in the literature concern time-changed univariate Lévy processes, where the time-changes are given by inverse stable subordinators. In this paper we present analogue results for multivariate Lévy processes; in particular the random time-changes are suitable linear combinations of independent inverse stable subordinators.
In this note, we are interested in the probability that two independent squared Bessel processes do not cross for a long time. We show that this probability has a power decay which is given by the first zero of some hypergeometric function. We also compute along the way the distribution of the location where the crossing eventually occurs.
In this paper, using the excursion theory, we provide new proofs of, firstly, Lehoczky’s formula (in an extended form allowing a lower bound for the underlying diffusion) for the joint distribution of the first drawdown time and the maximum before this time, and, secondly, of Malyutin’s formula for the joint distribution of the first hitting time and the maximum drawdown before this time. It is remarkable – but there is a clean explanation – that the excursion theoretical approach which we developed first for Lehoczky’s formula also provides a proof for Malyutin’s formula. Moreover, we discuss some generalizations and analyze the pure jump process describing the maximum before the first drawdown time when the size of the drawdown is varying.
We study the nonparametric estimation of both the potential and the interaction terms of a scalar McKean-Vlasov stochastic differential equation (SDE) in stationary regime from a continuous observation on a time interval [0, T], with asymptotic framework T -> +infinity. To estimate the two functions, we consider the observation of four i.i.d. sample paths. The observation of two sample paths could be enough at the cost of much more computations. Estimators of the potential and the interaction functions are built using a combination of a moment method and a projection method on sieves. The potential and the interaction term do not belong to L-2(R), so we define a specific risk fitted to this estimation problem and obtain a bound for it. A nonparametric estimator of the invariant density also is proposed. The method is implemented on simulated data for several examples of McKean-Vlasov SDEs and a model selection procedure is experimented.