Motivated by the preeminence of Poisson point processes in various imaging modalities, we address the problem of recovering the intensity of a filtered Poisson point process on Rd. That is, we observe the convolution of the empirical measure underlying the point process with a known filter h. We investigate both the case where the full process is continuously observed, or the more realistic setting where it is observed on a grid. Upon choosing an orthornormal basis, we propose two adaptive nonparametric estimators of the intensity, either through a model selection approach or a thresholding procedure, with a focus on the case of a tensorized Hermite basis. We provide convergence rates for both observation schemes based on appropriate concentration inequalities for Poisson processes. Extensive and elaborate numerical experiments show the implementability of the procedures and illustrate the theoretical results.
We consider nonparametric estimation of the regression function in a model where individuals share a common noise component and repeated measurements are available for each individual. We propose a projection estimator which minimizes a least-squares contrast that accounts for the covariance structure resulting from the common noise. We analyze its risk measured either as the expectation of the empirical norm or as the expectation of the theoretical norm associated with the contrast. We discuss how the number of repeated measurements affects the estimation rates in the common noise model, and precisely characterize the dependence on the number of repetitions. In addition, we propose a data-driven projection estimator and establish risk bounds in terms of the expected empirical norm. The results are illustrated with some simulation experiments.
This work is concerned with the problem of estimating the p-fold convolution of the densities of p independent random vectors in R-d. Two nonparametric estimators are proposed, a kernel and a projection estimator, and their integrated quadratic risk is studied. We use Fourier analysis to bound the variance and consider Sobolev classes to discuss the convergence rates for both cases. In addition, we propose a fully data-driven kernel estimator based on a thresholding procedure and study model selection for the projection estimator. Finally, we illustrate the results in simulation experiments.
We study the nonparametric estimation of both the potential and the interaction terms of a scalar McKean-Vlasov stochastic differential equation (SDE) in stationary regime from a continuous observation on a time interval [0, T], with asymptotic framework T -> +infinity. To estimate the two functions, we consider the observation of four i.i.d. sample paths. The observation of two sample paths could be enough at the cost of much more computations. Estimators of the potential and the interaction functions are built using a combination of a moment method and a projection method on sieves. The potential and the interaction term do not belong to L-2(R), so we define a specific risk fitted to this estimation problem and obtain a bound for it. A nonparametric estimator of the invariant density also is proposed. The method is implemented on simulated data for several examples of McKean-Vlasov SDEs and a model selection procedure is experimented.
We consider N independent and identically distributed (i.i.d.) stochastic processes (Xj(t),t∈[0,T]), j∈{1,…,N}, defined by a one-dimensional stochastic differential equation (SDE) with time-dependent drift and diffusion coefficient. In this context, the nonparametric estimation of a general drift function b(t,x) from a continuous observation of the N sample paths on [0,T] has never been investigated. Considering a set Iϵ=[ϵ,T]×A, with ϵ≥0 and A⊂R, we build by a projection method an estimator of b on Iϵ. As the function is bivariate, this amounts to estimating a matrix of projection coefficients instead of a vector for univariate functions. We make use of Kronecker products, which simplifies the mathematical treatment of the problem. We study the risk of the estimator and distinguish the case where ϵ=0 and the case ϵ>0 and A=[a,b] compact. In the latter case, we investigate rates of convergence and prove a lower bound showing that our estimator is minimax. We propose a data-driven choice of the projection space dimension leading to an adaptive estimator. Examples of models and numerical simulation results are proposed. The method is easy to implement and works well, although computationally slower than for the estimation of a univariate function.
We assume that we observe N independent copies of a diffusion process on a time-interval [0,2T]. For a given time t, we estimate the transition density p_t(x,y), namely the conditional density of X_t + s given X_s = x, under conditions on the diffusion coefficients ensuring that this quantity exists. We use a least squares projection method on a product of finite dimensional spaces, prove risk bounds for the estimator and propose an anisotropic model selection method, relying on several reference norms. A simulation study illustrates the theoretical part for Ornstein-Uhlenbeck or square-root (Cox-Ingersoll-Ross) processes.
Consider an additive functional regression model where a one-dimensional response process Y(t) and a Kdimensional explanatory random process Xj(t), j = 1,.. ., K, are observed fort E [0, T], with fixed T. Examples of such explanatory processes are continuous or inhomogeneous counting processes. The linear coefficients of the model are K unknown deterministic functions t -> bj(t), j = 1,.. ., K, t E [0, T]. From N independent trajectories, we build a nonparametric least-squares estimator ( b1, ... , bK) of (b1, ... , bK), where each bj, 1 <= j <= K, is given by its expansion on a finite-dimensional space. We prove a bound on the mean-square risk of the estimator, from which rates of convergence are obtained and are established to be optimal. An adaptive procedure, achieving simultaneous and anisotropic selection of each space dimension, is then tailored and an oracle risk bound is proved. The procedure is studied numerically and implemented on a real dataset of electric consumption.
This paper deals with the process X = (X_t)_t∈ [0,T] defined by the stochastic differential equation (SDE) dX_t = (a(X_t) + b(Y_t))dt +σ(X_t)dW_1(t), where W_1 is a Brownian motion and Y is an exogenous process. The first task - of probabilistic nature - is to properly define the model, to prove the existence and uniqueness of the solution of such an equation, and then to establish the existence and a suitable control of a density with respect to the Lebesgue measure of the distribution of (X_t,Y_t) (t > 0). In the second part of the paper, a risk bound and a rate of convergence in specific Sobolev spaces are established for a copies-based projection least squares estimator of the ℝ^2-valued function (a,b). Moreover, a model selection procedure making the adequate bias-variance compromise both in theory and practice is investigated.
We consider the nonparametric estimation of the value of a quadratic functional evaluated at the density of a strictly positive random variable X based on an iid. sample from an observation Y of X corrupted by an independent multiplicative error U. Quadratic functionals of the density covered are the 𝕃^2 -norm of the density and its derivatives or the survival function. We construct a fully data-driven estimator when the error density is known. The plug-in estimator is based on a density estimation combining the estimation of the Mellin transform of the Y density and a spectral cut-off regularized inversion of the Mellin transform of the error density. The main issue is the data-driven choice of the cut-off parameter using a Goldenshluger–Lepski-method. We discuss conditions under which the fully data-driven estimator attains oracle-rates up to logarithmic deteriorations. We compute convergence rates under classical smoothness assumptions and illustrate them by a simulation study.
We consider N independent and identically distributed one-dimensional inhomogeneous diffusion processes ( X i ( t ) , i = 1 , ... , N ) with drift \mu ( t, x ) = \sum jK =1 \alpha j ( t ) g j ( x ) and diffusion coefficient \sigma ( t, x ), where K and the functions g j ( x ) and \sigma ( t, x ) are known. Our concern is the nonparametric estimation of the K -dimensional unknown function ( \alpha j ( t ) , j = 1 , ... , K ) from the continuous observation of the sample paths ( X i ( t )) throughout a fixed time interval [0 , \tau ]. A collection of projection estimators belonging to a product of finite -dimensional subspaces of L 2 ([0 , \tau ]) is built. The L 2 -risk is defined by the expectation of either an empirical norm or a deterministic norm fitted to the problem. Rates of convergence for large N are discussed. A data -driven choice of the dimensions of the projection spaces is proposed. The theoretical results are illustrated by numerical experiments on simulated data.
In this paper, we propose a nonparametric estimation strategy for the conditional density function of Y given X, from independent and identically distributed observations (Xi, Yi) 1=i=n. We consider a regression strategy related to projection subspaces of L 2 generated by non compactly supported bases. This rst study is then extended to the case where Y is not directly observed, but only Z = Y + e, where e is a noise with known density. In these two settings, we build and study collections of estimators, compute their rates of convergence on anisotropic space on non-compact supports, and prove related lower bounds. Then, we consider adaptive estimators for which we also prove risk bounds.
In this paper, we study the estimation of the derivative of a regression function in a standard univariate regression model. The estimators are defined either by derivating nonparametric least-squares estimators of the regression function or by estimating the projection of the derivative. We prove two simple risk bounds allowing to compare our estimators. More elaborate bounds under a stability assumption are then provided. Bases and spaces on which we can illustrate our assumptions and first results are both of compact or noncompact type, and we discuss the rates reached by our estimators. They turn out to be optimal in the compact case. Lastly, we propose a model selection procedure and prove the associated risk bound. To consider bases with a noncompact support makes the problem difficult.
In the present paper, we consider that N diffusion processes X 1 , … , X N are observed on [ 0 , T ], where T is fixed and N grows to infinity. Contrary to most of the recent works, we no longer assume that the processes are independent. The dependency is modeled through correlations between the Brownian motions driving the diffusion processes. A nonparametric estimator of the drift function, which does not use the knowledge of the correlation matrix, is proposed and studied. Its integrated mean squared risk is bounded and an adaptive procedure is proposed. Few theoretical tools to handle this kind of dependency are available, and this makes our results new. Numerical experiments show that the procedure works in practice.
We study the non-parametric estimation of the value θ(f ) of a linear functional evaluated at an unknown density function f with support on R_+ based on an i.i.d. sample with multiplicative measurement errors. The proposed estimation procedure combines the estimation of the Mellin transform of the density f and a regularisation of the inverse of the Mellin transform by a spectral cut-off. In order to bound the mean squared error we distinguish several scenarios characterised through different decays of the upcoming Mellin transforms and the smoothnes of the linear functional. In fact, we identify scenarios, where a non-trivial choice of the upcoming tuning parameter is necessary and propose a data-driven choice based on a Goldenshluger-Lepski method. Additionally, we show minimax-optimality over Mellin-Sobolev spaces of the estimator.
We consider a stochastic system of N interacting particles with constant diffusion coefficient and drift linear in space, time-depending on two unknown deterministic functions. Our concern here is the nonparametric estimation of these functions from a continuous observation of the process on [0, T] for fixed T and large N. We define two collections of projection estimators belonging to finite-dimensional subspaces of L-2([0, T]). We study the L-2-risks of these estimators, where the risk is defined either by the expectation of an empirical norm or by the expectation of a deterministic norm. Afterwards, we propose a data-driven choice of the dimensions and study the risk of the adaptive estimators. The results are illustrated by numerical experiments on simulated data.
In this paper, we consider the inverse problem of estimating the product fg of two densities, given a d-dimensional n-sample of i.i.d. observations drawn from each distribution. We propose a general method of estimation encompassing both projection estimators with model selection device and kernel estimators with bandwidth selection strategies. The procedures do not consist in making the product of each density estimator, but in plugging an overfitted estimator of one of the two densities, in an estimator based on the second sample. Our findings are a first step toward a better understanding of the good performances of overfitting in regression Nadaraya-Watson estimator.
In this paper, we consider the problem of estimating the d -th order derivative f^(d) of a density f , relying on a sample of n i.i.d. observations X_1,…,X_n with density f supported on ℝ or ℝ^+ . We propose projection estimators defined in the orthonormal Hermite or Laguerre bases and study their integrated 𝕃^2 -risk. For the density f belonging to regularity spaces and for a projection space chosen with adequate dimension, we obtain rates of convergence for our estimators, which are optimal in the minimax sense. The optimal choice of the projection space depends on unknown parameters, so a general data-driven procedure is proposed to reach the bias-variance compromise automatically. We discuss the assumptions and the estimator is compared to the one obtained by simply differentiating the density estimator. Simulations are finally performed. They illustrate the good performances of the procedure and provide numerical comparison of projection and kernel estimators
In a regression model, we write the Nadaraya-Watson estimator of the regression function as the quotient of two kernel estimators, and propose a bandwidth selection method for both the numerator and the denominator. We prove risk bounds for both data driven estimators and for the resulting ratio. The simulation study confirms that both estimators have good performances, compared to the ones obtained by cross-validation selection of the bandwidth. However, unexpectedly, the single-bandwidth cross-validation estimator is found to be much better than the ratio of the previous two good estimators, in the small noise context. However, the two methods have similar performances in models with large noise.
We study the non-parametric estimation of an unknown density f with support on R+ based on an i.i.d. sample with multiplicative measurement errors. The proposed fully data driven procedure is based on the estimation of the Mellin transform of the density f , a regularisation of the inverse of the Mellin transform by a spectral cut-off and a data-driven model selection in order to deal with the upcoming bias-variance trade-off. We introduce and discuss further Mellin-Sobolev spaces which characterize the regularity of the unknown density f through the decay of its Mellin transform. Additionally, we show minimax-optimality over Mellin-Sobolev spaces of the data-driven density estimator and hence its adaptivity.
In this article, we consider the problem of nonparametric hazard rate estimation in the presence of right‐censored observations. We provide a generalized risk bound for a regression‐type nonparametric estimator of the hazard function of interest. Under adequate integrability conditions, our bound is a generalization of estimation strategies specific to compactly supported bases to bases that are not necessarily compactly supported. We show that it encompasses previous compact‐support results and interestingly represents hazard rates as combinations of gamma functions. We discuss the model selection method, which comes out from the new terms of the risk bounds, and compare the performance of the new estimator to that of previous ones, when using a noncompact Laguerre basis. A real data example is also presented.