
In this paper, we propose a new variable selection method for Nadaraya-Watson-Frechet regression, which is a special case of local Fr & eacute;chet regression. We demonstrate both the practical performance and the theoretical properties of this novel variable selection method. To provide theoretical support, we first derive a refined uniform convergence result for the Nadaraya-Watson-Frechet regression estimator of the Fr & eacute;chet regression function, which then enables us to establish asymptotic selection consistency for the proposed variable selection method. Excellent performance of the proposed variable selection method for Nadaraya-Watson-Frechet regression in finite sample situations is demonstrated both with simulation studies and in real data examples.
Statistical methods for dynamic network analysis play an essential role in modelling the temporal dynamics of network structures. While many studies have modelled longitudinal networks based on continuous-time observations, research based on discrete-time observations remains relatively limited. This study introduces a statistical method that integrates dynamic stochastic block models with Poisson processes to analyse recurrent dyadic interaction events observed at discrete time points. A variational expectation-maximisation algorithm is employed for parameter estimation. The asymptotic properties of the proposed dynamic model are discussed. Simulation studies studies and real-world network applications demonstrate the effectiveness of the proposed model in modelling the temporal evolution of network structures and interaction patterns.
This paper explores the Bahadur representation for nonparametric estimation of expected shortfall under the dependent random variables, incorporating eight types of sample quantiles. The Bahadur representation is employed to establish the asymptotic normality and Berry-Esseen bounds of nonparametric estimators for expected shortfall under dependent random sequences. Furthermore, numerical simulations demonstrate that the combined use of the fourth ES estimator (E-4, corresponding to (mu) over bar (p) ) and the eighth ES estimator (E-8, also corresponding to (mu) over bar (p) ) yields more precise ES estimates relative to the other estimators evaluated in this study, outperforming the previously established seventh estimator (E-7 , corresponding to (mu) over bar (p)). These simulations successfully corroborated the Berry-Esseen bound issue associated with ES estimators in finite sample contexts, and the estimation of expected shortfall is applied to real data analysis.
Computer models are used to solve complex problems in many scientific applications, such as nuclear physics and climate research. Markov chain Monte Carlo-based Bayesian calibration of computer models although a popular approach, is computationally expensive. This work proposes a fast and scalable posterior approximation algorithm for Bayesian computer model calibration via Variational Inference. We provide the statistical guarantee of the proposed algorithm in the form of a posterior contraction theorem for the estimated physical process. To this end, we establish that the variational posterior concentrates in $ \epsilon _n $ & varepsilon;n neighbourhoods of the true physical process under regularity assumptions on the variational family. The main results are shown in the two widely used classes of Gaussian process priors, the Squared Exponential covariance class and the Mat & eacute;rn covariance class. Finally, we provide a simulation study to demonstrate the proposed method's computational efficiency and fidelity compared to the standard Markov chain Monte Carlo method.
This work addresses the problem of testing independence in r-tuples of orientations of symmetric objects in three-dimensional space. Such objects naturally occur in crystallography, materials science, and other fields. We propose a novel definition of covariance between two random orientations and utilise it in the construction of multivariate asymptotic tests for pairwise uncorrelatedness based on U-statistics. We evaluate the finite-sample performance of these tests through a simulation study that investigates their power for three models of r-tuples of random orientations and compares them with permutation tests based on sample covariances and total distance multivariance. The application of the proposed tests is demonstrated on a dataset of polycrystalline material with cubic symmetry of the crystal lattice.
This paper considers pointwise inference for nonparametric estimating equations models. The paper proposes two general test statistics that are based on a local version of the Generalised Empirical Likelihood (GEL) approach that can be used to test simple hypotheses about the unknown infinite dimensional parameters and to test for the correct specification of the chosen nonparametric estimating equations model. The paper shows that among the class of the proposed GEL test statistics, the empirical likelihood ratio is the only one admitting a Bartlett correction, however by appropriately modifying the other GEL based test statistics, it is still possible to obtain second order accurate inferences. The paper also proposes a new (local) version of the so-called efficient bootstrap that delivers the same level of second order accuracy as that of the (modified) GEL test statistics for the correct specification of the chosen nonparametric estimating equations model. Finally, the paper uses simulations and a real data example to illustrate the finite sample properties and applicability of the proposed inference methods.
This article introduces a model averaging prediction method for high-dimensional varying coefficient models. The proposed method can reduce efficiently the effect of misspecification of index variables. To improve the actual prediction performance, we combine the model averaging method with the Adaboost algorithm, which can obtain more accurate estimates of the response variables and model averaging weights. Compared with other competing model averaging methods and various popular machine learning methods, the proposed method outperforms the competitors and the prediction accuracy is improved with the assistance of Adaboost algorithm through simulation results. We analyse the house prices data set and provide more precisely predictions.
In this paper, we consider the heteroscedastic time varying-coefficient errors-in-variables models. We establish the wavelet estimator of the coefficient function by weighted local least-squares method. Under some suitable conditions, the asymptotic normality for the proposed wavelet estimators are considered under alpha-mixing errors. In addition, in order to illustrate the feasibility of theoretical results, the simulation study and real data analysis are provided on finite samples.
Measurement errors are common, yet most research focusses on a single error-prone covariate. In practice, multiple variables often exhibit correlated measurement errors. We propose a novel method for estimating a bivariate multiplicative measurement error model when both covariates are contaminated. Likelihood-based estimation in such models is computationally demanding due to repeated evaluation of two-dimensional integrals, especially for large samples. To address this, we exploit the lognormal error structure, which enables rewriting the likelihood integrals using cumulative distribution function values of the bivariate normal distribution. This significantly reduces computational cost and improves numerical stability. Simulation studies demonstrate the accuracy and efficiency of the proposed method. We further illustrate its practical utility by correcting a biased Pearson correlation and applying it to fetal biometry data.
Given an independent and identically distributed sample of angles from some absolutely continuous circular random variable with unknown probability density function f, in this work we study the problem of testing the hypothesis on whether f is the uniform distribution on the circle. For this purpose we consider a Bickel-Rosenblatt type test statistic ( $ L<^>2 $ L2 distance) based on the Parzen-Rosenblatt type estimator for circular data. The asymptotic behaviour of the proposed test procedure for fixed and non-fixed bandwidths is studied. From a finite sample point of view the power performance of the tests associated with different bandwidths depends on the considered bandwidth which acts as a tuning parameter. The automatic selection of this tuning parameter, the choice of which is crucial to obtaining a performing test procedure, is also addressed in this work, and comparisons are made with other existing uniformity tests through a simulation study.
The problem of bandwidth selection in nonparametric classification is considered when a large number of class variables may be missing, but not necessarily missing at random. This setup is generally acknowledged to be more challenging than the so-called missing at random setup. Our proposed approach starts by constructing a family of kernel-based classifiers where the members of the family are indexed by the kernel bandwidth h as well as the nonignorability parameter of the selection probability mechanism. Next, a search is performed to find the member of a finite cover of this family that has the smallest empirical misclassification error. To assess the performance of the resulting classifiers, we derive exponential performance bounds on the deviations of their errors from that of the theoretically optimal classifier. These bounds are then used to study strong convergence properties of the proposed classifiers. Our theoretical findings are further confirmed by numerical studies.
We consider the problem of building adaptive, rate optimal estimators for the mean and covariance functions of random curves in the context of streaming data. In general, functional data analysis requires nonparametric smoothing of curves observed at a discrete set of design points, which may be measured with error. However, classical nonparametric smoothing methods (e.g. kernels, splines, etc.) assume that the degree of smoothness is known. In many applications functional data could be irregular, even perhaps nowhere differentiable. Moreover, the (ir)regularity of the curves could vary across their domain. We contribute to the literature by providing estimators and inference procedures that use an iterative plug-in estimator of 'local regularity' which delivers a computationally attractive, recursive, online updating method that is well-suited to streaming data. Theoretical support and Monte Carlo simulation evidence are provided, and code in the R language is available for the interested reader.
This paper considers variable selection for mixed panel count data, which frequently occur in longitudinal studies and whose analysis is quite challenging due to their complex data structures. For the problem, we propose a penalised likelihood procedure with the use of Gaussian Seamless-L-0 penalty under a proportional mean model. For its implementation, a computationally efficient EM algorithm is developed that enables sparse variable selection while ensuring accurate parameter estimation. The resulting estimator is shown to have the oracle property, and a simulation study is performed and confirms the strong finite-sample performance of the proposed method. Finally we apply the proposed approach to a set of real data on medical non-adherence arising from the Sequenced Treatment Alternatives to Relieve Depression Study and identify some new risk factors.
In recent years, significant progress has been made in understanding the theoretical properties of estimators in deep neural networks. However, most of the existing research is limited to scenarios involving bounded loss functions or bounded input/output data. To the best of our knowledge, there is currently no work that addresses the strong convergence of DNN estimators in general settings. In this paper, we investigate the rate of strong convergence for generalisation bounds and excess risk in learning psi-weakly dependent processes and strong mixing processes, where the loss function and input/output are not necessarily bounded. Under some appropriate assumptions, the asymptotic learning rate approximates o(n(-mu+1/2 mu+3)) with mu >= 0 for psi-weakly dependent processes and o(n(-1/tau)) with tau > 2 for strong mixing processes. We also provide some simple simulation results to support our findings.
Kolmogorov-Smirnov (KS) statistic has been widely used in many areas to evaluate the performance of binary classification. However, almost no classification algorithm tries to optimise it directly at the training stage due to the computational and theoretical challenges brought by the special form of KS. In this paper, we propose a novel Kolmogorov-Smirnov neural Network (KSNet) using KS as the optimisation objective. The difficulty of non-smoothness of the empirical KS is overcame by introducing a smooth nonconvex surrogate function. The KSNet brings great potential to improve the KS in test data especially for imbalanced data and it shows inspiring robustness to data noise. Theoretically, we establish the non-asymptotic excess risk bound of KSNet with a ReLU activated feedforward neural network and show its Bayes-risk consistency. Experiments on a variety of real datasets confirm the advantages of KSNet over a lot of existing methods.
In this paper, we propose a novel robust semi-Gaussian kernel distance (RSGKD) covariance to measure dependence between a random vector $ {\bf X} $ X and a categorical random variable Y. Our proposed RSGKD covariance satisfies the independence-zero equivalence property, i.e. the RSGKD covariance is non-negative and equals zero if and only if $ {\bf X} $ X and Y is statistically independent. Since our proposed RSGKD covariance does not require moment restrictions on random vectors, it is robust to heavy-tailed distributions or outliers in the dataset. Under some conditions, the asymptotic properties of our proposed test statistic are established. Furthermore, both extensive numerical simulations and a real data analysis demonstrate the satisfactory finite-sample performance of our proposed test.