Given a sample of n observations with sample mean (X) over bar and standard deviation S drawn independently from a population with unknown mean It, it is well known that the skewness of the statistic n(-1/2)((X) over bar - mu)/S is of the opposite sign to the skewness of the population. As a consequence, an equal-tailed confidence interval for It may be used for descriptive purposes, since the relative position of (X) over bar in the interval provides visual information about the skewness of the population. In this paper, we are interested in confidence intervals for the mean which share this descriptive property. We formally define two simple classes of intervals where the degree of asymmetry around (X) over bar monotonically depends on the sample skewness through a parameter lambda with values between 0 and 1. These classes contain the symmetric ordinary-z (or the ordinary-t) confidence interval as a special case. We show how to determine this parameter lambda in order to obtain an equal-tailed confidence interval for mu which is second order accurate. While our first solution has already been investigated in the literature and has serious drawbacks, our second solution appears to be new and sound. Moreover, our method provides the sample skewness with a new and concrete interpretation.
Motivated by multivariate data on epicentres of earthquakes, we suggest nonparametric methods for analysis of point-process data. Our methods are based partly on nonparametric intensity estimation, and involve techniques for dimension reduction and for mapping the trajectory of temporal evolution of high-intensity clusters. They include ways of improving statistical performance by data sharpening, i.e. data pro-processing before substitution into a conventional nonparametric estimator. We argue that the 'true' intensity function is often best modelled as a surface with infinite poles or pole lines, and so conventional methods for bandwidth choice can be inappropriate. The relative severity of a cluster of events may be characterised in terms of the rate of asymptotic approach to a pole. The rate is directly connected to the correlation dimension of the point process, and may be estimated nonparametrically or semipaxametrically.
We consider methods for kernel regression when the explanatory and/or response variables are adjusted prior to substitution into a conventional estimator. This "data-sharpening" procedure is designed to preserve the advantages of relatively simple, low-order techniques, for example, their robustness against design sparsity problems, yet attain the sorts of bias reductions that are commonly associated only with high-order methods. We consider Nadaraya-Watson and local-linear methods in detail, although data sharpening is applicable more widely. One approach in particular is found to give excellent performance. It involves adjusting both the explanatory and the response variables prior to substitution into a local linear estimator. The change to the explanatory variables enhances resistance of the estimator to design sparsity, by increasing the density of design points in places where the original density had been low When combined with adjustment of the response variables, it produces a reduction in bias by an order of magnitude. Moreover, these advantages are available in multivariate settings. The data-sharpening step is simple to implement, since it is explicitly defined. It does not involve functional inversion, solution of equations or use of pilot bandwidths.
Summary Given a linear time series, e.g. an autoregression of infinite order, we may construct a finite order approximation and use that as the basis for confidence regions. The sieve or autoregressive bootstrap, as this method is often called, is generally seen as a competitor with the better-understood block bootstrap approach. However, in the present paper we argue that, for linear time series, the sieve bootstrap has significantly better performance than blocking methods and offers a wider range of opportunities. In particular, since it does not corrupt second-order properties then it may be used in a double-bootstrap form, with the second bootstrap application being employed to calibrate a basic percentile method confidence interval. This approach confers second-order accuracy without the need to estimate variance. That offers substantial benefits, since variances of statistics based on time series can be difficult to estimate reliably, and—partly because of the relatively small amount of information contained in a dependent process—are notorious for causing problems when used to Studentize. Other advantages of the sieve bootstrap include considerably greater robustness against variations in the choice of the tuning parameter, here equal to the autoregressive order, and the fact that, in contradistinction to the case of the block bootstrap, the percentile t version of the sieve bootstrap may be based on the ‘raw'’ estimator of standard error. In the process of establishing these properties we show that the sieve bootstrap is second order correct.
A 'skewing' method is shown to effectively reduce the order of bias of locally parametric estimators, and at the same time retain positivity properties. The technique involves first calculating the usual locally parametric approximation in the neighbourhood of a point x' that is a short distance from the place x where we wish to estimate the density, and then evaluating this approximation at x. By way of comparison, the usual locally parametric approach takes x' = x. In our construction, x' - x depends in a very simple way on the bandwidth and the kernel, and not at all on the unknown density. Using skewing in this simple form reduces the order of bias from the square to the cube of bandwidth; and taking the average of two estimators computed in this way further reduces bias, to the fourth power of bandwidth. On the other hand, variance increases only by at most a moderate constant factor.
As an alternative to traditional, parametric approaches, we suggest nonparametric methods for analyzing spatial and temporal data on earthquake occurrences. Nonparametric techniques are particularly adaptive to anomalous behavior in the data and provide a new way of accessing a variety of different types of information about the way in which both intensity and magnitude of events evolve in time. They can be employed to estimate the spatial trajectory of event clusters as a function of time, and to define quiescent and active periods. The latter application suggests new approaches to forecasting high magnitude events. Our methods are founded on multivariate techniques for curve and surface estimation, particularly in contexts where curves or surfaces are unbounded at points or along lines.
The standard approach to local linear regression involves fitting a straight line segment to a curve in a symmetrical way, in that the segment is fitted directly above a small region whose midpoint is the abscissa, x, at which we wish to estimate the curve. In this paper we show that, if the segment is fitted in a skew manner, with its centre a little to the left or right of x, then bias can be reduced by an order of magnitude, without affecting the order of magnitude of variance. The amount by which the centre should be shifted depends only on the kernel function, and not at all on the unknown regression mean or on the design density. The average of two similarly but oppositely shifted estimators has two orders of magnitude less bias, again at the expense of a slight increase in variance. This particular estimator may be viewed as a limiting form of a convex combination of three local linear estimators, the two oppositely shifted estimators and the symmetric one, which has two orders of magnitude less bias and, depending on the kernel function, less variance as well.