Document clustering has not been well received as an information retrieval tool. Objections to its use fall into two main categories: first, that clustering is too slow for large corpora (with running time often quadratic in the number of documents); and second, that clustering does not appreciably improve retrieval. We argue that these problems arise only when clustering is used in an attempt to improve conventional search techniques. However, looking at clustering as an information access tool in its own right obviates these objections, and provides a powerful new access paradigm. We present a document browsing technique that employs docum-ent clustering as its primary operation. We also present fast (linear time) clustering algorithm.
An efficient method for the calculation of the interactions of a 2' factorial experiment was introduced by Yates and is widely known by his name. The generalization to 3' was given by Box et al. [1]. Good [2] generalized these methods and gave elegant algorithms for which one class of applications is the calculation of Fourier series. In their full generality, Good's methods are applicable to certain problems in which one must multiply an N-vector by an N X N matrix which can be factored into m sparse matrices, where m is proportional to log N. This results inma procedure requiring a number of operations proportional to N log N rather than N2. These methods are applied here to the calculation of complex Fourier series. They are useful in situations where the number of data points is, or can be chosen to be, a highly composite number. The algorithm is here derived and presented in a rather different form. Attention is given to the choice of N. It is also shown how special advantage can be obtained in the use of a binary computer with N = 2' and how the entire calculation can be performed within the array of N data storage locations used for the given Fourier coefficients. Consider the problem of calculating the complex Fourier series N-1
We propose a technique for combining estimates of the same quantity into a single number and its estimated variance. Several series (or experiments) provide estimates,yi, of the quantity of interest and of the variance ofyi, estimated internally to the series, namelysi2. In the special case where allyiestimate the same number, Cochran has proposed a method called partial weighting, which groups together the half to two-thirds of the estimates that have the smallest estimated variances and gives them equal weights, and weights each of the otheryiseparately, inversely as their variance.
This paper presents a method for detecting and nominating outliers based on the multihalver, or the delete-half jackknife. Since considering all possible half-samples is unpractical and unfeasable even for a moderate sample size, we present an algorithm for choosing a good set of half-samples. We also present an outlier detection method based on this algorithm. Simulations are given to show the effectiveness of our method and an example is also presented.
John Tukey connected the theory underlying simple random sampling without replacement, cumulants, expected mean squares and spectrum analysis. He gave us one degree of freedom for nonadditivity, and he pioneered finite population models for understanding ANOVA. He wrote widely on the nature and purpose of ANOVA, and he illustrated his approach. In this appreciation of Tukey’s work on ANOVA we summarize and comment on his contributions, and refer to some relevant recent literature.
In large two-way tables an additive fit typically leaves too much discernable structure in the residuals. Multipolishing refers to the fitting of more complex models allowing for nonlinear effects of the factors and thus partially describing the interaction between the factors. This paper introduces multipolishing and shows how the two-way plots, familiar in the case of additive fits, can be adapted to these more general and more detailed descriptions of the data table. Additional ideas such as the fragmentation of data tables, the inclusion of residuals into the plots, or the visual comparison of different fits are briefly discussed.
SUMMARY A distinctive feature of analysis of variance is the common occurrence of more than one error term. This feature calls attention to the two distinct potential roles of a single mean square. As a "numerator" it measures the variability visible at a given level in a design hierarchy, and as a "denominator" it measures how much variability has been "passed up" to higher levels, and may, if appropriate, serve as part of an error term. We propose a straightforward multiphase procedure that explicitly recognizes these two roles, and argue that, in general, such considerations preclude naive use of robust regression techniques for analysis of factorially designed experiments. Instead, an upsweeping-by-medians decomposition of the data is followed by a comparison-within-subtable analysis to flag exotic ("relatively large") entries in each of the subtables associated with the dif- ferent sorts of variation. A classical analysis by means, after replacing each identified exotic entry by an algorithmically specified value, yields a decomposition of the data that is used to construct an analysis of variance table in which for each sort of variation there is both a list of any exotic entries and an inner ("denominator") mean square that 'excludes' those exotic entries. The analysis can then be completed by downsweeping the inner subtables that are insuciently prominent, and providing (formally) appropriate error terms for analyzing table entries that remain. The results are displayed as a decomposition of the data into exotic values and those inner subtables, both simple and composite, that survive downsweeping. The exploratory nature of the approach is emphasized, and the method is applied to an example of a factorial experiment in which all factors have three or more versions.
Cox, J. L., Heyse, J. F., and Tukey, J. W. 2000. Efficacy estimates from parasite count data that include zero counts. Experimental Parasitology96, 1–8. Because of the positive skewness of parasite distributions and the greater constancy of percentage of response of therapy in animal populations, parasite count data are conventionally transformed logarithmically before combining results from different animals, either all controls or all treated. Observations of zero counts raise difficulties, since the logarithm of zero is not useful. In this study, several types of zero count adjustments are compared. Two systems for assigning values to zero counts were considered: a fixed system, which assigns the same value to all zero counts regardless of the proportion of such counts in a treatment group, and a variable system, which replaces zero counts with a value based on the proportion of zero counts in the group. The values assigned by either system are then adjusted to reflect aliquot size. An evaluation was performed by using 32 compound Poisson lognormal distributions, three sample sizes, and three representatives of each zero count adjustment system. The Poisson lognormal distribution provides a convenient method with which to provide variability greater than Poisson. Expected values of the sample estimate of the (known) population mean were calculated for each of the 576 combinations of these factors, and the bias associated with each combination was derived. The bias associated with the three representatives of the variable adjustment system was similar. The variable adjustment system had a lower overall bias than any representatives of the fixed adjustment system.
This article introduces a family of distributional shapes which is flexible in the sense that it contains skewed and symmetric laws as well as heavy-tailed and light-tailed laws. The proposed family is also practically convenient because it is easy to fit to a table of quantiles from any distribution. Inversely, for each of the distributional shapes it is trivial to compute quantiles for any desired probability, and it is possible to compute the corresponding densities.