Many theorists have dismissed a priori the idea that distributional information could play a significant role in syntactic category acquisition. We demonstrate empirically that such information provides a powerful cue to syntactic category membership, which can be exploited by a variety of simple, psychologically plausible mechanisms. We present a range of results using a large corpus of child-directed speech and explore their psychological implications. While our results show that a considerable amount of information concerning the syntactic categories can be obtained from distributional information alone, we stress that many other sources of information may also be potential contributors to the identification of syntactic classes.
The use of NLP techniques for document classification has not produced significant inprovements in performance within the standard term weighting statistical assignment paradigm (Fagan 1987; Lewis, 1992ab; Lewis and Sparck-Jones 1993; Buckley, 1993). This perplexing fact needs both an explanation and a solution if the power of recently developed NLP techniques are to be successfully applied in IR. This paper repeats results which show that standard statistical models are not particularly suitable for exploiting linguistically sophisticated representations, and offers another statistically based model of inference which provides significantly improved performance for sophisticated representations. It therefore shows that statistical systems can exploit sophisticated representations of documents, and lends some support to the use of more linguistically sophisticated representations for document classification. This paper describes a system which is being developed for the LRE project SISTA, and which is accompanied by a working PC prototype demonstration system.
We present a neural network which learns linguistic categories from unlabelled data. A Hebbian mechanism learns the distribution of contexts in which each item occurs, and these contexts are these clustered using a Kohonen network. In a pretest, the network separates vowels and consonsants into separate clusters, as expected. The network is then trained on a large corpus of noisy word-level text, and successfully finds syntactically interesting categories. Statistical analysis of the data set shows that additional syntactic information is implicit in the data, suggesting that network performance can be further improved.