This computerized study of the homonyms of elementary words (roughly equivalent to monosyllabic words) has allowed the compilation of exhaustive lists of homonym sets, using phonetic transcriptions from five different dictionaries. Of the 5,757 elementary words, 2,966 were involved in at least one homonym set, indicating that homonyms will pre-sent a significant problem in mechanized word recognition. The effects on the homonym sets of changing from the phonetic transcription of one dictionary to another were tabulated, as were the effects of removing dialectal pronunciations. Since the effects of dialectal variations turned out to be relatively small, it was possible to categorize and list for study the actual words whose dialectal pronunciations caused homonym-type confusion with other words.
The classifying of words according to syntactic usage is basic to language handling; this paper describes an algorithm for automatically classifying words according to thirteen commonly used parts of speech: noun, adjective, verb, past verb, adverb, preposition, conjunction, pronoun, interjection, present participle, past participle, auxiliary verb, and plural or collective noun. The algorithm was derived by a computerized study of the words in The Shorter Oxford English Dictionary. In its operation it utilizes a prepared dictionary of around nine hundred words to assign parts of speech to special or exceptional words. Other words are split into affix and kernel parts and assigned a part of speech on the basis of the part-of-speech implications of the affixes and the length of the remaining kernel. An accuracy of 95 per cent is achieved from the point of view of inclusive part of speech, where inclusive part of speech is defined as that string which contains all the parts of speech attributed to the word by the dictionary but which may also contain one or two more parts of speech.
This paper describes a systematic investigation of the extent to which the part of speech of words can be identified from their prefixes and suffixes. The results indicate that it is possible to determine, with 95 per cent accuracy, the part of speech of an affixed word from a consideration of its prefixes, suffixes, and length. By inclusive parts of speech we mean a string that will include all of the parts of speech assigned by both dictionaries considered but that may include one or two extraneous parts of speech. The extra parts of speech will differ according to the class of words, as adjectives may have an extra part-of-speech noun or adverb, while nouns may have an extra part-of-speech verb. The part-of-speech implications of seventy-two prefixes and of eightyseven suffixes are given. In a highly inflected language, the structure of a word is indicative of its syntactic role. A relationship between form and part of speech might also be expected in English, a language not highly inflected but closely related to more inflected languages. Such a relationship was noted by J. Dolby and H. Resnikoff, 1 who show that a high percentage of a set of words called “elementary words” (roughly equivalent to the set of onesyllable words) can be used as nouns, adjectives, or verbs, while a high percentage of the remaining multisyllable words can be used only as nouns or adjectives. If this relation can be regarded as a general rule, and if subrules can be developed to cover the considerable number of exceptions to the general rule, it will be possible to identify part of speech by algorithm. Intuitively, it would be expected that prefixes and suffixes are key structural elements; this expectation is reinforced by the structure of the European languages whose beginnings and endings indicate the grammatical properties of words. A logical step in an effort to classify words from their structure is to examine the relationship between the affixes of words and their part-of-speech possibilities as listed in a dictionary. The part-of-speech information from The Shorter Oxford Dictionary 2 and from the Merriam Webster New International Dictionary 3 was recorded on magnetic tape. A computer was used to correlate the affixes of words with their part-of-speech possibilities. A total of 73,582 words was recorded, but, of course, not all of these words contain affixes.