
A stemming algorithm, a procedure to reduce all words with the same stem to a common form, is useful in many areas of computational linguistics and information-retrieval work. While the form of the algorithm varies with its application, certain linguistic problems are common to any stemming procedure. As a basis for evaluation of previous attempts to deal with these problems, this paper first discusses the theoretical and practical attributes of stemming algorithms. Then a new version of a context-sensitive, longest-match stemming algorithm for English is proposed; though developed for use in a library information transfer system, it is of general application. A major linguistic problem in stemming, variation in spelling of stems, is discussed in some detail and several feasible programmed so-lutions are outlined, along with sample results of one of these methods.
This paper describes the use of an on-line system to do word-sense ambiguity resolution and content analysis of English paragraphs, using a system of semantic analysis programmed in Q32 LISP 1.5. The system of semantic analysis comprises dictionary codings for the text words, coded forms of permitted message, and rules producing message forms in combination on the basis of a criterion of semantic closeness. All these can be expressed as a single system of rules of phrase-structure form. In certain circumstances the system is able to enlarge its own dictionary in a real-time mode on the basis of information gained from the actual texts analyzed.
First the notion "paraphrase" is defined, and then several different types of paraphrase are analyzed: transformational, attenuated, lexical, derivational, and real-world. Next, several different methods of retrieving information are discussed utilizing the notions of paraphrase defined previously. It is concluded that a combination keyword-keyphrase method would constitute the optimum procedure.
This paper first presents a theoretical interpretation of the translation process. It then analyzes existing machine-translation research strategies and points out that some of the generally accepted principles of these strategies are not optimal. Finally, an alternative strategy is proposed, based on the author's theoretical position and research results.
A system for semantic analysis of a wide range of English sentence forms is described. The system has been implemented in LISP 1.5 on the System Development Corporation (SDC) time-shared computer. Semantic analysis is defined as the selection of a unique word sense for each word in a natural-language sentence string and its bracketing in an underlying deep structure of that string. The conclusion is drawn that a semantic analyzer differs from a syntactic analyzer primarily in requiring, in addition to syntactic word-classes, a large set of semantic word-classes. A second conclusion is that the use of semantic event forms eliminates the need for selection restrictions and projection rules as posited by Katz. A discussion is included of the relations of elements of this system to the elements of the Katz theory.
To the best of my knowledge, no work has previously been carried out on the mechanical translation of any Bantu language. This note is therefore a first suggestion of a possible basis for a scheme for the mechanical translation of Swahili into English. Swahili, in common with other Bantu languages, makes great use of prefixes. This is its most distinctive feature when compared with European languages. All agreements between adjectives, nouns, and verbs are shown by means of prefixes. There are prefixes for the subject and object of a verb and for the verb tense. Negation of a verb is also shown by means of prefixes. Suffixes are also used, but a lot of Swahili can be spoken without using them. Suffixes are used to show motion to or from a place and, apart from this, are used almost exclusively in modifying the form of verbs. The passive, causative, prepositional, reciprocal, subjunctive, plural imperative, and some singular imperative forms are all constructed by adding a suffix to the verb stem. As is usually the case, addition of a suffix often causes modification of the stem itself. For example, the passive form of a verb ending with the letter a is made by changing the final a to wa, as in kuandika ("to write") and kuandikwa ("to be written"). However, kununua ("to buy") gives rise to kununuliwa ("to be bought"). Prefixes, on the other hand, are added with no amendment to the verb stem, and I see this as one of the reasons why the strong reliance on prefixes will make Swahili reasonably susceptible to mechanical translation. Other advantages of the prefix structure are: 1. There is less need for context-dependent analysis. For example, if the present tense of the verb "run" is recognized in English, one still does not know the final form of the word: it could be "they run" or "he runs." In Swahili, however, no such distinction is made:
In the study of a previously unrecorded language, a taxonomy of the sound system is the most useful starting point for developing the phonological component of a grammar. If the linguist makes at least tentative assumptions about segmentation and fixes the limits of supposedly relevant contexts, a computer can approximate this taxonomy. A program by Alsop reduces a concordance of phonetic segments in their contexts to a series of taxonomic statements about phoneme distribution by applying Bloch's criteria for contrast within limited contexts. When applied to data on Paipai, a Yuman language of the Colorado delta, collected by Wares on a survey trip, the program found contrast between segments Wares had identified as allophones in two parallel consonantal series, indicating a distinction of presumably low functional load with morphophonemic implications.
A computer grammar is described which includes most of the English relative-clause constructions. It is written in the form of a left-to-right phrase-structure grammar with discontinuous constituents and subscripts, which carry such syntactic restrictions as number and verb government category. The motivation for the hierarchy of syntactic choices and for the use of discontinuous constituents is discussed. Many examples are given, and special attention is given to complement constructions and to the relation of the relative pronoun to complex prenominal and post-nominal determiner constructions. Written in COMIT , the program runs as part of a larger grammar of English.
The purpose of this paper is to help linguists contruct a consistent, sufficient and less redundant syntax of language.An acceptable string corresponds to an expression or an utterance: it may be a natural text, a string of morphemes, a tree structure or any kind of representation. A sharp distinction is made between the syntactic function which is an attribute of string and the distribution class which is a set of strings. Syntactic function of a continuous or discontinuous string is defined as the set of all the acceptable contexts of the string, and is called a complete neighborhood. Two contexts are equivalent if they accept or reject any given string at the same time. An elementary neighborhood is the set of all contexts equivalent to one context.Four simple distribution classes are proposed and their properties are discussed.Concatenation rules of a language can be described in terms of concatenated complete neighborhoods or concatenated distribution classes. Some possible representations and their consequences are discussed.Transformational rules are also described in a similar way. However, there is another problem of correspondence of original strings to their transforms. It is useful to establish subsets of elementary neighborhoods and this subclassification may contribute to a simplification of the clumsy representation of derivational history.Finally, some trivial but practically useful conventions are described.
This computerized study of the homonyms of elementary words (roughly equivalent to monosyllabic words) has allowed the compilation of exhaustive lists of homonym sets, using phonetic transcriptions from five different dictionaries. Of the 5,757 elementary words, 2,966 were involved in at least one homonym set, indicating that homonyms will pre-sent a significant problem in mechanized word recognition. The effects on the homonym sets of changing from the phonetic transcription of one dictionary to another were tabulated, as were the effects of removing dialectal pronunciations. Since the effects of dialectal variations turned out to be relatively small, it was possible to categorize and list for study the actual words whose dialectal pronunciations caused homonym-type confusion with other words.
Some considerations are presented regarding certain aspects of automatically translating Russian predicative infinitives into English. Emphasis is placed on the analysis (decoding) of the pertinent infinitive constructions in the source language rather than on the synthesis (encoding) of their equivalents in the target language. The paper does not aim at an exhaustive treatment of the problem, but merely offers some tentative and peripheral suggestions as well as some criticism of previous endeavors to tackle the problem of Russian predicative infinitives in machine translation.
This paper examines the theory of translation in Quine's Word and Object and attempts to show that it involves tacit appeal to a premise concerning a regularity in the behavior of bilinguals. The regularity is one whose existence is neither explained nor rendered probable by the theory. The suggestion that the regularity could result from congenital dispositions to organize and pattern linguistic data in certain characteristic ways is considered and rejected as implausible. This leaves the conclusion that if the regularity does obtain, the most plausible explanation would be that people, when acquiring a language, pay attention to and are guided by information and evidence ignored by Quine's criteria of translation. Thus the novelty of the present discussion is this: if its principle contention is correct, then—even if one embraces the analysis in Word and Object, accepting all of its most controversial theoretical features, for example, its identification of a language with a set of behavioral dispositions and its requirement that analyticity and synonymy be operationally defined— one is still bound to recognize that its survey of relevant evidence is essentially incomplete, and one is logically committed to this recognition by a premise embodied in the very analysis one has embraced. That is, the soundness of the analysis entails its incompleteness, and, thus, the analysis is at best incomplete, at best an account of a fragment of the relevant evidence. Now the fact that theory in a given domain is undetermined by a fragment of the relevant evidence leaves wholly undecided the question whether theory in that domain is undetermined by all the relevant evidence. Thus, assuming the correctness of the contentions in this paper, the doctrine of translational indeterminacy does not follow from the analysis intended to support it, and one of the most elaborate expositions offered in support of Quine's misgivings over the analytic-synthetic distinction fails to make those misgivings plausible.
The classifying of words according to syntactic usage is basic to language handling; this paper describes an algorithm for automatically classifying words according to thirteen commonly used parts of speech: noun, adjective, verb, past verb, adverb, preposition, conjunction, pronoun, interjection, present participle, past participle, auxiliary verb, and plural or collective noun. The algorithm was derived by a computerized study of the words in The Shorter Oxford English Dictionary. In its operation it utilizes a prepared dictionary of around nine hundred words to assign parts of speech to special or exceptional words. Other words are split into affix and kernel parts and assigned a part of speech on the basis of the part-of-speech implications of the affixes and the length of the remaining kernel. An accuracy of 95 per cent is achieved from the point of view of inclusive part of speech, where inclusive part of speech is defined as that string which contains all the parts of speech attributed to the word by the dictionary but which may also contain one or two more parts of speech.
The internal structure of the locative predicate-complement form-class in German is described within the framework of a generative grammar consisting of a phrase-structure (PS) component, a semantic (S) component, and a transformation (T) component. The S-component is interposed between the PS-component and the T-component. The PScomponent generates the deep internal structure of the locative form-class as a function of the metaelement irgendwo, assigning hierarchical relationships and groupings in the process. The S-component translates the irgendwo-quantified syntactic patterns of the P-marker into their corresponding semantic denotational patterns, resulting in an S-marker, and then returns the derivation to its P-marker at the level of the locative class symbols. The T-component then operates on this level, if necessary, to obtain the derived P-marker and thus the surface grammar. The metaelement irgendwo proves to be more than a syntactic filter assigning locative structure. It proves to be a semantic filter that reveals the indexical symbolic nature of the locative adverbs and their symbolic relationships to each other as well as to the locative prepositional phrase.
This study used special reading-comprehension tests to compare the speed and accuracy with which the same Russian technical articles in physics, earth sciences, and electrical engineering could be read by technically sophisticated readers when they were presented in English translated from the original Russian by machine only, by machine plus postediting, and by normal manual procedures. Thus, the emphasis was on the transmission of the technical message rather than on linguistic characteristics. In general, the results consistently showed that manual translations exceeded post-edited translations, which exceeded machine translations across all three disciplines and various types of questions. Losses in speed and efficiency were substantially greater than in accuracy, and differences between machine alone and post-edited generally exceeded differences between post-edited and manual translations. However, it was concluded that machine-alone translations were surprisingly good and well worth further consideration under the proper circumstances.
This paper discusses the task of formulating a model of linguistic performance and proposes an approach toward this goal that is oriented toward an embodiment of the model as a digital-computer program. The methodology of current linguistic theory is criticized for several of its features that render it inapplicable to a realistic model of performance, and remedies for these deficiencies are proposed. The syntactic- and conceptual-data structures, inference rules, generation and understanding mechanisms, and learning mechanisms proposed for the model are all described. The learning process is formulated as a series of five stages, and the roles of non-linguistic feedback and inductive generalization relative to these stages are described. Finally, the implications of a successful performance model for linguistic theory, linguistic applications of computers, and psychological theory are discussed.
This is a continuation of the authors' paper of the same title which appeared in Volume 8 of this journal. The present part extends the authors' definitions of prefix and suffix (in written English) to corpora of three-vowel-string words, and implements them on a corpus K consisting of 19,329 graphemically distinct three-vowel-string words from the Shorter Oxford Dictionary. The notion of a parasitic affix is introduced, and the parasitic suffixes for K are determined.
Automatic syntactic analysis is simplified by disengaging the grammatical rules, by means of a parsing logic, from the computer routines that apply them. A case in point is the John Cocke logic. It iterates on five simple parameters and finds all structures permitted by the grammar, thus testing the rules, which can then be changed without changing the routines. The rules themselves need not be ordered so far as the logic of the system is concerned. However, in operating with an IC grammar, rules for bracketing endocentric constructions must be made quite complex merely to avoid multiple analyses of unambiguous or trivially ambiguous expressions. The rules can be simplified if they are classified and if the system is provided with an additional capability for applying them in a specified order. Although an additional parameter is introduced into the system, the disengagement of grammar from routine is preserved. The additional parameter controls the direction, left-to-right or right-to-left, in which constructions are put together. The decision as to which direction should be specified is a grammatical decision, and is related to Yngve's hypothesis of asymmetry in language. It does not affect the operation of the parsing logic.
The purpose of this paper is to examine the oft-repeated assertion regarding the efficiency of a simple parsing algorithm combinable with a variety of different grammars written in the form of appropriate tables of rules. The paper raises the question of the increasing complexity of the tables when more than the most elementary natural-language conditions are included, as well as the question of the ordering of the rules within such nonelementary tables. Some concrete examples from the field of machine translation will be given in the final version of the paper. Some conclusions are presented.