Article Free Access Share on A directed random paragraph generator Authors: Stanley Y. W. Su The RAND Corporation, Santa Monica, California The RAND Corporation, Santa Monica, CaliforniaView Profile , Kenneth E. Harper The RAND Corporation, Santa Monica, California The RAND Corporation, Santa Monica, CaliforniaView Profile Authors Info & Claims COLING '69: Proceedings of the 1969 conference on Computational linguisticsSeptember 1969 Pages 1–36https://doi.org/10.3115/990403.990416Online:01 September 1969Publication History 1citation3,097DownloadsMetricsTotal Citations1Total Downloads3,097Last 12 Months658Last 6 weeks7 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Article Free Access Share on Syntactic and semantic problems in automatic sentence generation Author: Kenneth E. Harper View Profile Authors Info & Claims COLING '67: Proceedings of the 1967 conference on Computational linguisticsAugust 1967 Pages 1–9https://doi.org/10.3115/991566.991572Online:23 August 1967Publication History 0citation468DownloadsMetricsTotal Citations0Total Downloads468Last 12 Months82Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
A study was made of the degree of similarity between pairs of Russian nouns, as expressed by their tendency to occur in sentences with identical words in identical syntactic relationships. A similarity matrix was prepared for forty nouns; for each pair of nouns the number of shared (i) adjective dependents, (ii) noun dependents, and (iii) noun governors was automatically retrieved from machine-processed text. The similarity coefficient for each pair was determined as the ratio of the total of such shared words to the product of the frequencies of the two nouns in the text. The 780 pairs were ranked according to this coefficient. The text comprised 120,000 running words of physics text processed at The RAND Corporation; the frequencies of occurrence of the forty nouns in this text ranged from 42 to 328.The results suggest that the sample of text is of sufficient size to be useful for the intended purpose. Many noun pairs with similar properties (synonymy, antonymy, derivation from distributionally similar verbs, etc.) are characterized by high similarity coefficients; the converse is not observed. The relevance of various syntactic relationships as criteria for measurement is discussed.
The present study is a practical guide to editors who refine partially machine-translated text as a basis for linguistic analysis. The posteditors' tasks are: to code preferred English equivalents, to code English structural symbols, to resolve grammatic properties, and to code syntactic connections (dependencies). A general introduction to the field of machine translation is contained in The RAND Corporation RM-2060.
Callaham's Russian-English Technical and Scientific Dictionary was consulted for English equivalents of the multiple-valued words. The average number of equivalents for these words was 8.6. Many of these equivalents may be considered as synonyms; when there is a fairly distinct change in meaning, the listing is divided into groups, separated by a semicolon. The average number of such groups, representing distinct meanings, is 3. 0 (see Fig. II). Conclusion: 30% of the total words in the passage analyzed should be represented by three English equivalents. This figure has no meaning as applied to any given word; it is perhaps an indication of the extent of the problem of multiple meaning. The problem can be partially solved by an arbitrary selection of a given equivalent for certain fields (the Idioglossary). Even here, there are definite limits, as the list of typical multiple meaning words in Fig. II shows.
IN THE VARIOUS PROPOSALS for word-forword machine translation of Russian scientific literature into English, each word in the sentence is considered as a separate entity. If a word has more than one English equivalent, or more than one possible syntactic value, the alternatives must be listed. The chief difficulty with the resulting translation is its prolixity: the reader finds himself confronted with numerous alternatives, both syntactic and semantic, in every sentence. The extent of the problem of ambiguity is suggested by the following figures: from a sample Russian scientific text, 43% of the running words were found to be polysemantic (this in addition to syntactic ambiguities which the reader must solve on the basis of numerous alternatives given him in every sentence).
: The present study is a practical guide to editors who refine partially machine-translated text as a basis for linguistic analysis. The posteditors' tasks are: to code preferred English equivalents, to code English structural symbols, to resolve grammatic properties, and to code syntactic connections (dependencies). A general introduction to the field of machine translation is contained in RM-2060.