International Encyclopedia of Social & Behavioral Sciences(1994)
被引用3147|浏览27
摘要
Corpus linguistics encompasses the compilation and analysis of collections of spoken and written texts as the source of evidence for describing the nature, structure, and use of languages. This work typically brings a quantitative dimension to the description of languages by including information on the probability with which linguistic items or processes occur in particular contexts. Corpora vary greatly in size and design but most are nowadays in electronic form with purpose-built computer software to support analysis. Present-day corpus linguistics has grown out of a long tradition of using texts as the empirical basis for linguistic description, studying all levels of language, including phonology, lexis, grammar, and discourse. Corpora are often annotated to show grammatical classes and functions. Software to analyze grammatical structures or to identify collocations by means of concordancing has revolutionized text analysis. In focusing on the company which particular words tend to keep and on the ways in which language users habitually express themselves, corpus linguistics has thrown new light on how languages vary systematically in different historical, regional, and sociolinguistic contexts, genres, and registers. Probabilistic descriptions of languages can complement other methodologies used by linguists and have implications for work in a number of fields in addition to linguistic description. These include natural language processing, language education, and cognitive linguistics.