
This paper explores aspects of diachronic change in a non-native variety of English, Philippine English. It uses the Philippine section of the International Corpus of English (sampling period early 1990s) and a new corpus 'Phil-Brown', parallel in its design and sampling date (early 1960s) to the LOB and Brown corpora. Comparison is made between PhilE and the two super-varieties, British and American English, drawing from the pioneering work by Leech et al. (2009) on grammatical change in contemporary written English. The study focuses on relative clauses, more particularly that-relatives and wh-relatives. It was found that Philippine English has followed the two super-varieties in experiencing a decline in wh-relatives and an increase in that-relatives, but differs from them in the rapidity with which the changes have occurred, reflecting an attempt to approximate the patterns of its 'colonial parent, American English. When we compare the rates of change for the frequencies of that-relatives and wh-relatives across the genres we find more indications of Philippine English progressively aligning itself with American English, and in turn further evidence that the linguistic orientation of Philippine English remains predominantly exonormative.(1)
The present paper contrasts strategies of cohesive conjunction in English and German system and text. We clarify the notion of cohesive conjunction by discussing conceptualizations in the literature and by comparing cohesive conjunctions to other cohesive strategies. Using theory-informed methodologies we contrast the resources available in the two languages for explicitly establishing conjunctive relations of cohesion. Moreover, we discuss the first findings from our analysis of an English - German corpus of translations and originals, which reveal differences in the textual realizations in terms of frequencies and functions. Our study complements insights about other types of cohesion investigated in the frame of a larger research project.(1)
The crosslinguislic phenomenon of faux amis has been extensively studied in different fields of language research, such as translation (Granger and Swallow 1988, Vernal 2002, Malkiel 2006, Ruiz Mezcua 2008), lexicography (Hill 1982, Cuenca Villarejo 1987, Prado 2001, Postigo Pinazo, 2007), and second language acquisition research (Lengeling 1995, Frutos Martinez 2001, Wagner 2004, Chacon Beltran 2006). Faux amis (Koessler and Derocquigny 1928), also referred to as "false friends" (Zethsen 2004, Chacon Beltran 2006) or "deceptive cognates" (Lado 1957, Batchelor. and Offord 2000) are words which share similar forms in two or more languages but have different meanings and/or uses in each language (e.g. English carpet 'rug' versus Spanish carpeta 'folder'; English fabric 'cloth' versus French fabrique 'factory'; German Gift 'poison' versus English gift present). Despite the wide range of surveys, there is a conspicuous scarcity of studies which apply a corpus-based methodology to the investigation of these words; and none of the existing corpus-based studies explore the presence of English false friends in spoken learner language (Granger 1996, Palacios Martinez and Alonso 2005). The present study aims at filling this void by examining 100 high:frequency English false friends in the spoken and written performance of Spanish learners of English through an analysis of two learner corpora, namely the International Corpus of Learner English (ICLE) and the Louvain International Database of Spoken English Interlanguage (LINDSEI). The data obtained from these corpora allow us to draw conclusions about the learners' active-use of these lexical items in speech and writing. A total amount of 1403 sample sentences have been closely examined. My analysis reveals that EFL learners make more errors with false friends in their written than in their spoken production and it also shows that certain English false friends are especially difficult for learners (e.g. actually, pretend, argument). Thus, the findings of this study certainly shed some light on students' problems in this lexical area which should be addressed in an EFL context.(1)
In the present study we use the South Asian and Southeast Asian components of the International Corpus of English (ICE) as well as a larger set of web-derived newspaper corpora in order to account for similarities and differences in the use of the progressive in eight non-native varieties of English (Singapore, Hong Kong, Indian, Sri Lankan, Bangladeshi, Maldivian, Nepali and Pakistani English). The analysis is two-fold: First, for the varieties under scrutiny, we provide a quantitative overview of the different use of progressive aspect marking according to tense and voice. Second, we apply a cluster analysis in order to determine respective variety-based clusters so that a comparison of those verbs of which progressive aspect marking is said to differ most significantly will be feasible. In a third step we present examples of some of the most influential verbs that are responsible for the differences in use across the corpora.
The study of EFL writing has so far not been able to go beyond the observation and discussion of group characteristics. While it has been possible to study developmental data, in the sense that writing products of students at various levels of proficiency were available for research, longitudinal EFL data were virtually non-existent. The LONGDALE project seeks to find an answer to the question how EFL writing develops over time. This article reports on a quantitative and a qualitative study of Dutch EFL writing, based on a modest amount of longitudinal data. Its aim is to find out if and how non-native writing develops over time, whether it develops in the direction of native writing, and whether individual students display individual developmental patterns. The answer to all three questions appears to be affirmative.
This study of word-stress variation relies on a sub-corpus of more than 2,000 word entries in which stress variants appear. They were extracted from a computer-searchable version of John Wells' first edition of the Longman Pronunciation Dictionary (henceforth LPDI, Wells 1990). With the help of a selection of examples, I want to demonstrate that word-stress variation is the result of conflicting rules that are indicators of simplification in the phonology of British English. The theoretical framework adopted here is Lionel Guierre's study of word stress (Guierre 1979) that includes a close examination of exhaustive results automatically obtained from a computer-searchable version of Daniel Jones's twelfth edition of the English Pronouncing Dictionary (Jones 1963). More often than not, Guierre's Normal Stress Rule (hence NSR) is involved in the conflicts.' A close examination of the variations shows that NSR stressing appears to be the new variant that challenges traditional stress patterns. We may call this process "NSR regularisation". Word stress variation is undoubtedly symptomatic of ongoing changes that are phonetically abrupt, as exemplified by the poll preferences in LPD, and lexically gradual since some words in identified lexical paradigms are not affected by the changes. In this study, directions of the changes are determined with the help of diachronic data extracted from pronouncing dictionaries of the 18th and 19th centuries. After such identifications of stress variations within the lexicon, further research needs to be carried out to put dictionary data to the test.
Drawing on the British Academic Spoken English (BASE)1 corpus, this paper presents an overview of how lecturers mark important and less important discourse using verbal cues. Such relevance markers (e.g. the point is, remember, that is important, essentially) and markers of lesser relevance (e.g. anyway, a little bit, not go into, not write down) combine discourse organisation with evaluation and can help students discern the relative importance of points, thus aiding comprehension, note-taking and retention. However, until the research reported here was undertaken, little was known about this metadiscursive feature of lecture discourse, and markers found in the existing literature and EAP materials were rather few and typically not based on corpus linguistic evidence. Combining corpus based and corpus-driven methods, the research started from a close reading of 40 lectures to identify potential markers. These were next retrieved from the whole corpus and in the case of relevance markers supplemented with other approaches yielding further markers. Relevance markers were mainly classified into lexicogrammatical verb, noun, adjective and adverb patterns, while the markers of lesser relevance were classified pragmatically as indications of message status, topic treatment, lecturer knowledge, assessment, and attention and note-taking directives. This account of relevance marking is valuable for EAP practitioners, will interest lecturer trainers, and provides input for experimental research on lecture listening, note-taking and lecture effectiveness. Furthermore, the paper offers insights into the use of discourse markers such as "the thing is", "anyway", "I don't know" and "et cetera" and illuminates the understudied linguistic phenomenon of relevance marking. Finally, it illustrates the importance of corpus linguistic research and touches upon some difficulties pertaining to assigning discourse functions based on an examination of transcripts only.
This paper gives an account of plans for constructing a searchable database of eighteenth century English phonology. The project incorporates data from pronouncing dictionaries and other texts dealing with pronunciation published in the second half of the 18th century. The data will be recorded in the form of Unicode transcriptions of as many of the approximately 1,700 words used to exemplify John Wells' (1982) Standard Lexical Sets as appear in the eighteenth-century texts. Although all the eighteenth-century texts purported to describe the 'best' English, they were compiled by authors from different parts of the English-speaking world (mainly different regions of England, Scotland and Ireland but including some from North America) and so can provide evidence for geographical diffusion of innovations. (Beal 1999, C. Jones 2006). This paper provides an account of the design of this database and presents the results of a pilot study demonstrating how such a database can be used to answer questions concerning the chronological, social, geographical and phonological distribution of variation between /hw/ similar to /w/ similar to /h/ in WHICH, WHO, NOWHERE, etc. which is of interest to sociolinguists, dialectologists and historical phonologists.
This paper deals with the way learners make use of the demonstratives this and that. NLP tools are applied to classes occurrences of native and non-native uses of the two forms. The objective of the two experiments is to automatically identify expected and unexpected uses. The textual environment of all the instances is explored at text and PoS level to uncover features which play a role in the selection of a particular form. Results of the first experiment show that the PoS features predeterminer and determiner, which are found in the immediate context, help identify; unexpected learner uses among many occurrences, including native uses. The second experiment provides evidence that the PoS features plural noun and coordinating conjunction influence the unexpected uses of the demonstratives by learners. This study shows that NLP tools can be used to explore texts and uncover underlying grammatical categories that play a role in the selection of specific words.
This paper presents a collaborative project that focuses on letters of artisans and the labouring poor in England, c. 1750-1835 (LALP). The project’s objective is to create a corpus that allows for new research perspectives regarding the diachronic development of the English language by adding data representing language of the lower classes. An opportunity for an insight into the language use of the labouring poor has been provided by the laws for poor relief, which permitted people in need to apply for out-relief from parish funds during the period 1795-1834. For the last 18 years, the independent scholar Tony Fairman has collected and transcribed more than 2000 poor relief application letters and other letters by artisans and the labouring poor. In this project Fairman’s letter collection is being converted into an electronic corpus. Apart from converting the material into electronic form, the transcribed texts will be supplemented with contextual information and manuscript images. This paper presents the letter material, it describes the conversion of the letter collection into a corpus and discusses some of the problems and challenges in the conversion process.1
Dutch students of English at Radboud University, Nijmegen, the Netherlands are believed to enter university at CEFR B2 and, on graduation, are expected to have reached CEFR C2 in reading, writing, listening, spoken production and interaction. There is, however, preciously little evidence that links students' proficiency to the actual CEFR. This is hardly surprising as the difficulties of linking language users' production to specific CEFR levels are well-known. The English department does not really systematically chart students' progress in language proficiency, so that no documentation of the development of students' language production is available.As a first attempt at finding out whether or not it is possible to measure students' progress in spoken English over the first two years of their degree course objectively, a small pilot study of a corpus of spoken English was undertaken. The participants were 31 students from a single cohort who were recorded in their first week at university and at the end of their first and second year. On all occasions they were asked to respond to a set of written general questions, such as "have you read a good book lately?".