Abstract The present study investigates four types of core fluencemes (filled and unfilled pauses, discourse markers and repeats) in four components of the Louvain International Database of Spoken English Interlanguage (Gilquin et al. 2010). We test these fluencemes for (1) possible transfer effects, (2) the effect of task type, (3) the influence of the communicative behavior of the interlocutor and (4) other linguistic and non-linguistic variables. Our findings reveal L1 transfer effects for all variables, while learning context variables have a stable positive effect on learner fluency. Task type does not reveal unidirectional effects, while the communicative behavior of the interlocutor turned out to have a significant effect on the learners’ performance, indicating that processes of ‘confluence’ are at play.
To overcome planning phases in spontaneous speech production, learners and native speakers use strategies such as (un)filled pauses, smallwords or discourse markers. Small scale studies in this vein have demonstrated that learners differ from native speakers in that they underuse smallwords and discourse markers, and rely on other fluency-enhancing strategies instead. In the present paper, we present a corpus-based study, which investigates fluency-enhancing strategies in four components of the Louvain International Database of Spoken English Interlanguage (LINDSEI; Gilquin et al. 2010), covering four learner English varieties, namely Spanish, German, Bulgarian and Japanese. We investigate 216 different fluencemes (i.e. fluency-enhancing features; Götz in Fluency in native and nonnative English speech, John Benjamins, Amsterdam, 2013) in 200 transcribed interviews with advanced learners of English. An online coding application, which was specially designed and programmed for this project, enables us to cover such a large amount of data. We report on the design, functionality and (dis-)advantages of the online application, the multilevel-coding system we implemented, and the methodological challenges we face in detail. We will also present the findings of one first pilot study where we exhibit considerable variation between and within learners of particular native languages concerning fluenceme frequencies, while distributional patterns of fluencemes are rather similar across varieties.
Researchers in dialectometry have begun to explore measurements based on fundamentally quantitative metrics, often sourced from dialect corpora, as an alternative to the traditional signals derived from dialect atlases. This change of data type amplifies an existing issue in the classical paradigm, namely that locations may vary in coverage and that this affects the distance measurements: pairs involving a location with lower coverage suffer from greater noise and therefore imprecision. We propose a method for increasing robustness using generalized additive modeling, a statistical technique that allows leveraging the spatial arrangement of the data. The technique is applied to data from the British English dialect corpus FRED; the results are evaluated regarding their interpretability and according to several quantitative metrics. We conclude that data availability is an influential covariate in corpus-based dialectometry and beyond, and recommend that researchers be aware of this issue and of methods to alleviate it.
Traditional dialects have been encroached upon by the increasing mobility of their speakers and by the onslaught of national languages in education and mass media. Typically, older dialects are “leveling” to become more like national languages. This is regrettable when the last articulate traces of a culture are lost, but it also promotes a complex dynamics of interaction as speakers shift from dialect to standard and to intermediate compromises between the two in their forms of speech. Varieties of speech thus live on in modern communities, where they still function to mark provenance, but increasingly cultural and social provenance as opposed to pure geography. They arise at times from the need to function throughout the different groups in society, but they also may have roots in immigrants’ speech, and just as certainly from the ineluctable dynamics of groups wishing to express their identity to themselves and to the world. The future of dialects is a selection of the papers presented at Methods in Dialectology XV, held in Groningen, the Netherlands, 11-15 August 2014. While the focus is on methodology, the volume also includes specialized studies on varieties of Catalan, Breton, Croatian, (Belgian) Dutch, English (in the US, the UK and in Japan), German (including Swiss German), Italian (including Tyrolean Italian), Japanese, and Spanish as well as on heritage languages in Canada.
This article explores measures, operationalisations and effects of rhythm and weight as two constraints on the variation between the s-genitive and the of-genitive. We base the analysis on interchangeable genitives in the news and letters sections of ARCHER (A Representative Corpus of Historical English Registers), which covers the period between 1650 and 1999. Thus, we are ultimately concerned with the applicability of two factors that have their roots in speech (rhythm: phonology; weight: online processing) to an ‘unconventional’, written data set with a historical dimension. As for weight, we focus on the comparison of simple single-constituent and more complex multi-constituent measurements. Our notion of rhythm centres on the ideally even distribution of stressed and unstressed syllables. We find that in our data set, both rhythm and weight show theoretically unexpected quadratic effects: rhythmically better-behaved s-genitives are not necessarily preferred over of-genitives, and short constituents exhibit odd weight effects. In conclusion, we argue that while rhythm is only a minor player in our data set, the quadratic quirks it exhibits should inspire further study. Weight, on the other hand, is a crucial factor which, however, likewise comes with measurement and modelling complications.
Historically speaking, the of-genitive is of course the incoming form, which app eared during the ninth century. According to Thomas (1931, 284), the inflected genitive vastly outnumbered the periphrasis with of up until the twelfth century. In the Middle Englis h period, we begin to witness “a strong tendency to r eplace the inflectional genitive by periphrastic constructions, above all by periphrasi s with the preposition of” (Mustanoja 1960, I:70). The Early Modern English period, however, se es a revival of the s-genitive, “against all odds” (Rosenbach 2002, 184). While we know that the s-g nitive is comparatively – and increasingly – popular in Present-Day English, espe cially American English (Rosenbach 2002; Rosenbach 2003), the literature about genitiv variability in the Late Modern English period is somewhat sketchy (but see Szmrecsanyi 201 3; Wolk et al. 2013).
We present a cross-constructional approach to the history of the genitive alternation and the dative alternation in Late Modern English (AD 1650 to AD 1999), drawing on richly annotated datasets and modern statistical modeling techniques. We identify cross-constructional similarities in the development of the genitive and the dative alternation over time (mainly with regard to the loosening of the animacy constraint), a development which parallels distributional changes in animacy categories in the corpus material. Theoretically, we transfer the notion of `probabilistic grammar' to historical data and claim that the corpus models presented reflect past speakers' knowledge about the distribution of genitive and dative variants. The historical data also helps to determine what is constant (and timeless) in the effect of selected factors such as animacy or length, and what is variant.
Historically speaking, the of-genitive is of course the incoming form, which appeared during the ninth century. According to Thomas (1931, 284), the inflected genitive vastly outnumbered the periphrasis with of up until the twelfth century. In the Middle English period, we begin to witness “a strong tendency to replace the inflectional genitive by periphrastic constructions, above all by periphrasis with the preposition of” (Mustanoja 1960, I:70). The Early Modern English period, however, sees a revival of the s-genitive, “against all odds” (Rosenbach 2002, 184). While we know that the s-genitive is comparatively – and increasingly – popular in Present-Day English, especially American English (Rosenbach 2002; Rosenbach 2003), the literature about genitive variability in the Late Modern English period is somewhat sketchy (but see Szmrecsanyi 2013; Wolk et al. 2013).
Acquiring English dative verbs: proficiency effects in German L2 learners Christoph Wolk * (christoph.wolk@frias.uni-freiburg.de), Sascha Wolfer † , Peter Baumann † , Barbara Hemforth ‡ , Lars Konieczny *† * Freiburg Institute for Advanced Studies, Starkenstr. 44 D-79104 Freiburg i. Br., Germany † Center for Cognitive Science, University of Freiburg, Friedrichstr. 50 D-79098 Freiburg i. Br., Germany ‡ Laboratoire de Psychologie et de Neuropsychologie Cognitives, CNRS, Universite Paris Descartes, 71 ave Edouard Vaillant, 92100 Boulogne-Billancourt, France Abstract This paper investigates the influence of probabilistic informa- tion in the second language on the processing of English da- tive alternation constructions in German learners of English. We present two eye-tracking studies (visual world and reading) with evidence that the probabilistic patterns of the target lan- guage influence L2 processing when the initial preference is vi- olated, and indications that these patterns have a greater effect on more experienced speakers. We also observed a constrast- effect of L1, such that comprehenders expected constructions that occur more often in L2 than in L1, even if L2 lexical statis- tics suggested otherwise. Keywords: Sentence processing, dative alternation, second language acquisition, expectation-based language processing Introduction In many languages, semantically dative sentences can be re- alized with two different object orders that only slightly dif- fer in meaning, one in which the recipient comes before the theme and one where the reverse is true. In English, the for- mer ordering is achieved by two bare noun phrases, as in (1-a), and the latter by having the recipient as a prepositional phrase, as in (1-b). (1) a. b. double object dative (DO) I gave [her] recipient [the book] theme . prepositional dative (PO) I gave [the book] theme to [her] recipient . The dative alternation has received considerable attention from first- and second language acquisition researchers dur- ing the 1980s, especially from the perspective of genera- tive grammar. These studies focused primarily on investigat- ing the following two questions by means of grammatical- ity judgments and sentence completion tasks: First, how well do learners acquire hard constraints on the possibility of al- ternation, such as the fixed prepositional realization of most verbs of Latin origin such as donate; second, what is the order in which speakers acquire the possible realizations for verbs that do alternate. Major results (e.g. in Mazurkewich, 1985; Mazurkewich & White, 1984) were that verb-specific con- straints are acquirable as hard constraints for first language learners with rare errors, but are only learned as softer con- straints — or sometimes not learned at all — for second lan- guage learners. With regard to acquisition order, the preposi- tional dative realization tends to be acquired earlier and easier for second language learners. Recent research on first language (L1) dative alternation patterns, however, has switched the focus from presumably ’hard’ constraints on the possibility of alternation to the softer, probabilistic determinants of actually observed vari- ation. This was motivated by cross-linguistic similarities in grammatical preferences (Bresnan, Dingare, & Manning, 2001) as well as the fact that in both naturally occurring lan- guage and experimental investigation ’hard’ constraints turn out to be surprisingly violable (Bresnan & Nikitina, 2008; Bresnan, 2007) while simultaneous consideration of multi- ple ’soft’ constraints led to considerable success in predic- tion of realizations, reading time, and fluency of production (Bresnan, Cueni, Nikitina, & Baayen, 2007; Bresnan & Ford, 2010; Tily et al., 2009). With regard to acquisition, children have been shown to mirror the probabilistic realization pat- terns of their environment (Marneffe, Grimm, Arnon, Kirby, & Bresnan, to appear). Second language studies within this probabilistic paradigm, however, are still rare; one exception is the study by Frishkoff, Levin, Pavlik, Idemaru, and Jong (2008), who used the results of (Bresnan et al., 2007) to in- vestigate how both native and second language (L2) speak- ers learn to predict dative choice from examples, and found that L2 learners improve quickly when presented with stimuli containing a high degree of contrast between alternation pref- erences. Individual factors that were found to be reliable pre- dictors for L1 speakers in corpus models have, however, also been shown to influence L2 learning at various proficiency levels. These include among others pronominality (Le Com- pagnon, 1984), givenness and persistence (Marefat, 2005), and weight (Tanaka, 1968; Callies & Szczesniak, 2008). The goal of this paper is to investigate how attuned L2 learners are to fine probabilistic details of their target lan- guage. One predictor that is, due to its inherently probabilis- tic nature, especially suited for this research question is verb bias. More specifically, each dative verb has a specific id- iosyncratic degree of preference in alternation choice; this preference is in general not predictable from semantics or morphology. A learner’s acquisition of verb bias should thus be seen as direct instances of fundamentally experience-based learning. The experiments reported here are based on English L2 learners with German as L1. Like English, and unlike most L1s of previous studies, German has a double object da- tive; in contrast, the use of the prepositional dative is limited
This paper is concerned with sketching future directions for corpus-based dialectology. We advocate a holistic approach to the study of geographically conditioned linguistic variability, and we present a suitable methodology, 'corpusbased dialectometry', in exactly this spirit. Specifically, we argue that in order to live up to the potential of the corpus-based method, practitioners need to (i) abandon their exclusive focus on individual linguistic features in favor of the study of feature aggregates, (ii) draw on computationally advanced multivariate analysis techniques (such as multidimensional scaling, cluster analysis, and principal component analysis), and (iii) aid interpretation of empirical results by marshalling state-of-the-art data visualization techniques. To exemplify this line of analysis, we present a case study which explores joint frequency variability of 57 morphosyntax features in 34 dialects all over Great Britain.
Ziel des vorliegenden Beitrags ist es, die Ergebnisse einer vorangehenden, informellen Studie zur Nutzung von Vorlesungsaufzeichnungen durch ”harte Fakten“ zu uberprufen und gegebenenfalls neue Erkenntnisse zu gewinnen. Es wird kurz ein Werkzeug vorgestellt, mit welchem wir die Zugriffe der Studierenden auf die Vorlesungsaufzeichnungen untersucht haben. Die aufschlussreichsten Analysen werden in diesem Beitrag vorgestellt. Es ergeben sich hierbei interessante Ergebnisse bezuglich der Verwendung der Materialien durch die Studierenden oder der Nachfrage nach verschiedenen Medienformaten. Auch das immer wieder kontrovers diskutierte Thema, ob Vorlesungsaufzeichnungen mit Dozentenvideo besser geeignet sind als Aufzeichnungen ohne das Video, wird von uns aufgegriffen und die Position unserer Studierenden zu dieser Thematik anhand der Logfileanalyse dargelegt. Des Weiteren diskutieren wir das Thema der Archivierung von Vorlesungsaufzeichnungen und untersuchen, zu welchen Zeitpunkten Studierende besonders auf Vorlesungsaufzeichnungen als Lernmaterial zuruckgreifen.
We present three approaches to corpus-based dialectometry and apply them to morphosyntactic variation in the Freiburg Corpus of English Dialects, which covers 34 counties throughout Great Britain. Two of these are top-down approaches that start with a predefined feature list; one using a straightforward frequency-based analysis, the other enhancing the raw numbers using probabilistic modeling. Both methods are able to detect the structure of areal variation in Great Britain, and the second approach is able to reduce the influence of textual coverage as a nuisance factor. The final approach is a bottom-up method that eschews pre-specified lists and evaluates potential features directly from the data using a permutation-based metric. Again, we find that simple frequency-based metrics are biased, but that derivedmetrics yield a clearer pattern. Using thesemethods, we are able to uncover significant geolinguistic structure in Great Britain.