This article presents a case study on the contributions of prepositional particles to the meanings of German particle verbs (such as anstrahlen ‘to beam/smile at’ and aufgeben ‘to give up’). Based on a set of 16 “concept images”, two-dimensional directional arrow pictographs, 60 experiment participants selected one or more concept images for a systematically composed set of 270 German particle verbs and their 30 base verbs. We formulate a series of hypotheses for the meanings of nine constituent particle types (ab, an, auf, aus, ein, mit, nach, vor, zu) and investigate them in the light of the concept image selections. Qualitative and quantitative analyses indicate that our hypotheses are largely confirmed, across three source domains varying in their abstractness (Machines & Tools, Force, Sound), as well as across well-known vs. unknown particle verbs. The particles exhibit individual concept image profiles, and they vary in their flexibility to provide predominant directions; for example, while auf is rather consistently perceived as contributing an upward/right direction to a particle verb meaning, an shows similarly strong preferences for a set of concept images; in both cases, these tendencies are observed across source domains.
Abstract Most German particle verbs (PVs) are composed of a prepositional particle (P) and a base verb (BV). For instance, anstrahlen is formed from the P an and the BV strahlen. The meaning of a PV results from often systematic interactions between the P and BV meanings. But many Ps and BVs are ambiguous and, moreover, a single P meaning and a single BV meaning can be combined in several ways. Finally, the interactions between P and BV meanings depend on the context. This chapter presents a case study of how two particles, auf and ab, interact with certain BVs and contextual factors, while focusing on the difference between abstract and concrete BV/PV concepts, and between abstract and concrete contexts.
Multiword expressions (MWEs), such as noun compounds (e.g. nickname in English, and Ohrwurm in German), complex verbs (e.g. give up in English, and aufgeben in German) and idioms (e.g. break the ice in English, and das Eis brechen in German), may be interpreted literally but often undergo meaning shifts with respect to their constituents. Theoretical, psycholinguistic as well as computational linguistic research remain puzzled by when and how MWEs receive literal vs. meaning-shifted interpretations, what the contributions of the MWE constituents are to the degree of semantic transparency (i.e., meaning compositionality) of the MWE, and how literal vs. meaning-shifted MWEs are processed and computed. This edited volume presents an interdisciplinary selection of seven papers on recent findings across linguistic, psycholinguistic, corpus-based and computational research fields and perspectives, discussing the interaction of constituent properties and MWE meanings, and how MWE constituents contribute to the processing and representation of MWEs. The collection is based on a workshop at the 2017 annual conference of the German Linguistic Society (DGfS) that took place at Saarland University in Saarbrucken, Germany
This paper presents a collection to assess meaning components in German complex verbs, which frequently undergo meaning shifts. We use a novel strategy to obtain source and target domain characterisations via sentence generation rather than sentence annotation. A selection of arrows adds spatial directional information to the generated contexts. We provide a broad qualitative description of the dataset, and a series of standard classification experiments verifies the quantitative reliability of the presented resource. The setup for collecting the meaning components is applicable also to other languages, regarding complex verbs as well as other language-specific targets that involve meaning shifts.
This paper discusses an extension of the V-measure (Rosenberg and Hirschberg, 2007), an entropy-based cluster evaluation metric. While the original work focused on evaluating hard clusterings, we introduce the Fuzzy V-measure which can be used on data that is inherently ambiguous. We perform multiple analyses varying the sizes and ambiguity rates and show that while entropy-based measures in general tend to suffer when ambiguity increases, a measure with desirable properties can be derived from these in a straightforward manner.
This paper presents a token-based automatic classification of German perception verbs into literal vs. multiple non-literal senses. Based on a corpus-based dataset of German perception verbs and their systematic meaning shifts, we identify one verb of each of the four perception classes optical, acoustic, olfactory, haptic, and use Decision Trees relying on syntactic and semantic corpus-based features to classify the verb uses into 3-4 senses each. Our classifier reaches accuracies between 45.5% and 69.4%, in comparison to baselines between 27.5% and 39.0%. In three out of four cases analyzed our classifier’s accuracy is significantly higher than the according baseline.
This paper provides a corpus-based study on German particle verbs. We hypothesize that there are regular mechanisms in meaning shifts of a base verb in combination with a particle that do not only apply to the individual verb, but across a semantically coherent set of verbs. For example, the syntactically similar base verbs brummen ‘hum’ and donnern ‘rumble’ both describe an irritating, displeasing loud sound. Combined with the particle auf, they result in near-synonyms roughly meaning ‘forcefully assigning a task’ (in one of their senses). Covering 6 base verb groups and 3 particles with 4 particle meanings, we demonstrate that corpus-based information on the verbs’ subcategorization frames plus conceptual properties of the nominal complements is a sufficient basis for defining such meaning shifts. While the paper is considerably more extensive than earlier related work, we view it as a case study toward a more automatic approach to identify and formalize meaning shifts in German particle verbs.
Compositionality of German particle verbs German particle verbs (PVs) are highly productive combinations of a base verb and a prefix particle. Concerning their semantics, there is an ongoing discussion whether the meaning of German particle verbs is in general compositional or not. For example, Kratzer (2003) claimed that German PVs are idiosyncratic; this stands in opposition to the semantic analyses by Lechler and Roßdeutscher (2009), Kliche (2011), among others. who demonstrated that each particle has several different readings which however form regular patterns depending on the contexts. Our position is in-between: we agree that not every PV composition is transparent, but with a fine-grained sub-lexical analysis and taking analogy and meaning shift mechanisms into account, the majority of combinations can be explained by patterns. Our research focuses on how speakers of German combine particle senses with base verb senses. Questions which come along with this focus are: (i) how applicable, (ii) how available and (iii) how common or prototypical is a semantic pattern of a meaning composition?
For many NLP applications such as Information Extraction and Sentiment Detection, it is of vital importance to distinguish between synonyms and antonyms. While the general assumption is that distributional models are not suitable for this task, we demonstrate that using suitable features, differences in the contexts of synonymous and antonymous German adjective pairs can be identified with a simple word space model. Experimenting with two context settings (a simple windowbased model and a ‘co-disambiguation model’ to approximate adjective sense disambiguation), our best model significantly outperforms the 50% baseline and achieves 70.6% accuracy in a synonym/antonym classification task.
This paper presents a methodology to identify polysemous German prepositions by exploring their vector spatial properties. We apply two cluster evaluation metrics (the Silhouette Value (Kaufman and Rousseeuw, 1990) and a fuzzy version of the V-Measure (Rosenberg and Hirschberg, 2007)) as well as various correlations, to exploit hard vs. soft cluster analyses based on Self-Organising Maps. Our main hypothesis is that polysemous prepositions are outliers, and thus represent either (i) singletons or (ii) marginals of the clusters within a cluster analysis. Our analyses demonstrate that (a) in a subset of the clusterings, singletons have a tendency to contain polysemous prepositions; and (b) misclassification and cluster membership rate exhibit a moderate correlation with ambiguity rate.
The current study works at the interface of theoretical and computational linguistics to explore the semantic properties of an particle verbs, i.e., German particle verbs with the particle an. Based on a thorough analysis of the particle verbs from a theoretical point of view, we identified empirical features and performed an automatic semantic classification. A focus of the study was on the mutual profit of theoretical and empirical perspectives with respect to salient semantic properties of the an particle verbs: (a) how can we transform the theoretical insights into empirical, corpus-based features, (b) to what extent can we replicate the theoretical classification by a machine learning approach, and (c) can the computational analysis in turn deepen our insights to the semantic properties of the particle verbs? The best classification result of 70% correct class assignments was reached through a GermaNet-based generalization of direct object nouns plus a prepositional phrase feature. These particle verb features in combination with a detailed analysis of the results at the same time confirmed and enlarged our knowledge about salient properties.
The compositionality of particle verbs is a matter of dispute. The following paper will show that the semantics of German particle ver bs is the combination of verb semantics and particle semantics. For this purpose, the verb particle an was analyzed and turned out to be highly ambiguous. Still, it ca n be systematically reconstructed. The different meanings resulting out of this ambigu ity were grouped into semantic classes that were then formalized by using the Disc ourse Representation Theory. The classification of these meanings is based on a deta iled case study that thoroughly defines the distinct components and strengthens my above mentioned thesis.