
The dependency surface-syntactic structure is proposed, within the Meaning-Text framework, for binary conjunctions of the IF–THEN type; e.g.: IF→Y, THEN←X A universal typology of conjunctions is sketched, and three examples of English binary conjunctions are given. Binary conjunctions are “discontinuous” phrasemesidioms, collocations and formulemes that have to be considered together with their actants, since there are no direct syntactic links between their components. Full lexical entries for two Russian binary conjunctions are presented, supplied with linguistic comments, and deep-syntactic rules ensuring the expansion of a deep-syntactic binary conjunction node into the corresponding surface-syntactic tree are illustrated. 1 The Syntactic Structure of a Binary Conjunction This paper examines subordinating and coordinating binary conjunctions (or correlative subordinators/coordinators, as they are known in the literature: Quirk et al. 1991: 935–941, 999– 1001). The typical examples are the subordinating conjunction IF..., THEN... and the coordinating conjunction EITHER..., OR... The discussion is carried out within the Meaning-Text approach (see Mel’čuk 1974, 2012, 2016b). In sentence (1) dependency relations between lexemes are obvious, except for THEN, the second component of the conjunction IF..., THEN...: (1) If A→and→B are→equal, then B←follows→C. The dependency for THEN is proposed in what follows. Without THEN the superordinate clause can linearly precede or follow the subordinate clause with IF; but with THEN it can only follow. This gives the idea to make this THEN dependent on IF: IF–r→THEN; as a result, the binary conjunction IF..., THEN... can be stored in the lexicon exactly in the form of this syntactic subtree. Such a description had been tacitly accepted for almost half a century: • In Mel’čuk 1974: 231, No. 31, (e), the surfacesyntactic relation [SSyntRel] r between IF and THEN was called “1st auxiliary.” • In Mel’čuk & Pertsov 1987: 331, No. 19.1, it was rebaptized “binary-junctive.” • In Iomdin 2010: 43, it appears under the name of “correlative SSyntRel.” • In Mel’čuk 2012a: 143, No. 51, it is “correlativeauxiliary.” However, this syntactic description of binary conjunctions contradicts the definition of surface-syntactic dependency (or, more precisely, that of surface-syntactic relation), which was advanced in Mel’čuk 1988: 130–144 and has been used as such since; see its newer formulations, for instance, in Mel’čuk 2009: 25–40 and Mel’čuk 2015b: 411–433. In order to lay bare this contradiction, only the first part of this definition—namdely Criterion A—is needed, strictly speaking. Nevertheless, to facilitate the task of the reader I will cite here the whole definition—that is, the full set of criteria for SSyntRels. (Of course many substantial explanations and interesting special cases have to be bypassed.) Proceedings of the Fourth International Conference on Dependency Linguistics (Depling 2017), pages 127-134, Pisa, Italy, September 18-2
Copula constructions are problematic in the syntax of most languages. The paper describes three different dependency syntactic methods for handling copula constructions: function head, content head and complex label analysis. Furthermore, we also propose a POS-based approach to copula detection. We evaluate the impact of these approaches in computational parsing, in two parsing experiments for Hungarian.
This paper introduces UDLex , a computational framework for the automatic extraction of argument structures for several languages. By exploiting the versa-tility of the Universal Dependency annotation scheme, our system acquires subcategorization frames directly from a dependency parsed corpus, regardless of the input language. It thus uses a universal set of language-independent rules to detect verb dependencies in a sentence. In this paper we describe how the system has been developed by adapting the LexIt (Lenci et al., 2012) framework, originally designed to describe argument structures of Italian predicates. Practical issues that arose when building argument structure representations for typologically different languages will also be discussed.
Neural network (“deep learning”) models are taking over machine learning approaches for language by storm. In particular, recurrent neural networks (RNNs), which are flexible non-markovian models of sequential data, were shown to be effective for a variety of language processing tasks. Somewhat surprisingly, these seemingly purely sequential models are very capable at modeling syntactic phenomena, and using them result in very strong dependency parsers, for a variety of languages.
The paper looks into the expression of intensification with parametric nouns such as PRICE, COST, FEE, RATE, etc., focusing on collocations these nouns form with intensifying adjectives, inchoative and causative intensifying verbs and corresponding de-verbal nouns. Degrees of intensification possible with these nouns are discussed, as well as analytical vs. synthetic expression of intensification (a steep increase in prices ~ a spike in prices). Sample lexicalization rules are proposed—namely, rules that map semantic representations of intensifier collocations headed by nouns of this type to their deepsyntactic representations. The theoretical framework of the paper is Meaning-Text linguistic theory. 1 The Problem Stated The paper looks into the expression of intensification with parametric nouns such as PRICE, COST, FEE, RATE, etc., hereafter PRICE type nouns, or {NPRICE} for short (see Table 1, Section 3 below). More precisely, it describes collocations these nouns form with intensifying adjectives, as well as with inchoative and causative intensifying verbs and corresponding deverbal nouns. A cursory comparison is provided with antonymic, i.e., attenuating, expressions entering in collocations with {NPRICE}. A parametric noun (cf. Mel’čuk, 2013: 214) corresponds to (at least) a two-place predicate, ‘P of X is α’, with X being the thing parameterized and α, the value of the parameter: the priceP [of gas]X is [$1.85 per gallon]α, the speedP [of the vehicle]X is [70 miles per hour]α, the quantityP [of oil]X is [30 tons]α, etc. The α value may not be explicitly quantified, but characterized as being big or small (on some scale): The price of gas is high. | The speed of the vehicle is low. | The quantity of oil is huge. | Etc. I will be interested namely in the case where α of an NPRICE, without being explicitly quantified, is qualified as high, or ‘big’ [STATIVE], or rising—‘getting bigger’—[INCHOATIVE], or else being caused to rise [CAUSATIVE]. These cases are illustrated, respectively, in (1), (2) and (3); the examples come from Google searches (some have been slightly modified). (1) STATIVE: ‘⟦P of X being α,⟧ α is (very) big’, etc. a. Post-paid service plans often charge steep 〈astronomical, prohibitive〉 overage FEES. b. California divorce COST is high 〈whooping high, exorbitant〉. (2) INCHOATIVE: ‘⟦P of X being α,⟧ α begins to be bigger than αʹ by β (β being big)’ a. Electricity COSTS went up 〈rose sharply, surged, skyrocketed〉 in August. b. Make sure your mortgage payments do not increase1 if there is a rise 〈a major hike, a spike〉 in interest RATES. 1 An NPRICE parametric noun typically has additional dependents; thus, the person who determines the price of something corresponds to an argument (in our terms, semantic actant) of PRICE; similarly, the person who incurs the cost of something corresponds to a semantic actant of COST; FEE has two additional semantic atants: the one who sets it and the one who pays it; and so on. These actants are not directly relevant for the present discussion. Proceedings of the Fourth International Conference on Dependency Linguistics (Depling 2017), pages 145-153, Pisa, Italy, September 18-2
According to the Menzerath-Altmann law, there is a relation between the size of the whole and the mean size of its parts. The validity of the law was demonstrated on relations between several language units, e.g., the longer a word, the shorter the syllables the word consists of. In this paper it is shown that the law is valid also in syntactic dependency structure in Czech. In particular, longer clauses tend to be composed of shorter phrases (the size of a phrase is measured by the number of words it consists of).
Word embeddings induced from large amounts of unannotated text are a key resource for many NLP tasks. Several recent studies have proposed extensions of the basic distributional semantics approach where words form the context of other words, adding features from e.g. syntactic dependencies. In this study, we look in a different direction, exploring models that leave words out entirely, instead basing the context representation exclusively on syntactic and morphological features. Remarkably, we find that the resulting vectors still capture clear semantic aspects of words in addition to syntactic ones. We assess the properties of the vectors using both intrinsic and extrinsic evaluations, demonstrating in a multilingual parsing experiment using 55 treebanks that fully delexicalized syntax-based word representations give a higher average parsing performance than conventional word2vec embeddings.
This contribution presents a dependency grammar (DG) analysis of the so-called descriptive and resultative V-de constructions in Mandarin Chinese (VDCs); it focuses, in particular, on the dependency analysis of the noun phrase that intervenes between the two predicates in a VDC. Two methods, namely chunking data collected from informants and two diagnostics specific to Chinese, i.e. bǎ and bèi sentence formation, were used. They were employed to discern which analysis should be preferred, i.e. the ternary-branching analysis, in which the intervening NP (NP2) is a dependent of the first predicate (P1), or the small-clause analysis, in which NP2 depends on the second predicate (P2). The results obtained suggest a flexible structural analysis for VDCs in the form of “NP1+P1-de+NP2+P2”. The difference in structural assignment is attributed to a semantic property of NP2 and the semantic relations it forms with adjacent predicates.
This paper investigates the seminal texts on Immediate Constituent Analysis and the associated diagrams. We show that the relations between the whole and its parts, that are typical of current phrase structure trees, were less prominent in the early di-agramming efforts than the relationships between units of the same level. This can be observed until the beginning of the 1960's, including in Chomsky's Syntactic Structures (1957). We discuss whether such analyses could be said dependency-based , according to an attempt to define this term.
Universal Dependency (UD) annotations, despite their usefulness for cross-lingual tasks and semantic applications, are not optimised for statistical parsing. In the paper, we ask what exactly causes the decrease in parsing accuracy when training a parser on UD-style annotations and whether the effect is similarly strong for all languages. We conduct a series of experiments where we systematically modify individual annotation decisions taken in the UD scheme and show that this re-sults in an increased accuracy for most, but not for all languages. We show that the encoding in the UD scheme, in particular the decision to encode content words as heads, causes an increase in dependency length for nearly all treebanks and an increase in arc direction entropy for many languages, and evaluate the effect this has on parsing accuracy.
This paper presents insights into non-projective relations in Serbian based on the analysis of an 81K token gold-standard corpus manually annotated for dependencies. We provide a formal profile of the non-projective dependencies found in the corpus, as well as a linguistic analysis of the underlying structures. We compare the observed properties of Serbian to those of other languages found in existing studies on non-projectivity.
Previous work on Korean language processing has proposed different basic segmentation units. This paper explores different possible dependency representations for Korean using different levels of segmentation granularity — that is, different schemes for morphological segmentation of tokens into syntactic words. We provide a new Universal Dependencies (UD)-like corpus based on different levels of segmentation granularity for Korean. The corpus contains 67K words in 5,000 sentences which are split into training, development and evaluation data sets. We report parsing results using the new dependency corpus for Korean and compare them with the previous Korean UD corpus.
The aim of this paper is to study some characteristics of dependency flux, that is the set of dependencies linking a word on the left with a word on the right in a given position. Based on an exploration of the whole set of UD treebanks (12M word corpus), we show that what we have called the flux weight, which measures center embeddings, is less than 3 in 99.62 % of the inter-word positions and is bounded by 6, which could be due to short-term memory limitations.
This contribution introduces a novel unit of syntactic analysis, which is called the component. The validity and utility of the component unit are established in terms of chunking. When informants organize the words of sentences into groups, they are creating chunks, and these chunks then qualify as components in dependency syntax. By acknowledging the nature of chunking and the component unit, it is possible to cast light on controversial aspects of dependency hierarchies. In particular, the component unit, informant data, and the reasoning based on these provide an argument in favor of the traditional DG assumptions about hierarchical status of many function words (auxiliary verbs, prepositions, subordinators, etc.), and in so doing, they contradict the Universal Dependencies (UD) annotation scheme. The data discussed here are from English, but the methodology and reasoning employed are easily extendable to other languages.
In this paper we present a cross-genre study on word order variation in Italian based on automatically dependency– parsed corpora. A comparative analysis focused on dependency direction and dependency distance for major constituents in the sentence is carried out in order to assess the influence of both textual genre and linguistic complexity on the distribution of phenonemena of syntactic markedeness.
Comunicacio presentada a: the Fourth International Conference on Dependency Linguistics (Depling 2017), celebrada a Pisa, Italia, del 18 al 20 de setembre 2017.
This paper describes an automatic procedure, the Semgrex-Plus tool, we developed to convert dependency treebanks into different formats. It allows for the definition of formal rules for rewriting dependencies and token tags as well as an algorithm for treebank rewriting able to avoid rule interference during the conversion process. This tool is publicly available1.
This paper presents our work to apply non linear neural network for parsing five resource poor Indian Languages belonging to two major language families - Indo-Aryan and Dravidian. Bengali and Marathi are Indo-Aryan languages whereas Kannada, Telugu and Malayalam belong to the Dravidian family. While little work has been done previously on Bengali and Telugu linear transition-based parsing, we present one of the first parsers for Marathi, Kannada and Malayalam. All the Indian languages are free word order and range from being moderate to very rich in morphology. Therefore in this work we propose the usage of linguistically motivated morphological features ( suffix and postposition ) in the non linear framework, to capture the intricacies of both the language families. We also capture chunk and gender, number, person information elegantly in this model. We put forward ways to represent these features cost effectively using monolingual distributed embeddings. Instead of relying on expensive morphological analyzers to extract the information, these embeddings are used effectively to increase parsing accuracies for resource poor languages. Our experiments provide a comparison between the two language families on the importance of varying morphological features. Part of speech taggers and chunkers for all languages are also built in the process.