The Indo-European Cognate Relationships (IE-CoR) dataset is an open-access relational dataset showing how related, inherited words (‘cognates’) pattern across 160 languages of the Indo-European family. IE-CoR is intended as a benchmark dataset for computational research into the evolution of the Indo-European languages. It is structured around 170 reference meanings in core lexicon, and contains 25731 lexeme entries, analysed into 4981 cognate sets. Novel, dedicated structures are used to code all known cases of horizontal transfer. All 13 main documented clades of Indo-European, and their main subclades, are well represented. Time calibration data for each language are also included, as are relevant geographical and social metadata. Data collection was performed by an expert consortium of 89 linguists drawing on 355 cited sources. The dataset is extendable to further languages and meanings and follows the Cross-Linguistic Data Format (CLDF) protocols for linguistic data. It is designed to be interoperable with other cross-linguistic datasets and catalogues, and provides a reference framework for similar initiatives for other language families.
Exploring complex lexemes cross-linguisticallyEditors' introduction 1 Towards a typology of complex lexemes Concept-naming is one of the most fundamental activities performed by speakers, who need either ready-made labels to talk about entities or devices to build new labels (be they rules or processes, schemas or analogical mechanisms).Knowing how languages perform the basic function of creating labels to name concepts, especially complex concepts, is crucial to understanding their creative potential in building new (potentially stable) categories and, more generally, to understanding how they (may) categorize reality, and refer to it.What are the strategies employed by languages for naming complex concepts?How do they differ cross-linguistically, and what are the limits of their variation?Are there strategies that are more widespread than others, or even universal? 1 These are questions for lexical typology and/or word-formation typology, but what we know about the typology of complex concept naming is very limited compared to what we know about domains like word order or inflectional morphology.There may be different reasons behind this state-of-affairs.Analysing all of them falls outside the scope of the present introduction: we will just discuss some factors that we deem relevant for our current purposes.Complex concept naming is definitely related to word-formation.The domain of word-formation can count on an extremely rich and ever-growing body of literature, which would be impossible to credit here (suffice it to mention the collections edited by Booij, Lehmann & Mugdan 2000;2004; Lieber & Štekauer 2009a;2014; Müller et al. 2015; Lieber et al. 2021).However, quite surprisingly, word-formation has rarely been the subject of large-scale, thorough typological 1 We wish to thank the many anonymous referees who generously agreed to review the chapters included in this volume, including this introduction: their insightful comments significantly improved the quality of the volume.Heartfelt thanks are also due to Jean-Christophe Verstraete for his guidance and constant support, which were essential to bring this project to conclusion.We are also grateful to the audience of the When "noun" meets "noun" workshop at the
This chapter starts by demonstrating the need for the comparative concept 'binominal lexeme' in order to cover both 'noun-noun compounds' and their 'functional equivalents' (§1). To complement this informal definition, four different, but compatible definitions of binominal lexeme are developed: functional, onomasiological, formal and typological (§2). Although couched in a variety of terms based on different theoretical frameworks, these have essentially identical extensions. In §3 a nine-way classification of binominal strategies is presented, together with the mnemonics used throughout this volume: jxt, cmp, der, cls; prp, gen, adj, con, and dbl. These nine types are represented on a two-dimensional grid that captures the number of markers, the locus of marking and the degree of fusion. The grid reveals two lacunae or "missing types": prn and nml. Whereas the first of these probably exists somewhere in the world's languages, the second seems to be a logical impossibility. §4 discusses types that are intermediate between the nine main types and the grammaticalization pathways that produce them. It then goes on to examine the relationship between binominal constructions and adnominal possessives, and introduces a new methodology, based on the Pwav scale, for comparing two non-binary constructions. This leads to the formulation of two Greenbergian universals concerning binominals and nominal modification.
A key feature of binominal lexemes is the unstated (or underspecified) relation, ℜ, that pertains between the two major constituents. The nature of ℜ - the kinds of relations - has been the topic of considerable research during recent decades. While early studies focused almost exclusively on English, the last few years have seen a spate of work on other languages. Unfortunately, this work has been uncoordinated and each researcher entering the field has tended to devise their own classification, making it difficult to compare results and advance our understanding of the phenomenon. This is a pity, because such an understanding has the potential to provide insights into the nature of concept combination and the associative character of human thought. The purpose of this chapter is to present a well-documented, systematic classification of semantic relations that operates at multiple levels of granularity and is suitable for reuse across languages. Hatcher- Bourque is based on revisions of two earlier classifications, those of Anna Granville Hatcher and Yves Bourque, which operate at different levels of granularity. These are integrated into a single, coherent system, with automatic mapping from one level to the other. The classification is applied to a set of 3,650 binominals from 106 languages, and an analysis is presented of the frequency and distribution of semantic relations at both a highly abstract level and a more granular level. The Hatcher-Bourque classification, and an accompanying, Excel-based tool, the Bourquifier, are offered to the research community in order to encourage collaboration, and researchers are invited to participate in the Hatcher-Bourque Cake Challenge.
There have been many attempts at classifying the semantic modification relations (R) of N + N compounds but this work has not led to the acceptance of a definitive scheme, so that devising a reusable classification is a worthwhile aim. The scope of this undertaking is extended to other binominal lexemes, i.e. units that contain two thing-morphemes without explicitly stating R, like prepositional units, N + relational adjective units, etc. The 25-relation taxonomy of Bourque (2014) was tested against over 15,000 binominal lexemes from 106 languages and extended to a 29-relation scheme ("Bourque2") through the introduction of two new reversible relations. Bourque2 is then mapped onto Hatcher's (1960) four-relation scheme (extended by the addition of a fifth relation, similarity, as "Hatcher2"). This results in a two-tier system usable at different degrees of granularities. On account of its semantic proximity to compounding, metonymy is then taken into account, following Janda's (2011) suggestion that it plays a role in word formation; Peirsman and Geeraerts' (2006) inventory of 23 metonymic patterns is mapped onto Bourque2, confirming the identity of metonymic and binominal modification relations. Finally, Blank's (2003) and Koch's ( 2001) work on lexical semantics justifies the addition to the scheme of a third, superordinate level which comprises the three Aristotelean principles of similarity, contiguity and contrast.
Advances in computer-assisted linguistic research have been greatly influential in reshaping linguistic research. With the increasing availability of interconnected datasets created and curated by researchers, more and more interwoven questions can now be investigated. Such advances, however, are bringing high requirements in terms of rigorousness for preparing and curating datasets. Here we present CLICS, a Database of Cross-Linguistic Colexifications (CLICS). CLICS tackles interconnected interdisciplinary research questions about the colexification of words across semantic categories in the world's languages, and show-cases best practices for preparing data for cross-linguistic research. This is done by addressing shortcomings of an earlier version of the database, CLICS2, and by supplying an updated version with CLICS3, which massively increases the size and scope of the project. We provide tools and guidelines for this purpose and discuss insights resulting from organizing student tasks for database updates.
Abstract There have been many attempts at classifying the semantic modification relations (ℜ) of N + N compounds but this work has not led to the acceptance of a definitive scheme, so that devising a reusable classification is a worthwhile aim. The scope of this undertaking is extended to other binominal lexemes, i.e. units that contain two thing-morphemes without explicitly stating ℜ, like prepositional units, N + relational adjective units, etc. The 25-relation taxonomy of Bourque (2014) was tested against over 15,000 binominal lexemes from 106 languages and extended to a 29-relation scheme (“Bourque2”) through the introduction of two new reversible relations. Bourque2 is then mapped onto Hatcher’s (1960) four-relation scheme (extended by the addition of a fifth relation, similarity, as “Hatcher2”). This results in a two-tier system usable at different degrees of granularities. On account of its semantic proximity to compounding, metonymy is then taken into account, following Janda’s (2011) suggestion that it plays a role in word formation; Peirsman and Geeraerts’ (2006) inventory of 23 metonymic patterns is mapped onto Bourque2, confirming the identity of metonymic and binominal modification relations. Finally, Blank’s (2003) and Koch’s (2001) work on lexical semantics justifies the addition to the scheme of a third, superordinate level which comprises the three Aristotelean principles of similarity, contiguity and contrast.
There have been many attempts at classifying the semantic modification relations () of N + N compounds but this work has not led to the acceptance of a definitive scheme, so that devising a reusable classification is a worthwhile aim. The scope of this undertaking is extended to other binominal lexemes, i.e. units that contain two thing-morphemes without explicitly stating , like prepositional units, N + relational adjective units, etc. The 25relation taxonomy of Bourque (2014) was tested against over 15,000 binominal lexemes from 106 languages and extended to a 29-relation scheme (“Bourque2”) through the introduction of two new reversible relations. Bourque2 is then mapped onto Hatcher’s (1960) four-relation scheme (extended by the addition of a fifth relation, SIMILARITY, as “Hatcher2”). This results in a two-tier system usable at different degrees of granularities. On account of its semantic proximity to compounding, metonymy is then taken into account, following Janda’s (2011) suggestion that it plays a role in word formation; Peirsman & Geeraerts’ (2006) inventory of 23 metonymic patterns is mapped onto Bourque2, confirming the identity of metonymic and binominal modification relations. Finally, Blank’s (2003) and Koch’s (2001) work on lexical semantics justifies the addition to the scheme of a third, superordinate level which comprises the three Aristotelean principles of similarity, contiguity and contrast.
I discuss the history of research on this topic and various approaches that have been taken to understand the nature of . I then present a selection of classification schemes, focusing on those of Hatcher (1960) and Bourque (2014), which operate at different levels of granularity. Following Arnaud (2016), I show how these two systems can be mapped together into a twotiered system (the “Hatcher-Bourque classification”) and describe the advantages of so doing.
The nature of has been the subject of considerable research, often with each new researcher reinventing the classificatory wheel (see Hacken 2016 for a recent summary). We focus on two classification schemes of the “reductionist” type (Søgaard 2005) which operate at different levels of granularity: Hatcher’s (1960) system of four logical relations and Bourque’s (2014) 25-way empirically-derived classification. Following Arnaud (2016), we show how these two systems can be mapped together into a two-tiered system (the “HatcherBourque classification”). We argue that this resolves the dispute regarding the number of relations involved. That number depends on the requirements of the analysis, and the degree of granularity can range from one (as suggested by Bauer 1979) to unlimited (as opined by Jespersen 1942). Our resulting two-tiered system has been tested against a database of over 3,700 noun-noun compounds and their functional equivalents from 106 languages.
A key feature of binominal lexemes is the unstated (or underspecified) relation, R, that pertains between the two major constituents. Understanding the nature of R – the kinds of relations and their frequency – has the potential to provide insights into the nature of concept combination and the associative character of human thought. In this chapter I approach the question from a cross-linguistic, onomasiological perspective. I develop a two-tier classification called the Hatcher-Bourque system based on work by earlier researchers. The classification operates at two levels of granularity, with 29 and five relations respectively. The classification is applied to a data set consisting of over 3,700 binominal lexemes from 106 languages. This permits a statistical analysis of the frequency and distribution of semantic relations at both an highly abstract level and a more granular level. The Hatcher-Bourque classification, and an accompanying tool called the Bourquifier, are offered to the research community in order to encourage more collaboration than has hitherto been the case in the field of semantic relations.
This paper reports our contribution to the 2013 NLI Shared Task. The purpose of the task was to train a machine-learning system to identify the native-language affiliations of 1,100 texts written in English by nonnative speakers as part of a high-stakes test of general academic English proficiency. We trained our system on the new TOEFL11 corpus, which includes 11,000 essays written by nonnative speakers from 11 native-language backgrounds. Our final system used an SVM classifier with over 400,000 unique features consisting of lexical and POS n-grams occurring in at least two texts in the training set. Our system identified the correct nativelanguage affiliations of 83.6% of the texts in the test set. This was the highest classification accuracy achieved in the 2013 NLI Shared Task.
The Dublin Core Metadata Initiative is an open organization engaged in the development of interoperable online metadata standards that support a broad range of purposes and business models. Its most important standard is the Dublin Core Metadata Element Set [1], a vocabulary of fifteen proper ties for use in resource description, which was approved as ISO 15836:2003 [2]. This and other vocabularies developed by Dublin Core are defined as abstract models which may be expressed in any number of different syntaxes. This paper presents a proposal for expressing such metadata using Topic Maps.
This paper describes the need for a simple mechanism for defining and assigning unique global identifiers for arbitrary subjects on the World Wide Web in order to solve the problem of information overload. It presents the case for Published Subjects and published subject indicators (PSIs) being the best solution to this problem, and briefly characterizes the strengths and weaknesses of alternative approaches. It ends with a call to action. It might look like a scientific paper, but it is not. It does not represent scholarly work that is being published for the first time, and it ought to be understandable by anyone into whose hands it is likely to fall. Nor is it a standards document (although parts of it may read like one) because the ideas and proposals it contains are so simple and obvious as to hardly seem worth standardizing. Rather, it is a call for action, aimed at absolutely anyone who aids and/or abets in the publication of information, or dissemination of knowledge, on the World Wide Web, especially those concerned with semantic interoperability. (Yes, that does indeed mean you.)
The Semantic Web relies on the presence of semantic annotations which describe information in a machine readable form. There are two standard formalisms which are suitable for this aim: RDF and Topic Maps. This paper presents an analysis of the issues that need to be addressed in order to define a set of rules for performing automated and consistent translations between RDF and Topic Maps, i.e. semantic mapping issues. The analysis is based on existing approaches to the RDF and Topic Maps translation in the light of the new official formal models released for both RDF and Topic Maps.
This document provides an introduction to Published Subjects and basic requirements and recommendations for publishers of Published Subjects.
Lars Marius Garshol合作论文数ISO2