A system of semantic classification is presented, casting a net of 25 broadly defined situational features over nearly 2000 verbs of Norwegian, with equivalents for English. The feature specifications are formally connected to valence representations in the valence catalogue NorVal, to facilitate investigations into the relation between verb meaning and verb valence.
The present paper investigates multiple Correlative Constructions (CC) in Odia and sketches a combined semantic and syntactic analysis. The paper describes Correlative Constructions and related constructions in Odia, with a view especially to its quantificational systems, one residing in lexical quantifiers, and one in the clause combinations which constitute CCs. Over the last decades, a growing literature has addressed similarities between CCs as instantiated in languages on the Indian subcontinent and types of Free Relatives, e.g., in English, as they occur in positions adjoined to clauses, here to be called Adjoined Free Relatives (AFRs). AFR constructions supplement lexical quantification in English in a parallel way to CCs in Odia, and we explore possibilities of representing CCs and AFR constructions within a common semantico-syntactic frame of analysis. We show how the quantificational effects of CCs can be derived from their character as relative constructions, residing in what we call co-targeted predicates, as opposed to lexical encoding of quantificational meaning through items such as ‘each’, ‘every’ and the like. We thereby describe two distinct strategies for obtaining partially similar quantificational effects, a finding which applies to CC/AFR constructions cross-linguistically.
One characteristic of so-called Light Verb Constructions (LVCs) is that they unfold, mostly over a sequence ‘Subject V (P) N’, a content that could in principle be carried by some verb V alone, where the N of the sequence expresses a situational content close to that of V (N thus being some kind of ‘deverbal’ variant of V). A typical role of (the ‘light’ verb) V in the LVC is to connect the Subject to this situational content as some kind of role bearer, and add aspectual content to the situational content expressed by N. In languages where LVCs are seen as constituting a major category, the number of verbs serving as ‘Light Verb’ is low, and LVCs constitute formally recurrent patterns. In Norwegian the picture is different, with many verbs serving as possible Light Verbs, where these verbs select different nouns, and the nouns in turn select different verbs. In illustrating this, we outline a format for representation of LVCs in corpora. We in addition outline a format for representing LVCs in a valence lexicon, and for representing them in a sentence parsing formalism. The unification between the meaning of the Light Verb and the meaning of the noun is formally represented. While a sizable number of LVC selection relations between verbs and nouns have been identified, and a large valence lexicon and parser constitute the frames for the formalizations mentioned, the formalizations have only to very small extents been implemented, being presented here only for formal consideration.
A 'Valence Catalogue' is a design for succinctly representing information about the valence frames of a language. It can be coordinated with in-depth annotated corpora and computational parsers. Such a cluster is realized by the verb Valence Catalogue NorVal of Norwegian Hellan (Natural Language Processing in Artificial Intelligence—NLPinAI 2021, Springer, 2022) together with a computational parser for Norwegian, NorSource, and the multi-lingual grammar system TypeGram. Another such cluster is realized by a verb Valence Catalogue of the West African language Ga (GaVal) together with the language specific parser GaGram and with coverage also by TypeGram. The resources are all based on the general formalism of Typed Feature Structures, whereby grammars and valence catalogues are formally tied together, and specifications in the resources are succinctly comparable also cross-linguistically. In describing the multilingual aspects of these architectures, we thereby demonstrate that systems of this formal nature are fully attainable no matter where on the ladder of 'high-low-resourced' a language finds itself. In a comparative perspective, we also demonstrate discrepancies between the valence systems of the two languages, made visible through the succinctness and coverage of Valence Catalogues, and perspectives they open for typological research and for cross-linguistically informed resource generation.
Essential aspects of a verb’s usage reside in its valence environments. The Norwegian valence resource here presented, called NorVal, has 6,300 verb lemmas. About 3,360 of them are associated with sets of frames, and the organization of entries is divided into one enumeration of the total number of frame-specific entries, which is about 15,750, and one enumeration of lemmas, counting 6,300. About 300 frame types are distinguished inducing the 15,750 frame specific entries, taking into account most grammatical factors distinguishing verb frames and verb-headed construction types. Both the frame types and the two dimensions of entries are represented in string-based formalisms, enabling simple procedures for comparing individual valence frames, frame-specific entries, and entries representing lemmas, and for doing statistics over types and combinations of all of these. The paper illustrates the resources relative to their representation of light reflexives, verb particles, and frames including sentential constituents.
The paper presents an annotation schema with the following characteristics: it is formally compact; it systematically and compositionally expands into fullfledged analytic representations, exploiting simple algorithms of typed feature structures; its representation of various dimensions of semantic content is systematically integrated with morpho-syntactic and lexical representation; it is integrated with a ‘deep’ parsing grammar. Its compactness allows for efficient handling of large amounts of structures and data, and it is interoperable in covering multiple aspects of grammar and meaning. The code and its analytic expansions represent a cross-linguistically wide range of phenomena of languages and language structures. This paper presents its syntactic-semantic interoperability first from a theoretical point of view and then as applied in linguistic description.
Traditionally, a lexicographer identifies the lexical items to be added to a dictionary. Here we present a corpus-based approach to dictionary compilation and describe a procedure that derives a Twi dictionary from a TypeCraft corpus of Interlinear Glossed Texts. We first extracted a list of unique words. We excluded words belonging to different dialects of Akan (mostly Fante and Abron). We corrected misspellings and distinguished English loan words to be integrated in our dictionary from instances of code switching. Next to the dictionary itself, one other resource arising from our work is a lexicographical model for Akan which represents the lexical resource itself, and the extended morphological and word class inventories that provide information to be aggregated. We also represent external resources such as the corpus that serves as the source and word level audio files. The Twi dictionary consists at present of 1367 words; it will be available online and from an open mobile app.
The paper describes aspects of an HPSG style computational grammar of the West African language Ga (a Kwa language spoken in the Accra area of Ghana). As a Volta Basin Kwa language, Ga features many types of multiverb expressions and other particular constructional patterns in the verbal and nominal domain. The paper highlights theoretical and formal features of the grammar motivated by these phenomena, some of them possibly innovative to the formal framework. As a so-called deep grammar of the language, it hosts a rich lexical structure, and we describe ways in which the grammar builds on previously available lexical resources. We outline an environment of current resources in which the grammar is part, and lines of research and development in which it and its environment can be used.
Abstract This paper investigates constructions in Norwegian and German with an expletive pronoun in subject position, and for Norwegian also in object position. The discussion covers presentational, impersonal and extrapositional constructions in both languages, and in Norwegian also the ‘light reflexive’ seg in its interaction with presentationals. We relate the discussion to a parameter of theticity, whereby sentences with an expletive subject will count as thetic while sentences with a content-full NP subject will count as categorical. Also sentences with expletive object are argued to have a thetic value. Categorical sentences on their side are ranked according to a parameter of transitivity, accounting for constraints on presentational constructions in Norwegian, and seen as constituting an opposite dimension of constructional values to that of theticity.
We present a procedure for generating a valence resources for Norwegian (Bokmål) from a deep grammar. The corpus is presented in the form of IGT (interlinear glossed text) augmented by valence information. Our deep parser is the HPSG-based computational grammar Norsource (Hellan and Bruland 2015), our online IGT repository is TypeCraft (Beermann and Mihaylov 2014), while the sentences of the corpus are taken from the Leipzig Corpus Collection (Goldhahn et al. 2012). We create a common structure for the resources. Our aim is to make the grammatical information encoded in a deep parser more readily accessible for humans and for further processing.
We describe a methodology by which verb valence information can be derived from corpora by using subcorpora of typical sentences that are constructed in a language independent manner based on frequent POS structures. The inspection of typical sentences with a fixed verb in a certain position can show the valence information directly. Using verb fingerprints, consisting of the most typical sentence patterns the verb appears in, we are able to identify standard valence patterns and compare them against a language's valence profile. With a very limited number of training data per language, valence information for other verbs can be derived as well. Based on the Norwegian valence patterns, we are able to find comparable patterns in German where typical sentences are able to express the same situation in an equivalent way, and can so allow for the construction of verb valence pairs for a bilingual dictionary. This contribution discusses this application with a focus on the Norwegian valence dictionary NorVal.
West Africa as an area of linguistic diversity and unification processesWest Africa as a sub-region of the African continent is defi ned mostly on geographic and political criteria which exclude Northern Africa and the Maghreb covering the Sub-Saharan countries from Senegal to Nigeria.As a region of linguistic studies, West Africa adheres to these limits, though genetic relationship and historical contacts between languages make these conventional boundaries vague in a number of respects.The region is characterized by linguistic diversity which determines the prominence of research oriented at multilingualism and language contact.The works conducted so far has focused on identifying convergence zones rather than providing the proof of the linguistic coherence in the entire region.The term convergence zone refers to a region where many linguistic features are shared across the language boundaries.The two largest units, i.e.Macro-Sudan Belt extending from Senegal to Ethiopia (Güldemann 2008) and Wider Lake Chad Region overlap to some extent, especially in Nigeria where genetically distinct and structurally diff erent languages meet (Ziegelmeyer 2015;Wolff & Löhr 2005;Zima 2009; Cyff er & Ziegelmeyer 2009).West Africa is characterized by extensive societal multilingualism (Lüpke & Chambers 2010).Along with indigenous languages superimposed foreign languages such as French, English or Portuguese are used.The region has always been an area where languages brought by scholars, traders or travelers were in constant confrontation with those used locally.Muslim teachers and traders moving along West Africa brought Arabic to this region.The emergence of political centers such as Ghana in 12 th century, Mali in 14 th century, Songhay in 15 th century or the Sokoto Caliphate in 19 th century strengthened the position of Arabic as the language of courts, written correspondence, religious and legal teaching.Arabic also became an important contact language among educated and infl uential people living in towns and it had an impact on the major languages of the Sahel and northern savannah.The historical empires and city-states also promoted languages spoken by the ruling class such as Hausa, Fulfulde or Mande languages, pushing many other local languages aside.Due to globalization, urbanization and economic development, the number of languages
The paper presents a system for construction classification representing multiple levels of specification, such as grammatical functions, grammatically reflected actants, and lexical semantics, aligned with a compositional system of sign combination mediating between a construction perspective and a valence perspective. The system uses a feature structure formalism based on Head-Driven Phrase Structure Grammar (HPSG) but with essential elements from Lexical Functional Grammar (LFG; cf. Bresnan in Lexical functional syntax. Blackwell, Oxford, 2001), and has as implementation background large scale HPSG grammars. While on the one extreme being able to encode word level selection in multi-word patterns, the system on the other provides a compact format for construction specification, allowing for cross-language comparison both in construction and valence frame inventories. Pivotal in these capacities as well as in sign formalization in general are the grammatical functions. The paper motivates the usefulness of the various functionalities and illustrates the way in which they work together in a formally uniform system.
Our project draws on linguistic resources for German and Norwegian in support of initiatives that try to make public language more accessible. We focus on "Leichte Sprache" for German and "Klart Språk" for Norwegian. The former refers to usage forms of German which are easily understood also by users of German with a lower competence in text processing, the latter refers to a governance project which has as its aim to aid institutions in their effort to communicate with the public. While different in their goals both initiatives seek to increase the general access to public information. A central concern is to identify the factors which affect language complexity and set up linguistic resources such as parallel text corpora, linguistic rules and terminology databases. Based on these results and resources, we are building software that supports authors and translators, in close collaboration with them.
We describe existing resources of the Kwa languages Akan and Ga, with a view to transfer of resources well developed for one to the other. While we can build on an Interlinear Glossed Text (IGT) corpus for Akan we have a modern digital lexicon for Ga, something we still lack for Akan, while we have only very limited IGT data for Ga. While it is normally the case that annotations from a resource rich language are transferred to a resource poor language, we are here preparing our resources to allow for a transfer approach between two resource-low but closely related languages. We envisage this to be a viable strategy also for other pairs of closely related under resourced languages.
Daniel Flickinger合作论文数School of Humanities and Sciences, Stanford University3