The LiLa: Linking Latin project involves the creation of a Knowledge Base of linguistic resources for Latin based on the Linked Data framework. The ultimate goal is to reach full interoperability on the web between distributed lexical and textual resources. LiLa integrates all types of annotation applied to a particular word text into a common representation where all linguistic information contained in a linguistic resource becomes accessible. The LiLa Knowledge Base is thus a collection of resources represented with a shared vocabulary of (meta)linguistic knowledge description. The inclusion in the Knowledge Base of information on word formation, extracted from the Word Formation Latin lexical resource, raised a number of theoretical and practical issues concerning its treatment and representation. This paper discusses such issues, presents how they were addressed in the project with the help and implementation of a Word Paradigm theoretical model, and describes how the word formation data were included in the LiLa ontology.
Word Formation Latin (WFL; Litta et alii, 2016) is a linguistic resource providing explicitly recorded word formation relations between items in the Latin lexicon. The lexical basis of WFL comprises around 43,000 lemmas resulting from the collation of three Latin dictionaries, namely the Oxford Latin Dictionary (Glare, 1984), Georges & Georges (1913-1918) and Gradenwitz (1904). The WFL data are stored in a MySQL database, where input and output lexical items are connected via Word Formation Rules (WFRs). Each WFR provides information about (a) its input and output items, (b) their Part of Speech, (c) the rule type (derivation, compounding) and (c) the affix at work when predication or suffixation processes are concerned.WFL data were included into the database of the Latin morphological analyser Lemlat (Passarotti et alii, 2017), thus enhancing its inflectional morphological analysis (lemmatisation and morphological features) with information about word formation.WFL is freely available as part of the Lemlat database (https://github.com/CIRCSE/LEMLAT3) and accessible through a graphical web application (http://wfl.marginalia.it). Such application enables users to query WFL via WFR, affix, input/output PoS and lemma. While querying a specific lemma, users are provided with its full word formation cluster, i.e. a tree graph showing the full derivation of the lemma in question, consisting of its ancestor(s) as well as its descendant(s). In case the lemma is not morphologically derived, it is considered to be the ancestor of a "morphological family", i.e. the set of lemmas morphologically derived from a common ancestor. In the tree graphs of the WFL web application, nodes are lexical items and edges are WFRs.Beside this tree-like representation of word formation based lexical relations, we are experimenting also with a kind of visualisation where those lexical items that belong to the same morphological family are cells of a word formation paradigm instead of nodes of a derivation tree. The paradigm of a specific family is then connected with those of the other families in WFL and can be compared with these in terms of shared (and not shared) cells.Such an alternative view on WFL data follows the more recent approaches to derivational morphology based on Word & Paradigm models (Stekauer, 2014). We believe that a linguistic resource aiming to support studies in derivational morphology must provide access to data from both perspectives, thus enabling users to exploit the empirical lexical evidence made available in the resource either by following one single approach or by joining (and possibly comparing) the node-based and the paradigm-based one.Our contribution wants (1) to introduce WFL, by detailing both the theoretical and the practical aspects behind its building and (2) to raise a discussion with the workshop attendees about the two approaches to word formation, by showing examples of running queries on WFL in both kinds of visualisations available in the web application. Not only will this help us to refine the WFL application by understanding the needs coming from the community of the users, but it will provide the workshop attendees with enough expertise to use WFL in their research work about Latin derivational morphology.ReferencesGeorges Karl E. and Georges Heinrich. 1913-1918. Ausfuhrliches Lateinisch-Deutsches Handworterbuch. Hahn, Hannover. Glare, Peter G.W. 1982. Oxford Latin Dictionary. At the Clarendon Press, Oxford. Gradenwitz, Otto. 1904. Laterculi vocum latinarum. Hirzel, Leipzig. Litta, Eleonora, Marco Passarotti and Chris Culy. 2016. Formatio formosa est. Building a Word Formation Lexicon for Latin. In Proceedings of the Third Italian Conference on Computational Linguistics (CLiC–it 2016), aAccademia University Press. Passarotti, Marco, Marco Budassi, Eleonora Litta, and Paolo Ruffolo. 2017. “The Lemlat 3.0 Package for Morphological Analysis of Latin.” In Proceedings of the NoDaLiDa 2017 Workshop on Processing Historical Language, 24–31. Linkoping University Electronic Press. Stekauer, Pavol. 2014. Derivational Paradigms, in The Oxford Handbook of Derivational Morphology, Lieber, Rochelle and Stekauer, Pavol (eds.). Oxford University Press.
Despite more than a century of research, the origin of the Insular Celtic double system of verbal inflection is still debated. In this paper, we defend the thesis that the set of absolute endings originated by the agglutination of a subject clitic to the verb form. This clitic marked the declarative (vs. relative) use of verbs, since its distribution was complementary to that of the relative marker (star)yo. The present indicative as well as the preterite (in both the absolute and conjunct inflection) of one strong verb (berid 'bring') and one weak verb (lecid 'leave') are reconstructed according to this theory. For compound verb forms, the clitic similar to (star)yo alternation can be posited as well. The cases in which the distribution of initial mutations on the verb stem after preverbs does not follow the diachronic phonological rules of Old Irish (that is, there is no lenition after preverbs originally ending in a vowel) are accounted for from a synchronic standpoint. This "anomalous" behaviour can be explained by positing that a functionally relevant (morphological) system of mutations had replaced the previous phonology-based system.
English. This paper aims at examining the diachronic distribution of one of the richest classes of nouns in Latin, namely those ending in -io. The work is performed through the combined use of a morphological analyser for Latin (Lemlat), and a database collecting all word forms occurring through different periods of Latin language (TF-CILF). Italiano. Questo articolo presenta un’analisi della distribuzione diacronica di una delle più ricche classi di nomi in latino, ossia quelli che terminano in -io. Metodologicamente, il lavoro viene condotto attraverso l’uso incrociato di un analizzatore morfologico per il latino (Lemlat) e di una risorsa lessicale contenente tutte le forme di parole latine che occorrono in testi che vanno dall’antichità al neo-latino (TF-CILF).
The recent enhancement of the morphological analyser for Latin Lemlat with a large Onomasticon enables us to analyse both the morphology and the distribution of loanwords in the Latin lexicon. In this paper, first we describe the categories of proper names that were not possible to insert into Lemlat automatically, showing that a large part of them are loanwords. Then, we present the results of a qualitative analysis of loanwords to detect those 'exceptional' endings that identify loanwords featuring inflectional properties not assimilated to those regular in the morphological system of Latin. In the end, we report a quantitative analysis of data to study the frequency of such loanwords in Latin texts.
This paper introduces the main components of the downloadable package of the 3.0 version of the morphological analyser for Latin Lemlat. The processes of word form analysis and treatment of spelling variation performed by the tool are detailed, as well as the different output formats and the connection of the results with a recently built resource for derivational morphology of Latin. A light evaluation of the tool’s lexical coverage against a diachronic vocabulary of the entire Latin world is also provided.
We present a study on the degree of homonymy between the lexicon of a morphological analyser for Latin and an Onomasticon. To understand the impact of homonymy, we discuss an experiment on four Latin texts of different era and genre. L'articolo presenta uno studio sul grado di omonimia tra il lessico di un analizzatore morfologico per il latino e un Onomasticon. Al fine di comprendere l'impatto dell'omonimia, viene descritto un esperimento condotto su quattro testi latini di diversa epoca e genere.
Lemlat is a morphological analyser for Latin, which shows a remarkably wide coverage of the Latin lexicon.However, the performance of the tool is limited by the absence of proper names in its lexical basis.In this paper we present the extension of Lemlat with a large Onomasticon for Latin.First, we describe and motivate the automatic and manual procedures for including the proper names in Lemlat.Then, we compare the new version of Lemlat with the previous one, by evaluating their lexical coverage of four Latin texts of different era and genre.