Non-standardised early vernaculars present a problem for search tools due to the high degree of variation. The challenge lies in the variation found in orthography, syntax, and lexicon between titles, incipits, and explicits in manuscript copies of the same work. Traditional search methods relying on exact string matching or regular expressions fail to address these variations comprehensively. This project presents a web-based search tool specifically designed to handle linguistic and textual variation. The software is made available as a part of the Index of Middle English Prose (IMEP). The search tool addresses the issue of variation by utilizing a database of incipits and explicits, character-based n-gram language models (LMs) built with the Stanford Research Institute Language Modelling (SRILM) toolkit, and a fuzzy search script (IMEP: FSS) written in Python. The tool optimizes for recall, retrieving multiple potential matches for a search string, without attempting to identify the ‘correct’ one. The search process involves looking up exact matches in the database while simultaneously using the fuzzy search script to evaluate the incipits and explicits against a model of the search string, followed by a match of the search string against models of the incipits and explicits. This two-step process shortens the processing time, which would otherwise be unreasonably long, because while using SRILM to match the search string against each incipit or explicit in the IMEP for precision could be time-consuming, running a first step where all texts are matched against a single LM built from the search string allows for faster processing. A web application, built using Django and Docker, combines the results of the direct database lookup and the fuzzy search script, presenting them as a list with exact matches followed by fuzzy matches ordered by increasing model perplexity. The tool is made available Open Access and can be adapted to other datasets.
Non-standardised early vernaculars present a problem for search tools due to the high degree of variation. The challenge lies in the variation found in orthography, syntax, and lexicon between titles, incipits, and explicits in manuscript copies of the same work. Traditional search methods relying on exact string matching or regular expressions fail to address these variations comprehensively. This project presents a web-based search tool specifically designed to handle linguistic and textual variation. The software is made available as a part of the Index of Middle English Prose (IMEP). The search tool addresses the issue of variation by utilizing a database of incipits and explicits, character-based n-gram language models (LMs) built with the Stanford Research Institute Language Modelling (SRILM) toolkit, and a fuzzy search script (IMEP: FSS) written in Python. The tool optimizes for recall, retrieving multiple potential matches for a search string, without attempting to identify the 'correct' one. The search process involves looking up exact matches in the database while simultaneously using the fuzzy search script to evaluate the incipits and explicits against a model of the search string, followed by a match of the search string against models of the incipits and explicits. This two-step process shortens the processing time, which would otherwise be unreasonably long, because while using SRILM to match the search string against each incipit or explicit in the IMEP for precision could be time-consuming, running a first step where all texts are matched against a single LM built from the search string allows for faster processing. A web application, built using Django and Docker, combines the results of the direct database lookup and the fuzzy search script, presenting them as a list with exact matches followed by fuzzy matches ordered by increasing model perplexity. The tool is made available Open Access and can be adapted to other datasets.
In this paper, we report on efforts to improve the Oslo-Bergen Tagger for Norwegian morphological tagging by using a hybrid system that combines the output of the rule-based Constraint Grammar tagger with a neural sequence-to-sequence model trained for tagging. The results are very promising for cases where the two systems intersect in tokenisation and morphological analysis, but problems remain in integrating the two systems in many cases.
This paper presents the NDC Treebank of spoken Norwegian dialects in the Bokm degrees al variety of Norwegian. It consists of dialect recordings made between 2006 and 2012 which have been digitised, segmented, transcribed and subsequently annotated with morphological and syntactic analysis. The nature of the spoken data gives rise to various challenges both in segmentation and annotation. We follow earlier efforts for Norwegian, in particular the LIA Treebank of spoken dialects transcribed in the Nynorsk variety of Norwegian, in the annotation principles to ensure interusability of the resources. We have developed a spoken language parser on the basis of the annotated material and report on its accuracy both on a test set across the dialects and by holding out single dialects.
We present the Norwegian Anaphora Resolution Corpus (NARC), the first publicly available corpus annotated with anaphoric relations between noun phrases for Norwegian. The paper describes the annotated data for 326 documents in Norwegian Bokmål, together with inter-annotator agreement and discussions of relevant statistics. We also present preliminary modelling results which are comparable to existing corpora for other languages, and discuss relevant problems in relation to both modelling and the annotations themselves.
Language documentation, including the development and use of corpora, is frequently linked to revitalisation. This is also the case for the Kven language, a Finnic minoritised language, traditionally spoken in the two northernmost counties of Norway. Kven is a recognised minority language in Norway, protected by the European Charter for Regional or Minority Languages. This status led to increased efforts to document Kven, including the development of the Ruija Corpus, consisting of recordings of interviews in Kven. The corpus was an important tool for the standardisation of Kven. In this article we describe how the corpus was developed and account for search functions, including a discussion of the limitations of the corpus. We also discuss the role of corpora and other online tools for language revitalisation, with a particular focus on the standardisation of Kven and conclude by reflecting on how expertise also resides with the speakers of an endangered language and that they have a right to be involved in efforts of language documentation and revitalisation.
In this article we show how the search interface Glossa has been developed in step with the various corpora that have been built at the Text Laboratory. Furthermore, we present statistics on what kind of searches people do – single words or longer phrases, with or without specifications for phonetic form or grammatical features etc. – focusing on the Nordic Dialect Corpus and the Corpus of American Nordic Speech. Finally, we demonstrate how researchers have searched for data in these corpora and used them in published articles – both simple and extended search, in smaller or larger language areas – within several different branches of linguistics.
Denne artikkelen rapporterer om ein studie av geografiske og demografiske trekk ved 46 etterstilte uttrykk i norske talemål, mellom anna gitt, sant og kan du skjønne. I fyrste del av studien er spørjeskjema nytta som metode. Resultata frå denne undersøkinga viser i kva grad informantar frå ulike stader i Noreg rapporterer om bruk av dei etterstilte uttrykka. Andre del av studien er ei undersøking av førekomstar av etterstilte uttrykk i korpusa Nordisk dialektkorpus og LIA norsk. Samla viser studien at i) mange av dei 46 etterstilte uttrykka er avgrensa geografisk, ii) ein del av uttrykka blir nytta i større grad av dei eldre informantane enn av dei yngre, og omvendt, iii) nokre få uttrykk har ein meir frekvent bruk hos menn enn hos kvinner, og omvendt, og iv) yngre språkbrukarar ser ut til å plassere uttrykk i den etterstilte posisjonen meir hyppig enn eldre når fleire posisjonar er mogleg. I tillegg gjev undersøkinga auka innsikt om dei to metodane som vart nytta.
This article reports on a study of geographical and demographical aspects of 46 final particles in Norwegian dialects, among them gitt (< lit. ‘boy’), sant (< lit. ‘true’) and kan du skjonne (< lit. ‘can you realize’). In the first part of the study, a questionnaire was used as the method. The results from this study show to what extent informants from different regions in Norway report on use of the given final particles. The second part of the study is an investigation of the use of the 46 final particles in the corpora Nordisk dialektkorpus and LIA norsk. In sum, the study shows that i) many of the 46 final particles are delimited geographically; ii) some of the expressions are used to greater extent by the older informants than by the younger ones, and the other way around; iii) a few expressions have a more frequent use among men than among women, and the other way around, and iv) overall, younger people seem to place expressions in the final (tag) position more often than do older people when more than one syntactic position is possible. In addition, the study contributes insights into the two methods that were used.
The present article presents four experiments with two different methods for measuring dialect similarity in Norwegian: the Levenshtein method and the neural long short term memory (LSTM) autoencoder network, a machine learning algorithm. The visual output in the form of dialect maps is then compared with canonical maps found in the dialect literature. All of this enables us to say that one does not need fine-grained transcriptions of speech to replicate classical classification patterns.
Denne artikkelen er en introduksjon til Leksikografisk bokmålskorpus (LBK). Vi starter med en historisk oversikt over ordboksarbeid som er utført for norsk språk, og forklarer bakgrunnen for at LBK ble bygd opp på den måten det ble. Deretter gir vi en oversikt over innholdet i korpuset, før vi til slutt viser hvordan man kan søke i korpuset ved hjelp av korpussøkeverktøyet Glossa.
In this article, we present the Nordic Word Order Database (NWD), with a focus on the rationale behind it, the methods used in data elicitation, data analysis and the empirical scope of the database. NWD is an online database with a user-friendly search interface, hosted by The Text Laboratory at the University of Oslo, launched in April 2019 (https://tekstlab.uio.no/nwd). It contains elicited production data from speakers of all of the North Germanic languages, including several different dialects. So far, 7 fieldtrips have been conducted, and data from altogether around 250 participants (age 16–60) have been collected (approx. 55 000 sentences in total). The data elicitation is carried out through a carefully controlled production experiment that targets core syntactic phenomena that are known to show variation within and/or between the North Germanic languages, e.g., subject placement, object placement, particle placement and verb placement. In this article, we present the motivations and research questions behind the database, as well as a description of the experiment, the data collection procedure, and the structure of the database
This paper describes an evaluation of five data-driven part-of-speech (PoS) taggers for spoken Norwegian. The taggers all rely on different machine learning mechanisms: decision trees, hidden Markov models (HMMs), conditional random fields (CRFs), long-short term memory networks (LSTMs), and convolutional neural networks (CNNs). We go into some of the challenges posed by the task of tagging spoken, as opposed to written, language, and in particular a wide range of dialects as is found in the recordings of the LIA (Language Infrastructure made Accessible) project. The results show that the taggers based on either conditional random fields or neural networks perform much better than the rest, with the LSTM tagger getting the highest score.
This article presents the LIA treebank of transcribed spoken Norwegian dialects. It consists of dialect recordings made in the period between 1950-1990, which have been digitised, transcribed, and subsequently annotated with morphological and dependency-style syntactic analysis as part of the LIA (Language Infrastructure made Accessible) project at the University of Oslo. In this article, we describe the LIA material of dialect recordings and its transcription, transliteration and further morphosyntactic annotation. We focus in particular on the extension of the native NDT annotation scheme to spoken language phenomena, such as pauses and various types of disfluencies, and present the subsequent conversion of the treebank to the Universal Dependencies scheme. The treebank currently consists of 13,608 tokens, distributed over 1396 segments taken from three different dialects of spoken Norwegian. The LIA treebank annotation is an on-going effort and future releases will extend on the current data set.
This paper presents and describes a modernised version of Glossa, a corpus search and results visualisation system with a user-friendly interface. The system is open source and can be easily installed on servers or even laptops for use with suitably prepared corpora. It handles parallel corpora as well as monolingual written and spoken corpora. For spoken corpora, the search results can be linked to audio/video, and spectrographic analysis and visualised geographical distributions can be provided. We will demonstrate the range of search options and result visualisations that Glossa provides.
An untested assumption behind the crowdsourced descriptions of the images in the Flickr30k dataset (Young et al., 2014) is that they “focus only on the information that can be obtained from the image alone” (Hodosh et al., 2013, p. 859). This paper presents some evidence against this assumption, and provides a list of biases and unwarranted inferences that can be found in the Flickr30k dataset. Finally, it considers methods to find examples of these, and discusses how we should deal with stereotype-driven descriptions in future applications.
In this paper we discuss and evaluate machine learning-based optimization of a Constraint Grammar for Norwegian Bokmal (OBT). The original linguistwritten rules are reiteratively re-ordered, re-sectioned and systematically modified based on their performance on a handannotated training corpus. We discuss the interplay of various parameters and propose a new method, continuous sectionizing. For the best evaluated parameter constellation, part-of-speech F-score improvement was 0.31 percentage points for the first pass in a 5fold cross evaluation, and over 1 percentage point in highly iterated runs with continuous resectioning.