
This paper introduces the system ParliamentSampo – Parliament of Finland on the Semantic Web , a Linked Open Data (LOD) service, data infrastructure, and semantic portal for studying Finnish political culture, language, and networks of the Members of Parliament (MP). The article presents the vision behind the system, the LOD service, and explores the possibilities to utilize it in research and application development. A knowledge graph of linked data has been created based on ca. 962 000 speeches in all plenary sessions of the Parliament of Finland in 1907—2021; the data is also available in XML format, utilizing the new international Parla-CLARIN format. For the first time, the entire time series of the Finnish parliamentary speeches has been converted into data and a data service in a unified format. In addition, the speeches have been interlinked with another knowledge graph created from the database of the MPs and enriched from other data sources into a broader ontology-based data service. The paper shows how the LOD service SPARQL endpoint can be used to research parliamentary culture, the use of political language, and networks of politicians through data analysis. The service endpoint can also be used to develop applications for different user groups without programming skills, such as the ParliamentSampo semantic portal introduced in the paper, too. This application aims to make political decision making more transparent to the general public, media, politicians, and other end users.
This paper presents the vision of aggregating, harmonizing, and publishing letter catalog metadata (information e.g. of senders, receivers and datings of letters) from cultural heritage (CH) institutions in Finland as a single reconciled Linked Open Data (LOD) service and a semantic portal providing data analytical tools for researchers. The research is conducted as part of the consortium research project Constellations of Correspondence (CoCo). The target of the project is to study – for the first time – scattered, heterogeneous epistolary metadata regarding the period of the Grand Duchy of Finland (1809–1917) as one, integrated dataset and make it interoperable and available. This will enable scholars to ask ambitious research questions in the field of computer science and to conduct empirical, bottom-up case studies e.g. on epistolary culture, communicative networks, and heritagization processes. This paper discusses one of the first datasets acquired by the project, the letter collection of the Board of the Finnish Art Society (1846–1901), provided by the Finnish National Gallery, which contains details of c. 1150 letters sent or received by c. 400 actors.
This paper shows how various prosopographical phenomena can be highlighted and visualized in the WarSampo Knowledge Graph, which contains rich data about Finland in the Second World War in Linked Open Data format. WarSampo Portal contains tools for simple prosopographical data analysis of the person registers, and accessing the SPARQL endpoint directly opens up further possibilities in using the ontology infrastructure for enhanced information retrieval and pursuing comparative studies. This paper gives an overview of the possibilities of WarSampo, and presents examples of how WarSampo, and by extension Linked Data more generally, can be used to created various analyses to help with historical research.
This work focuses on questions of knowledge organization related to literary fiction. How can LGBTQI fictional literature become more accessible to readers and scholars? The project Queerlit Metadata Development and Searchability for LGBTQI Literary Heritage addresses this question in two ways: by the development of a thesaurus for the description of Swedish LGBTQI literature, and by building a curated bibliographical database for this material with flexible search options. Despite the community and scholarly interest in LGBTQI literature, relevant LGBTQI literature is hard to find both for readers and researchers. Subject indexing is underdeveloped for this topic, and subject headings have been historically inadequate and offensive. The paper focuses on how LGBTQI literature can be made more easily accessible through subject indexing. This will make new research possible, such as gaining overviews of the development of specific themes over time, the presence of LGBTQI literature within or outside of the literary canon or in different genres and changing ideas and perceptions concerning sexualities and gender identities. It will also accommodate user’s needs of better access to LGBTQI themed fictional literature.
This poster describes the creation and publication of a curated, cumulative, open – living – bibliography of the organisation Digital Humanities in the Nordic and Baltic Countries (DHNB). It recounts how the bibliographic material - conference abstracts, posters, peer-reviewed papers, journal articles, but also news items about the conferences, blog posts, forum posts - has been collected in a Zotero library. It discusses ways of documenting and archiving the bibliography both as a set of research data in a Dataverse repository and as an open Zotero group library. The poster concludes with the prospect of opening up the bibliography to become a community driven project that will be expanded and maintained by the DHNB community and made explorable for research both as a live bibliography and as a data set for further analyses.
The DARIAH-EU infrastructure for Digital Humanities (DH) is often focusing on using structured data for quantitative studies, while the EU-CLARIN infrastructure deals primarily with unstructured natural language texts. However, in DH research both texts and structured data are often needed. It therefore makes sense to develop and use both infrastructures together, as suggested in the Dutch CLARIAH programme and the corresponding FIN-CLARIAH initiative in Finland, a new part of the Finnish research infrastructure road map of the Academy of Finland. This poster paper introduces work in FIN-CLARIAH relating to the idea of integrating natural language processing (NLP) tools with the Linked Open Data (LOD) Infrastructure for Digital Humanities in Finland (LODI4DH). We present a plan for NLP services to be opened as part of the Linked Data Finland (LDF.fi) platform. The new services are used for knowledge extraction from Finnish texts for weaving LOD, and on the other hand for language DH data analyses of the published datasets in applications in many domains, such as political culture. The extended LDF.fi platform will provide users with documented APIs for NLP services using unified output formats as well as software delivery as Docker containers, to lower the bar for deployment.
This study provides an exploratory attempt to develop a framework for how to semi-automatically annotate salient topics in Swedish parliamentary debate. The discussion is grounded in the ongoing digital humanities project SweTerror that studies the terrorism discourse in the Riksdag 1968–2018 through a mixed-methods approach. The paper presents our tentative framework through its three main categories: metadata, language data and frame data. While the first two categories are mostly generic and their data could mainly be automatically extracted, the third category is contextual and requires manual interpretation. We discuss the design of the latter through the theoretical concept of ‘framing’ and illustrate the framework’s overall principles through a case study of utterances in the debates 1968–1970 concerning terrorism. We conclude by suggesting that it may be more generally applicable for studies of parliamentary debates in HSS research if further modified for the particular research purposes.
This paper introduces a Sentiment and Emotion Lexicon for Finnish (SELF) and a Finnish Emotion Intensity Lexicon (FEIL). Sentiment analysis and emotion detection require annotated data regardless of the chosen approach, but most existing resources are for the English language. To overcome this, the SELF and FEIL lexicons use projected annotations from existing resources with carefully edited translations and domain adaptations. In this paper the creation process and translation issues are explained in detail to allow others to create similar lexicons for other languages. The usefulness of SELF and FEIL are demonstrated via several interdisciplinary affect-related projects. To our best knowledge, this is the first comprehensive sentiment and emotion lexicon for Finnish.
The temporal aspects of politics have been discussed extensively by political theorists, but have not been explored using grammatically parsed textual datasets. This paper explores the ways in which future, present and past are projected and referred to in speeches in the Finnish parliament that talk about ideologies. Ideologies are crucial categories of thinking about the political past and future and therefore serve as a case in which temporality is expressed in a variety of ways. We use a dataset drawn from Finnish parliamentary records from 1980 to 2021 and operationalize morpho-syntactic information on clause structures and grammatical tense system to explore the different temporal profiles of ideologies. We show how some isms, like communism and fascism, are much more likely to appear in the context of the past, whereas others, like capitalism and racism, tend to appear in the present tense. We further develop a framework for analyzing temporality based on clause structures and grammatical tense and relate that to how the study of politics has approached time in parliamentary speaking.
In this article I analyse parliamentary debates of the Finnish Parliament ( Eduskunta ) on European integration from 1990 to 2020. Finland joined the European Union (EU) in 1995, but Finland’s integration history dates back to the late 1950s and early 1960s. In the turbulent years following the end of the Cold War and the collapse of the Soviet Union in the early 1990s, European integration rose higher on the Finnish political agenda. The data used in this article consist of a machine-readable database of plenary protocols of the Finnish Parliament. The main database covers the whole lifespan of the modern Finnish Parliament since 1906. The dataset used for the analysis contains all plenary speeches with references to “Europe”, “European” and “Europeanism” (N=25,674), together with adequate metadata. The core analysis focuses on six time windows, each with a span of three years. These focus widows are linked to nationally important key European events. Methodologically, the article is rooted in Exploratory Data Analysis (EDA) and applies different text mining tools to explore, analyse and visualise how members of the Finnish Eduskunta politicise and debate European issues. The analysis is carried out in three steps. First, I use traditional term-based text mining methods to explore the vocabulary used in the debate, both across time and by parliamentary faction. In the second step, I use tf-idf analysis to explore the vocabulary differentiating parliamentary factions. The analysis is rounded out in the third step by the application of Text Network Analysis (TNA). I apply TNA to explore and visualise topics in the collection of plenary speeches and, thus, to evidence the power and usefulness of this novel method as a complement to other topic modelling methods. Overall, the results presented in this article find strong support when critically reflected against findings from previous studies. The results also significantly improve our knowledge and understanding of national parliamentary debates on European integration. Further, the article is encouraging when it comes to the application of computational methods and tools on large corpora of unstructured political texts.
Machine translation (MT) models have become increasingly accurate and widely accessible for multiple languages in recent years. They can potentially lift the barriers to applying NLP tools and methods to previously unsupported languages and boost comparative cross-lingual research in digital humanities. This study empirically contrasts results obtained with source and target Slovenian ParlaMint corpus of parliamentary debates on topic modelling. It qualitatively compares three steps in topic interpretation: topic description, topic significance in subcorpora, and marginal topic distribution. The results indicate that the topic modelling on the target corpus only partially replicates the topic modelling on the source corpus, but the overlap is sufficient to provide a starting point for the cross-country comparison.
REPUBLIC is a five-year project to compile a digital online edition of the resolutions of the Dutch States General issued between 1576 and 1796. In this article we compare the expected outcome of the project with the book and online editions that preceded it, and we assess the losses and gains of the new approach.
This paper explores the contexts of the keywords “propaganda”, “information” and “upplysning” in the Swedish parliamentary debate protocols, from 1920 to 2019. The digitized protocols have recently been annotated with metadata for speakers’ gender and party affiliation. Based on perspectives developed within conceptual history, we have traced the concepts in the parliamentary debates and used computational methods to cluster the contexts in which they occur. Word windows around the three keywords were compiled into a sub-corpus, and topic modelling was used to cluster the contexts. The findings show that the distribution of the topics gets more even over time, partly explained by the spread of the term “information” in various political areas in the mid-20th century. Furthermore, the only distinct topic shared between the three keywords relates to campaigns to limit and prevent the consumption of alcohol, narcotics and tobacco. While a conceptual shift takes place within this topic, from “upplysning” to “information”, it is also shown that it was possible to discuss and frame these issues in terms of “propaganda” post-WWII – it even became more common to do so in the 1950s and 1960s.
This article wish to make a case for scholarly digital editions (SDE’s). SDE’s can help to put the computational humanities into the center of humanistic scholarship. Large scaled projects pave the way for easy access to historical documents, but small, curated digital editions, strenuously enriched by philologists, will be the key player in the process of integrating computational humanities into traditional scholarship; first and foremost because this material is clean, reliable, and flexible, secondly, thorough markup leaves it open to comprehensive, fine-grained, hermeneutically complex explorations.
In this long paper, we use NLP techniques to explore two decades (1881-1899) of parliamentary debates of the French Third Republic (1870-1940), and more specifically to analyse the importance of the army in the political debate. We use Latent Dirichlet Allocation to partition the vocabulary into topics, and then study the distribution of the topic “army” over time. We also examine its connection with other topics, in relation to the main political and military events of the period.
Latvian Romani is a Northeastern Romani dialect with a limited number of publicly available sources. Two large archival collections of texts in Latvian Romani, compiled primarily in the 1930s in Latvia and Estonia, have been recently digitized as images and made available online for a wider public. In our study, we focus on one of these collections, the Latvian Romani folklore texts collected by Jānis Leimanis in interwar Latvia. In this paper, we describe how initial manual transcriptions, most of which have been created with the help of a special crowdsourcing platform, were integrated in the handwritten text recognition (HTR) workflow in Transkribus. We present two HTR models trained on the basis of Leimanis’ collection and discuss various issues related to the work on these texts.
Transcriptions in different languages are a ubiquitous data format in linguistics and in many other fields in the humanities. However, the majority of these resources remain both under-used and under-studied. This may be the case even when the materials have been published in print, but is certainly the case for the majority of unpublished transcriptions. Our paper presents a workflow adapted in the research project Language Documentation Meets Language Technology, which combines text recognition, automatic transliteration and forced alignment into a process which allows us to convert earlier transcribed documents to a structure that is comparable with contemporary language documentation corpora. This has complex practical and methodological considerations.
In the early modern period, the Imperial Diet (or Reichstag ) played a central role in the constitutional structure of the Holy Roman Empire and had a significant impact on European politics. This is documented by the variety of handed down source material, including negotiation files ( Verhandlungsakten ), minutes ( Protokolle ), reports of individual envoys to their princes ( Berichte ), and petitions ( Supplikationen ). The DFG and FWF funded project The Imperial Diet of Regensburg of 1576 – a Pilot Project on the Digital Edition of Sources on the Early Modern Era is breaking new ground and adds a new chapter to the editorial history of the Imperial Diet records which have been edited by the Historical Commission at the Bavarian Academy of Sciences since the 19th century. For the first time a database with an archival documentation of surviving manuscripts and edited texts will be made available as a digital edition. Furthermore, the editorial focus will be on a central aspect of the Imperial Diets, communication and interaction of the various political agents upfront and during the event. The resulting ontology described in this paper will be the basis for search operations and the integration of the edition´s RDF data into the Semantic Web.
The paper outlines initial experience of using transcription as a tool to promote deep reading and understanding of cultural heritage. The experience is a side activity of a project, which aims to improve the database of Latvian folk songs. Use of transcription appeared beneficial not only in terms of elaborating a data base, but also for stimulating deep reading and comprehension of folk songs. The analysis of experiences shows the benefits and challenges of using transcription as a tool for working with students in humanities and arts.
Different parliamentary activities allow Members of Parliament (MPs) varying amounts of autonomy. Previous studies have shown that, in parliamentary systems with strong parties and party-centered electoral rules, MPs have limited room for crossing the party line in the legislature both in voting and speech. Further, party-centered systems limit MP’s ability to address electoral concerns of their constituency; they are less responsive. In this paper, I combine these findings by showing that even within system variation in party control over institutions affects the levels of responsiveness in parliamentary questions. By linking MP’s constituency mentions with different types of questions, my results show that the institutional design in the Norwegian Storting affects the level of MP constituency signaling. Specifically, I show that questions with low levels of party control and public attention (written questions and question time) give MPs far more opportunity to raise constituency specific issues than the more party controlled activities (interpellations and question hours). Consequently, I argue that responsiveness does not disappear in party-centered systems; it is located at lower-level institutions. Particularly, some types of questions, where shirking from the party line is less consequential and the party organizations have less control over its members, allow for constituency signaling.