WarSampo Knowledge Graph contains data about Finland in the Second World War as Linked Open Data, including metadata of more than 100 000 people. A crucial part is the casualty register, containing detailed information about all Finnish soldiers who perished during the war, consisting of 94 676 person records. This register contains occupational labels, which have been manually harmonized into an ontology and linked to occupational classifications such as HISCO. This paper gives an overview of the harmonized occupation ontology and provides an outlook of how the ontology, the occupation classifications, and related social measures could be used for prosopographical study in the future to provide new insights into events of the war or of the surrounding society.
This volume is the final selection of papers presented at the Biographical Data in a Digital World 2022 workshop, co-located with the Digital Humanities 2022 (DH2022) conference, the leading conference series in the field of digital humanities, which took place in Tokyo 25-29 July 2022. The Biographical Data in a Digital World 2022 Workshop was an online event, held on July 25th. The papers in the conference and the proceedings cover three themes: network analysis and semantic web; finding and preparing biographical data for research; and use cases and advanced ways of working with biographies and biographical data.
To facilitate discovery of premodern manuscripts in U.S. memory institutions, Digital Scriptorium, a growing consortium of over 35 institutional members representing American libraries, museums, and other cultural heritage institutions, has developed a digital platform for an online national union catalog. The platform will allow low-barrier and efficient collection, aggregation, and enrichment of member metadata and sustainably publish it as Linked Open Data. This article describes the methods and principles behind the data model development and the decision to use Wikibase. The results of the prototype implementation and testing phase demonstrate the practicality and sustainability of Digital Scriptorium’s approach to building an online national union catalog based on Linked Open Data technologies and practices.
The paper documents the process of collecting, consolidating, and publishing epistolary metadata from Finnish cultural heritage organizations to create an archive for bottom-up analyses of 19th-century epistolary culture. We describe and discuss the data survey that was conducted to gather information about available letter collections across Finland, as well as the cleaning and harmonizing of over 350,000 letters from twelve different sources in various digital formats. We have also developed a data model that combines event-based and letter-based aspects of the metadata. Furthermore, the paper contributes to the ongoing discussion of the initial phases of data-intensive research and the importance of discussing the labor of cleaning data. We believe that our experiences described in this paper can have wider significance for other digital humanities projects in Europe.
As is the case with several north and west European countries, Finland’s legislation allows for the hobbyist discovery of archaeological material, most commonly through metal detecting. Following in the footsteps of such countries as England and Wales with the Portable Antiquities Scheme (PAS), the Netherlands with Portable Antiquities of the Netherlands (PAN), MEDEA in Belgium, and DIME in Denmark, Finland is also developing a linked open data finds database to record and disseminate the archaeological information that comes from these non-professional activities. This approach is based upon philosophies of citizen science and participatory research, and considers the democratization of archaeology as a central goal. The interdisciplinary and multi-organizational research team charged with building FindSampo face several challenges in their effort to develop the database in an effective way. In this chapter we document the process of developing FindSampo, paying particular attention to public and community engagement challenges, working with digital data, and professional ethics.
This paper presents the vision of aggregating, harmonizing, and publishing letter catalog metadata (information e.g. of senders, receivers and datings of letters) from cultural heritage (CH) institutions in Finland as a single reconciled Linked Open Data (LOD) service and a semantic portal providing data analytical tools for researchers. The research is conducted as part of the consortium research project Constellations of Correspondence (CoCo). The target of the project is to study – for the first time – scattered, heterogeneous epistolary metadata regarding the period of the Grand Duchy of Finland (1809–1917) as one, integrated dataset and make it interoperable and available. This will enable scholars to ask ambitious research questions in the field of computer science and to conduct empirical, bottom-up case studies e.g. on epistolary culture, communicative networks, and heritagization processes. This paper discusses one of the first datasets acquired by the project, the letter collection of the Board of the Finnish Art Society (1846–1901), provided by the Finnish National Gallery, which contains details of c. 1150 letters sent or received by c. 400 actors.
This paper addresses the problem of maintaining CIDOC CRM-based knowledge graph (KG) by non-expert users. We present a practical method using Wikibase and specific data input conventions for creating and editing linked data that can be exported as CIDOC CRM compliant RDF. Wikibase is a proven and maintained software for generic KG maintenance with a fixed but flexible data model and easyto-use user interface. It runs the collaboratively edited Wikidata KG, as well as increasing amount of domain specific services. The proposed solution introduces a set of data input conventions for Wikibase that can be used to generate CIDOC CRM compliant RDF without programming. The process relies on the aforementioned data input rules combined with generic mapping implementations and metadata stored as part of the KG. We argue that this convention over coding makes the system more easily approachable and maintainable for users that want to adhere to the CIDOC CRM principles, but are not ontology experts. As part of the preliminary evaluation of the proposed solution, an example on managing Cultural Heritage data in the military history domain with discussion on the limitations of the approach is presented.
This paper presents WarMemoirSampo, a portal that provides semantic search and navigation of video interviews with Finnish World War II veterans. The portal associates video fragments with contextual data extracted from the video transcriptions, enabling users to find suitable video segments via faceted search and highlighting relevant content in the video being watched. This is carried out by processing natural language texts in order to extract named entities, keywords and lemmas. The result is a Linked Data Knowledge Graph that underpins the portal. We describe the collaboration between Natural Language Processing and Semantic Web technologies used in order to produce these results.
This paper introduces the system ParliamentSampo – Parliament of Finland on the Semantic Web, a Linked Open Data (LOD) service, data infrastructure, and semantic portal for studying Finnish political culture, language, and networks of the Members of Parliament (MP). The article presents the vision behind the system, the LOD service, and explores the possibilities to utilize it in research and application development. A knowledge graph of linked data has been created based on ca. 962 000 speeches in all plenary sessions of the Parliament of Finland in 1907—2021; the data is also available in XML format, utilizing the new international Parla-CLARIN format. For the first time, the entire time series of the Finnish parliamentary speeches has been converted into data and a data service in a unified format. In addition, the speeches have been interlinked with another knowledge graph created from the database of the MPs and enriched from other data sources into a broader ontology-based data service. The paper shows how the LOD service SPARQL endpoint can be used to research parliamentary culture, the use of political language, and networks of politicians through data analysis. The service endpoint can also be used to develop applications for different user groups without programming skills, such as the ParliamentSampo semantic portal introduced in the paper, too. This application aims to make political decision making more transparent to the general public, media, politicians, and other end users.
This paper discusses building lightweight ontologies for faceted search user interfaces with Named Entity Recognition (NER) from textual data. This is studied in the context of building a Knowledge Graph for the textual indexing of interview videos in the in-use WarMemoirSampo system, consisting of a Linked Open Data service and an open semantic web portal for contextualized video viewing. It is shown that state-of-the-art NER tools are able to find entities from textual data and categorize them with high enough recall and precision to be useful for building facet ontologies, without involving considerable manual domain ontology engineering. To enable entity disambiguation and to be able to show relevant contextual information and useful links for the users of the portal, also Named Entity Linking techniques are employed.
Manuscripts are a crucial form of evidence for research into all aspects of premodern European history and culture, and there are numerous databases devoted to describing them in detail. This descriptive information, however, is typically available only in separate data silos based on incompatible data models and user interfaces. As a result, it has been difficult to study manuscripts comprehensively across these various platforms. To address this challenge, a team of manuscript scholars and computer scientists worked to create “Mapping Manuscript Migrations” (MMM), a semantic portal, and a Linked Open Data service. MMM stands as a successful proof of concept for integrating distinct manuscript datasets into a shared platform for research and discovery with the potential for future expansion. This paper will discuss the major products of the MMM project: a unified data model, a repeatable data transformation pipeline, a Linked Open Data knowledge graph, and a Semantic Web portal. It will also examine the crucial importance of an iterative process of multidisciplinary collaboration embedded throughout the project, enabling humanities researchers to shape the development of a digital platform and tools, while also enabling the same researchers to ask more sophisticated and comprehensive research questions of the aggregated data.
This paper shows how various prosopographical phenomena can be highlighted and visualized in the WarSampo Knowledge Graph that contains rich data about Finland in the Second World War as Linked Open Data, including detailed metadata of more than 100 000 people. WarSampo Portal contains tools for simple prosopographical data analysis of the person registers, and accessing the SPARQL endpoint directly opens up further possibilities in using the ontology infrastructure for enhanced information retrieval and pursuing digital humanities studies. This paper overviews of these possibilities of WarSampo, and presents examples of how it, and by extension Linked Data more generally, can be used to created data analyses to support historical research.
This paper shows how various prosopographical phenomena can be highlighted and visualized in the WarSampo Knowledge Graph, which contains rich data about Finland in the Second World War in Linked Open Data format. WarSampo Portal contains tools for simple prosopographical data analysis of the person registers, and accessing the SPARQL endpoint directly opens up further possibilities in using the ontology infrastructure for enhanced information retrieval and pursuing comparative studies. This paper gives an overview of the possibilities of WarSampo, and presents examples of how WarSampo, and by extension Linked Data more generally, can be used to created various analyses to help with historical research.
This paper presents the FindSampo system for analyzing and disseminating archaeological object finds made by the public. The system is based on Linked Open Data (LOD), and consists of a web portal and an open data service. The underlying knowledge graph contains data of some 3000 archaeological object finds catalogued in the archaeological collection of the Finnish Heritage Agency (FHA) from 2015 to 2020. The portal and LOD service have been open to public use since May 2021.
This paper presents WarMemoirSampo, a portal that provides semantic search and navigation of video interviews with Finnish World War II veterans. The portal associates video fragments with contextual data extracted from the video transcriptions, enabling users to find suitable video segments via faceted search and highlighting relevant content in the video being watched. This is carried out by processing natural language texts in order to extract named entities, keywords and lemmas. The result is a Linked Data Knowledge Graph that underpins the portal. We describe the collaboration between Natural Language Processing and Semantic Web technologies used in order to produce these results.
This paper presents a new software framework, Sampo-UI, for developing user interfaces for semantic portals. The goal is to provide the end-user with multiple application perspectives to Linked Data knowledge graphs, and a two-step usage cycle based on faceted search combined with ready-to-use tooling for data analysis. For the software developer, the Sampo-UI framework makes it possible to create highly customizable, user-friendly, and responsive user interfaces using current state-of-the-art JavaScript libraries and data from SPARQL endpoints, while saving substantial coding effort. Sampo-UI is published on GitHub under the open MIT License and has been utilized in several internal and external projects. The framework has been used thus far in creating six published and five forth-coming portals, mostly related to the Cultural Heritage domain, that have had tens of thousands of end-users on the Web.
This paper argues for the idea of publishing legislation and case law as Linked Open Data (LOD) on the Semantic Web, to cater several user groups, including the general public, legislators, lawyers, researchers of legal informatics, and application developers. To support the argument, the proof-of-concept system LawSampo – Finnish Legislation and Case Law on the Semantic Web is introduced, including a semantic portal and a LOD service. Based on the Sampo Model, the main novelty of LawSampo is the provision of heterogenous distributed legal data through multiple application perspectives for faceted searching and exploring the data and for data analysis in legal informatics.
Manuscripts are a crucial form of evidence for research into all aspects of premodern European history and culture, and there are numerous databases devoted to describing them in detail. This descriptive information, however, is typically available only in separate data silos based on incompatible data models and user interfaces. As a result, it has been difficult to study manuscripts comprehensively across these various platforms. To address this challenge, a team of manuscript scholars and computer scientists worked to create “Mapping Manuscript Migrations” (MMM), a semantic portal, and a Linked Open Data service. MMM stands as a successful proof of concept for integrating distinct manuscript datasets into a shared platform for research and discovery with the potential for future expansion. This paper will discuss the major products of the MMM project: a unified data model, a repeatable data transformation pipeline, a Linked Open Data knowledge graph, and a Semantic Web portal. It will also examine the crucial importance of an iterative process of multidisciplinary collaboration embedded throughout the project, enabling humanities researchers to shape the development of a digital platform and tools, while also enabling the same researchers to ask more sophisticated and comprehensive research questions of the aggregated data.
This paper introduces the system ParliamentSampo – Parliament of Finland on the Semantic Web , a Linked Open Data (LOD) service, data infrastructure, and semantic portal for studying Finnish political culture, language, and networks of the Members of Parliament (MP). The article presents the vision behind the system, the LOD service, and explores the possibilities to utilize it in research and application development. A knowledge graph of linked data has been created based on ca. 962 000 speeches in all plenary sessions of the Parliament of Finland in 1907—2021; the data is also available in XML format, utilizing the new international Parla-CLARIN format. For the first time, the entire time series of the Finnish parliamentary speeches has been converted into data and a data service in a unified format. In addition, the speeches have been interlinked with another knowledge graph created from the database of the MPs and enriched from other data sources into a broader ontology-based data service. The paper shows how the LOD service SPARQL endpoint can be used to research parliamentary culture, the use of political language, and networks of politicians through data analysis. The service endpoint can also be used to develop applications for different user groups without programming skills, such as the ParliamentSampo semantic portal introduced in the paper, too. This application aims to make political decision making more transparent to the general public, media, politicians, and other end users.
Eetu Mäkelä合作论文数Laboratory of Media Technology, Aalto University9