
Psychological wellbeing has long been studied through surveys and interviews that rely on individuals’ self-reported experiences. With the rise of digital communication, however, new opportunities have emerged to observe how wellbeing is expressed organically in everyday language. This study investigates how the multidimensional components of a wellbeing framework -- the PERMA+H framework, are reflected in Finnish online discussions, using a large-scale dataset from the Suomi24 forum spanning from 2001 to 2020. By applying a lexicon-based approach to extract wellbeing-related features and employing Self-Organizing Maps (SOM) for clustering and visualisation, the study identifies distinct patterns of wellbeing expression across different subforums and over time. The analysis reveals thematic zones of wellbeing: Positive Emotion and Relationships often co-occur in emotionally expressive spaces, while Engagement and Accomplishment are prevalent in discussions of work and study. A “low wellbeing” region was also identified, alongside a gradual decline in wellbeing indicators across the two-decade period. These findings suggest that online discourse can meaningfully reflect psychological states and their temporal shifts. By bridging positive psychology, computational social science, and the emerging field of wellbeing informatics, this research offers new insights into how emotional life manifests and evolves within digital communities.
This paper examines community engagement, participatory methodologies, and inclusive practices in digital humanities as essential strategies for data collection during wartime. The Saint Sophia Cathedral in Kyiv, Ukraine, houses over 7,000 inscriptions spanning more than a millennium, composed in diverse languages and scripts. These inscriptions are multilingual, multimodal, and layered, offering a rich yet endangered source for historical and cultural research amid ongoing Russian aggression. Since February 2022, over 470 culturally significant sites have been damaged or destroyed. While Saint Sophia has not been a direct target, its proximity to key infrastructure renders it vulnerable to missile and drone strikes. In response, researchers from Sweden have employed stakeholder interviews, knowledge exchange, and participatory data collection methods to preserve this heritage. These approaches challenge the prevalent assumption within the heritage sector that documentation and digitisation are definitive processes, while also addressing the institutional divide between technical and heritage expertise. The “Digital Documentation of Inscriptions in the Saint Sophia Cathedral in Kyiv” project began by identifying key stakeholders and their data needs through interviews and the establishment of a reference group. This group enabled broader expert participation in an initial workshop aimed at evaluating and refining digitisation methods. Museum personnel at Saint Sophia have played a central role in testing and assessing these methods on site. They have also been trained in workflows that enable them to independently conduct high-quality digitisation, data preprocessing, and data management, ensuring local capacity building and sustainable heritage preservation.
Digitized historical newspapers are a promising source for economic, social, and political history. In spite of the fact that many historical newspapers are available today in some digitized form, their quantitative analysis has only sporadically found its way into the canon of scientific methods and is far from becoming a standard approach, mostly due to access limitations and unsuitable data formats for digital analysis. We propose to re-process large quantities of digital images of newspapers into a better machine-readable and paragraph-segmented form and use topic modeling techniques to identify and track topics in newspapers over time and to create topic-specific sub-corpora. This approach will serve to identify relevant articles for any number of further research questions in a mere matter of hours, eliminating months of flicking through web-viewers or copying results from keyword searches. Most topic models are designed for smaller corpora. Since historical newspapers are now available in enormous quantities, their applicability stands to be questioned. We create a large dataset by re-processing all digitally available pages of the Kölnische Zeitung, consisting of 425,194 pages from 1803--1945. Subsequently, we investigate the application of four different topic models (Gensim LDA, tomotopy LDA, LeetTopic and BERTopic) on this large dataset, to demonstrate their (un)suitability for processing datasets of this scale. Among these methods, the tomotopy LDA implementation proves most reliable on large datasets. We show that topic modeling can easily isolate sections of the newspaper which are of particular interest to researchers, like trade registry entries and various kinds of labor market ads, but can also identify and isolate prominent political topics in the newspaper articles.
This paper aims to establish a tentative, minimal framework for the implementation of Artificial Intelligence (AI) in heritage collections with a particular focus on collections management and curation. We begin by summarising relevant scholarship on the need for a diverse and inclusive understanding of heritage, including the complexities that AI introduces. We then conduct qualitative fieldwork by drawing upon a sample of semi-structured interviews conducted with Swedish cultural heritage professionals in curatorial roles. Building on a more comprehensive understanding of the strengths, weaknesses, opportunities, and threats that AI presents for cultural heritage collections, this article proposes a minimal working model to support successful technological implementation. We argue that such implementation requires a modular approach. Drawing on a small case study within the Swedish context, we outline a minimal framework composed of two core modules: (1) domain expertise and (2) collections as information and data.
This article presents a study on the use of artificial intelligence to analyze Finnish-language newspapers published in North America between 1876 and 1923. Using GPT-4 and LLaMA 3.1, we develop and evaluate a text segmentation method on large-scale digitized historical data, and we propose a hierarchical taxonomy for register (genre) classification based on manual annotation. We present a methodology for identifying text boundaries and, separately, a manually developed register taxonomy used to classify segments. Our analysis is grounded in a manually annotated stratified corpus of 500 documents selected from 312,300 digitized pages. Results demonstrate strong segmentation accuracy, contributing both to digital humanities methodology and the historical understanding of the Finnish-American immigrant press.
How can we create exciting and engaging immersive experiences with the help of cutting-edge technologies without compromising the primary educational outcomes? Within the framework of the research project “Evidence, source criticism and credibility in the digital transformation of museums” we are investigating how the museum landscape has been shifting due to the increasing popularity of the use of immersive technologies. We argue that these technological novelties offer new opportunities to enhance traditional educational practices and engage visitors. However, this does not automatically translate into desirable educational outcomes. The challenge lies in ensuring that while producing immersive experiences in the GLAM sector, the pedagogical learning output is the main focus rather than achieving a technologically impressive outcome.
The 9th Digital Humanities in the Nordic and Baltic Countries Conference “Digital Dreams and Practices” (DHNB 2025), held in Tartu, Estonia, showcased a broad range of research, methodological innovation, and institutional collaborations. This article provides an analytical overview of the conference programme, participation, and research profile and provides an analytical framing of the material represented in the book of abstracts. Drawing on conference metadata from the ConfTool system, author-supplied keywords and institutional affiliations, the article examines workshops, methodological trends, data types, and collaboration patterns. The paper analyses four interrelated dimensions: (1) the role of workshops and tutorials as laboratories for pedagogical innovation and infrastructural literacy; (2) dominant methodological trends, particularly the growing integration of artificial intelligence, linked open data, and multimodal cultural analytics; (3) the diversity of research materials, ranging from textual corpora to audiovisual, spatial, and networked data; and (4) institutional and international collaboration patterns shaping the regional research ecosystem. The article highlights how DHNB conferences function as a transnational hub where experimental “digital dreams” are translated into sustainable scholarly and infrastructural practices.
This paper introduces the Dawit Isaak Database of Censorship (DIDOC), a pilot initiative designed to advance research on censorship practices through the systematic documentation of banned and suppressed literature. Developed in collaboration among the Dawit Isaak Library, the Gothenburg Research Infrastructure for Digital Humanities (GRIDH), Swedish PEN, and Lund University, the project combines cultural heritage resources, scholarly expertise, and digital infrastructure. DIDOC builds on a curated selection of 160 titles from the Dawit Isaak Library’s larger collection of approximately 1,200 censored works. Each entry links bibliographic information to documented censorship events, categorized by type, reason, and geographical location. Implemented on the Omeka-S platform and structured according to Linked Open Data principles, the database supports semantic annotation, interoperability, and future integration with external resources such as the Norwegian Beacon for Freedom of Expression dataset. The project contributes methodologically by demonstrating how digital infrastructures and metadata standards can be applied to the study of censorship, while also addressing ethical questions concerning data selection, representation, and public accessibility. Beyond its role as a public resource for educators, journalists, and in libraries, DIDOC establishes a foundation for international comparative research on censorship across different historical and geographical contexts.
This paper introduces the Finnish Named Entity Linker (FINEL), a tool that leverages Deep Learning models, including Large Language Models (LLMs), to recognize, disambiguate, and link Named Entities in Cultural Heritage texts. FINEL is designed to enhance the metadata of textual documents by connecting them to Knowledge Graphs (KG). We propose a zero-shot classification method that resembles Retrieval-Augmented Generation (RAG) and discuss a prototype web service with a user interface that enables human intervention for final disambiguation decisions. This editing capability is crucial, particularly when automatic linking may be hindered by errors and hallucinations inherent in LLM-based tools. The paper also reflects on lessons learned from using FINEL in applications targeting Digital Humanities (DH) research. Since the focus is on Finnish texts, our methods accommodate the specific challenges posed by this highly inflectional language and the available processing resources. Preliminary evaluation results underscore the potential of FINEL: our named entity lemmatizer achieved an accuracy of 96.5% on the test dataset, while an LLM from the Llama family reached 97% accuracy for entities with only one candidate. However, accuracy decreased with each additional candidate.
This paper introduces and discusses the new research and educational resource Mapping Saints, built by the cultural heritage research and digitization project, Mapping Lived Religion: The Medieval Cults of Saints in Sweden and Finland (MLR), launched in November 2024. The goal of the project was to create a digital resource to enable new, interdisciplinary approaches to the study of the cults of saints, applying research-driven digitalization with a theoretical focus on “lived religion”. The project was also tasked with the digitization of two cultural heritage collections: the analogue card catalogue, Iconographical Index of Ecclesiastical Art in Sweden (Ikonografiska registret) at the Swedish National Heritage Board; and the collection of photographs of ecclesiastical art taken by Lennart Karlsson, held by the Swedish History Museum, and previously available as low-resolution images in the database The Medieval World of Images (Medeltidens bildvärld). The paper critically reflects on the creation of bespoke tools for research projects in the digital humanities in its discussion of the digital methods applied in the project, which are of importance to medieval religious studies, a long-established field. These include the foundation of the resource itself: the place register and spatial analysis of past phenomena, as well as the application of the principles of linked open data to ensure sustainability. The paper concludes by providing two concrete cases which show the usefulness of Mapping Saints for research into lived religion: the inclusion of altars as unique places and the analytical possibilities of identifying and mapping previously unknown medieval individuals.
The Dictionary of the Latvian Language [Latviešu valodas vārdnīca] is a unique general scientific dictionary of the Latvian language, which was started in the 1880s by Kārlis Mühlenbachs (1923–1932), then edited, supplemented, and completed by Jānis Endzelīns and Edīte Hauzenberga-Šturma (1934–1946). It is both an explanatory and a translation dictionary with features of an etymological and synonym dictionary (Nespore et al. 2006). In 1994, the digitization of the dictionary and its additional volumes was started at the Institute of Mathematics and Computer Science, University of Latvia. In 2002, the electronic version of the dictionary (MEV) was published. It contained 132,718 entries. To make the dictionary more accessible, in 2024–2025, the MEV application was modernized and integrated with the Tēzaurs.lv platform (Grasmanis et al. 2023; Mīlenbaha-Endzelīna vārdnīca 2000–2025). The article discusses the digitization process and methods of the resource, as well as outlines directions for future work.
Historical letters in archives are typically stored in fonds, based on letter recipients. To analyze the correspondences of a correspondent x, one has to aggregate the received letters in x’s own fond with the letters in the fonds of the correspondents yi who have received letters from x. To address this challenge, epistolary data services have been created by aggregating data from distributed heterogeneous archival data silos and fonds. This paper presents an overview of a new in-use data service and portal for this task, LetterSampo Finland – Finnish Nineteenth-Century Letters on the Semantic Web. In contrast to various legacy services online, this system is based on a Linked Open Data (LOD) Knowledge Graph (KG) in a SPARQL endpoint aggregated and harmonized from distributed heterogeneous data sources, in our case from 16 Finnish cultural heritage organizations and over 1,600 fonds. The new massive KG contains metadata about nearly 1.3 million letters sent or received in the Grand Duchy of Finland during 1809–1917, including also letter contents from four critical editions of correspondences of prominent Finns. The paper shows how this new KG and a portal on top of it can be used for searching and browsing letter data and for data analysis in digital humanities research. We show how the aggregated datasets are related to and enrich each other, pinpointing semantic challenges of data aggregation and linking processes needed. This kind of analysis is needed to make enriched LOD more transparent to the end user and to enhance data literacy for reliable computational analyzes.
This paper presents the design and early implementation of the Finno-Ugric Data Sharing Space (DSS), a multilingual, community-driven prototype for linking cultural heritage data across institutional and geographic boundaries. Rather than a finished infrastructure, the DSS should be read as a blueprint and exploratory model – developed as a thought experiment with minimal resources but grounded in our prior work on music metadata governance involving both public and private actors. We use this experimental setting to review structural problems in existing Finno-Ugric knowledge systems: the negative outcomes of Wikipedia’s Livonian and Mari initiatives, the dispersion of diasporic knowledge, and the limitations of national GLAM infrastructures. Building on empirical literature and our own governance practice, we propose a lightweight federated infrastructure built on Wikibase and open ontologies, which enables multilingual vocabularies, contextual annotation, and ethical data linking without flattening local epistemologies. Case studies of Seto textile collections and the Hõimulõimed multilingual song archive illustrate how the prototype supports cultural reconstruction and participatory enrichment. While not an institutional solution, the DSS demonstrates how a semantically rich, community-anchored model can serve as a testbed for broader applications in low-scale cultural ecosystems.
This paper introduces version 1.0 of riksdagsdebatter.se, a freely available and user-friendly interface for searching, filtering, and exploring Swedish parliamentary speeches since 1867. The underlying corpus is based on open-access data from the Swedish Parliament, further processed and enriched within the SWERIK project, including re-OCRing the debate records, segmenting the text into individual speeches, and adding structured metadata, with ongoing improvements to data quality. The corpus currently comprises over one million annotated speeches, totaling approximately 450 million words, and offers a comprehensive record of political debate in modern Swedish history. This resource has wide-ranging applications for researchers and students in digital humanities, history, political science, and sociology, as well as for journalists, policymakers, and the general public. Version 1.0 of riksdagsdebatter.se provides four primary tools: word trends, keyword in context, n-gram analysis, and the ability to create and download filtered sub-corpora.
Efforts to automate cataloging in libraries have progressed significantly, with AI tools like ChatGPT emerging as potential aids. However, automating the subject indexing of LGBTQ+ fiction poses unique challenges. Traditional fiction indexing often overlooks specific themes and characters sought by users, particularly those related to LGBTQ+ identity and issues. Generative Pre-trained Transformers (GPTs) like ChatGPT offer promise in producing detailed subject terms but face biases and inaccuracies. This study explores ChatGPT’s efficacy in generating subject index terms for LGBTQ+ fiction by comparing AI-generated terms with those assigned by professional information specialists in the Queerlit database. The Queerlit database, which uses the QLIT thesaurus for LGBTQ+ terms and general Swedish controlled vocabularies, provides a gold standard for this comparison. Using a sample of 20 full-text works and 20 metadata records from the Queerlit database, ChatGPT was tasked with generating subject index terms. The evaluation revealed that ChatGPT struggled to identify any LGBTQ+ themes, often producing broader and irrelevant terms, even when the index terms were given as input in the metadata. The precision and recall scores were low, highlighting AI’s limitations in this context. The study underscores the need for careful evaluation of AI tools in library and information science and professional practice, particularly for indexing fiction and minority representation. Future research should involve collaboration with both information and subject experts to examine the potential of automatically generated terms that were not previously assigned, as well as to explore the possibility of refining automated indexing methods, and to address inherent biases in AI models.
This paper presents and discusses parts of the entity model of the runological database “Sigrdrifa”. This database was developed with the aim to create the technological infrastructure for future born-digital editions of runic inscriptions. As such, great importance was placed on developing a system permitting runologists to store pertinent information not just about the inscription texts, but also the visual shapes and forms of the runes as they appear on the objects. The requirement to store information at two different levels (text and visual shape) turned out to require complex modelling, which is the topic of this paper. First, the form variations in runic corpora and the issues encountered when attempting to encode runic inscriptions digitally using the Unicode block Runic are described. Then a potential solution for encoding using a minimal character set based on Runic is outlined, which enables encoding of the visual rune shapes at two different levels of detail by using a two-digit code in addition to the Unicode character. Following this, the parts of the entity model designed to store this information are described. Lastly, the possibilities for adding free text descriptions in addition to the standardised encoded characters and the application of the encoding system on a smaller dataset of runic inscriptions are presented.
This paper introduces Quantitative Close Reading (QCR), a Human-in-the-Loop (HITL) methodology designed to structure and enrich noisy digitized historical corpora – particularly text archives compromised by poor Optical Character Recognition (OCR) quality. QCR integrates manual annotation with quantitative analysis, enabling researchers to trace semantic shifts and symbolic representations over time while preserving interpretive depth. We detail the methodological framework of QCR, including keyword selection, annotation manual development, coder training, and postprocessing techniques. Through three case studies that trace semantic shifts in the representations of fællessang [communal singing], the mythological figure Dana, and the Catholic shrine Lourdes, we demonstrate how QCR bridges close and distant reading, operationalizes conceptual history at scale, and transforms fragmented textual data into reliable, analyzable corpora. By combining the scalability of computational methods with the contextual sensitivity of human annotation, QCR offers a robust approach for historical semantic analysis across digitized texts.
The easy access to generative AI applications, such as ChatGPT, might have huge implications for teaching and learning in higher education. Therefore, more and more policies to regulate the use of AI are published by countries and institutions; however, there is no general overview of the existing policies within CLARIN, a Common Language Resources and Technology Infrastructure – at the member state and institutional level. There is also no overview of how the lecturers and teachers in the CLARIN network are applying generative AI applications in the classroom; furthermore, there is no data available on whether the community has specific concerns regarding the use of AI in education and research. To fill this gap, a small exploratory study was launched in April 2024 within the CLARIN community. This contribution shares the first insights from an online survey and highlights the implications for training and upskilling opportunities.
This paper presents the Uralic Trove, a collection of datasets related to the human past in the Uralic language speaker area with special focus on the area of Finland. All the datasets are made in large collaborations, and (apart from no 8) have initially been launched elsewhere – this paper aims at collecting a ‘trove’ where all these areal datasets are easily located. We briefly describe the contents of eight multidisciplinary open access datasets: The databases considering the whole Uralic family or its speaker area are 1) UraLex (a basic vocabulary database with cognate coding), 2) UraTyp (binary data of existence of given linguistic typological functions), 3) a Cartographical Database of Uralic Languages (GIS format) and another GIS database of Interdisciplinary Maps of North-West Eurasia. The Finnish datasets are more detailed and include 5) a digital, easy to use version of the Dialect Atlas of Finnish collected already 100 years ago, 6) AADA, Archaeological Artefact Database of Finland, 7) Historical Travel Environment model and 8) Historical Culture and Environment Database of Finland. Uralic Trove also includes two user interfaces, that are aimed for easy visualization and access of the data.
This article explores the transformative power of innovation and collaboration in the realm of cultural heritage management, particularly focusing on the conservation and documentation of art exhibitions. Beginning with the creation of Novart, a relational database designed to integrate collection management and exhibition production, the narrative delves into the complexities of contemporary art conservation. It examines the fluid nature of exhibitions and the evolving landscape of conservation practices, emphasizing the need for adaptability and interdisciplinary collaboration. Through practical experiences in institutions like Luma Arles, where limitations of existing tools spurred the quest for innovative solutions, the narrative highlights the persistence of custom database development driven by the imperative to bridge the gap between conservation practices and digital innovation. Amid setbacks and organizational challenges, Novart emerges as a tool to streamline exhibition documentation and enhance museum operations, serving as a repository of institutional memory and information-sharing platform. Looking ahead, the narrative envisions a transition to academic research and a commitment to promoting interdisciplinary approaches in digital humanities. Proposing an innovative methodology for exhibition conservation, the thesis proposal reflects a deep commitment to preserving complex and ephemeral media for present and future generations. In essence, this journey encapsulates the spirit of innovation, collaboration, and perseverance essential for addressing the challenges of cultural heritage management in the digital age.