
This study examines audience responses to “Proteggi I Fatti”, an Italian Facebook campaign developed within the European project Deconstruct: Deconstructing and Countering Holocaust Distortion via Campaign, Education and Training. The campaign aimed to raise awareness of Holocaust distortion and denial by disseminating historically accurate information and supporting critical engagement with online content. The analysis is based on 10,160 public comments collected from the comment sections of nine Facebook posts published between December 2024 and May 2025. Methodologically, the study combines lexical and bigram analysis, sentiment analysis, word embeddings, and LLM-assisted classification through the Gemini API, followed by manual validation. Findings show that many users reframed the campaign’s educational messages through the lens of the Israeli-Palestinian conflict, especially the war in Gaza. Recurring patterns include Holocaust distortion, Nazi analogies, victim-perpetrator reversal, and accusations that Holocaust memory is instrumentalized for political purposes. The paper contributes to research on digital Holocaust education by showing how social media campaigns may become sites where remembrance, political conflict, antisemitism, and distortion intersect. It also highlights the value and limits of AI-assisted methods for identifying implicit and context-dependent forms of Holocaust distortion in platformed publics.
Linked Open Data (LOD) is widely promoted as the next step for digital humanities projects, yet our experience directing FLAME (Framing the Late Antique and early Medieval Economy) counsels a tempered, pragmatic view of its potentialities and drawbacks. Drawing on five years of work to align FLAME’s coin-find database with LOD ecosystems, we present several case studies on difficulties we have encountered in (1) reconciling typologies with the main initiative for creating a shared ontology in our field; (2) importing external datasets; and (3) integrating a large but narrow dataset. We conclude that LOD delivers clear gains when robust, centralized shared ontologies exist, but we encountered persistent barriers: disciplinary path-dependence; gaps and asymmetries in vocabularies; missing cross-project identifiers for finds; inconsistent metadata and coverage; duplication across platforms; and the amplification of geographic and methodological biases that users rarely read about but readily visualize. Resource constraints and distributed governance further slow coordination and maintenance. We recommend adopting or co-developing standards at project inception; designating liaisons; instituting and coordinating unique identifiers for coin finds; combining automated duplicate detection with manual review; and prominently highlighting provenance, methods, and biases. LOD remains desirable and practicable, but only under sustained, field-level coordination rather than assumed universality.
Repetition, with and without variation, pervades language use. We see it, for example, in spontaneous conversation, in memes on social media, in traditional rhetoric and, of course, in poetry and oral tradition. This paper presents recurrence plotting as a visualisation tool for the analysis of exact and inexact repetition in language. It is a particularly powerful tool for the exploration of un(der)studied data, such as those arising through language documentation work, as its reliance on the visual sense offers instant, holistic insights into the structures of large amounts of data. Recurrence plotting can be used with any type of connected source material, including audio (e.g., full-spectra, f0 contours, etc.) and written data (e.g., transcripts, glosses and translations). We illustrate in this paper the diverse potential of recurrence plotting with data from two oral traditions: the Old Indo-Aryan Rigveda, and the oral tradition Igu of the eastern Himalayan society of the Kera’a.
The extent to which codes of ethics (and ethics more generally) are influenced by social, political and cultural factors have been discussed in the literature for some time. We can find several direct examples of this in codes themselves, with some explicit about the need to interpret their content in light of broader social and political forces. Despite these discussions however, we have not yet been able to quantify these differences or the factors that influence codes. Utilising natural language processing, this study sought to explore this, examining whether similarities between codes of ethics were related to a range of socio-political factors. Data for this study came from 202 medical and nursing codes of ethics from around the world. To explore associations, a number of common metrics were utilised, including macro-economic, political, cultural and healthcare related variables. To compare each document, cosine similarity was calculated (on the whole document text and TF-IDF) providing two measures of similarity between document pairs. The other variables included for modelling were also transformed into pairs, categorising them or calculating differences between scores. Generalised linear mixed models revealed that sharing a profession (medical or nursing) and language had the strongest associations with code similarity. Colonial history (i.e. countries with a coloniser-colonised relationship) and greater similarities between GDP and individualism/collectivism were also associated with greater similarity. The pattern when it came to political and healthcare system variables was less clear, showing no clear and consistent association across models. We discuss these findings in light of the broader literature and theoretical debates within bioethics.
This article examines how artificial intelligence (AI) challenges and enhances editorial practice in the long-term Digital Scholarly Edition of Grundtvig’s Works (2010–2030). Focusing on named entity recognition, semantic annotation, and biblical reference detection, the authors explore how computational methods – particularly large language models (LLMs) – are evaluated and applied within a Human-in-the-Loop (HITL) framework. The article argues that while such tools can ease repetitive tasks and support editorial workflows, they must be aligned with the interpretive, philological, and historical foundations of scholarly editing. Through concrete examples from annotation practices and manuscript transcription, the authors show how editorial work remains a critical interpretive act. The article also outlines efforts to integrate AI-assisted methods into the digitisation and enrichment of Grundtvig’s extensive manuscript archive. Rather than automation replacing expertise, it proposes a model of methodological convergence: where digital tools extend the possibilities of human judgment in scholarly editions.
This study applies computational methods for measuring semantic distances to born-digital drafts. Using text versions leading up to a short story by Flemish author Ellen Van Pelt, we are looking for relevant entry points into a born-digital genetic dossier. Several methods are applied and compared to consecutive pairs of drafts; a count of unique lemmas entering and leaving the working document, a document cosine similarity measure based on a BERT model of Dutch, and a narrativity measure. By comparing these different methods with close reading, we can assess whether such computational tools may be of help to textual scholars working with big corpora of digital draft material. The working process of Van Pelt could be divided into several stages with a different focus. The methods each partially picked up on these stages, and also highlighted specific drafts. It did not become clear which of the methods was most successful, in part because of the small corpus used, for a writing process that was characterised by gradual expansion and continuous revision.
The paper presents a comprehensive review of the research on computational methods of style change detection methods for intrinsic plagiarism detection. We systematically evaluate and summarize 120 research papers published between 2006 and 2024. We categorize the various methods of style change detection and examine papers devoted to related problems, such as style breach detection and author diarization. By compiling and analyzing information from the articles, we provide statistics on the key aspects of the solution pipeline. The assessment of datasets used in the articles, their languages and employed methods reveals trends in the field of intrinsic plagiarism detection along with changes in language and topic diversity. By using manually labeled keywords for the articles, our review also provides the list of most frequently used features and methods. We demonstrate that intrinsic plagiarism detection and style change detection are active research fields.
This paper critically engages with Mohammad Ayish’s (2025) meta-analysis of 1,226 publications and 88 curricula in Arab media studies to argue that the field’s digital turn necessitates not only methodological innovation but also an epistemological reconceptualization grounded in postcolonial and decolonial thought. While affirming Ayish’s findings of a paradigm shift toward digital topics (54.41
Hybrid technologies enable the blending of physical and digital elements, creating new ways to experience and interact with the world. Such technologies can transform engagement with relics—both secular and sacred—but they present challenges for capturing faith, belief, and representation responsibly. Given the complexities of digital representation and the ethical challenges inherent in digitising culturally significant objects, a transdisciplinary understanding of these issues is needed. To inform this discussion from a linguistic perspective, we examined the representation of relics in historical and contemporary texts. Using a corpus linguistic approach to extract modifiers of the word ‘relic’ in corpora of Early Modern English books and contemporary web-sourced texts from 2021, we examined the multifaceted ways in which relics have been perceived and evaluated over time. Early texts consider relics as both objects of moral and spiritual significance, and tools of religious and political control, while they are more often framed as heritage symbols, reflecting past events, places, and traditions in contemporary texts. We discuss how hybrid, sometimes AI-based technologies can enhance accessibility and engagement, whilst also challenging traditional sensitivities around authenticity and sensory experience, which are integral to the meaning and significance of relics.
This article presents a methodological framework for representing historical diaries as multi-layered, interactive, region-based RDF knowledge graphs, enabling precise and granular modeling of manuscripts, pages, textual regions, and annotations. It identifies five significant challenges in scholarly digital editions and demonstrates how these can be addressed through a framework that supports segment-level text representation, targeted annotation, and interoperability across editions, languages, and manuscript witnesses. This framework incorporates a segment-level citation system with resolvable, timestamped identifiers, ensuring reproducibility, stable referencing, and reliable scholarly citation over time. To describe the suggested methodologies and functionalities, this paper provides examples from digital editions of Jacob Bernoulli’s Meditationes and Reisbüchlein, highlighting the potential of RDF ontologies to capture complex editorial layers, including transcriptions, translations, and scholarly commentary.
In the Ghost Neighborhoods of Columbus project, we are developing and applying machine learning (ML) and geographic information science (GIS) methods to extract data from historical Sanborn fire insurance maps and build 3D urban models of how neighborhoods looked in the past. We are focusing on historically Black neighborhoods in Columbus that have been altered by urban highway construction, urban renewal and redlining practices. We are working with neighborhood members and community partners to identify neighborhoods for investigation, model use cases, design features and delivery modes. We are also collecting stories, memories and photos with the intent of using the 3D urban models as platforms for storytelling about lived experiences in these lost places. In this paper, we describe our approaches to 3D model development, model delivery and community engagement. We identify and discuss issues and bottlenecks we discovered, and strategies for resolving these problems.
The United Kingdom’s Maritime Heritage records hold vital information on heritage assets and archaeological contexts key to addressing current research priorities. These records stand alone, rarely connecting, demanding an arduous and time-consuming process of investigation and cross-referencing for synthesis to occur. Challenges intensify due to the qualities of this data; they stem partly from legacy information not initially intended for heritage recording purposes. Complications also arise from the varied methods of recording, storing, and disseminating this information deployed over time. However, archaeologists now possess Named Entity Recognition (NER) and Natural Language Processing (NLP) tools with the potential to organize and search such fragmented databases. This paper explores a method to enhance and enrich these datasets by employing Few Shot learning techniques to perform Multiclass Text Classification. The Welsh National Monuments Record (NMR), a glimpse into the U.K.’s maritime heritage, has been utilized to test these approaches.
The growing availability of digitized historical collections has enabled large-scale computational research; however, transforming heterogeneous and noisy textual data into structured and analyzable formats remains a major challenge within the Digital Humanities. This article presents a reproducible workflow for historical text processing that integrates Optical Character Recognition (OCR), Large Language Models (LLM)-assisted post-correction, and Named Entity Recognition (NER) into a unified pipeline. Implemented within the n8n automation framework, the workflow emphasizes transparency, modularity, and human-in-the-loop validation, enabling scholars to maintain interpretive control over data transformation. The pipeline is evaluated on five historical corpora in Spanish (eighteenth and nineteenth centuries), demonstrating significant reductions in Character and Word Error Rates (up to -93.5
This paper explores the use of large language models (LLMs) to enhance semantic annotation and annotation projection in digital scholarly editions (DSEs), focusing on historical ego-documents. Using the TEI/XML-encoded French memoirs of Countess Luise Charlotte of Schwerin (1684–1732) and their German translation as a case study, we evaluate LLMs for Named Entity Recognition (NER), annotation transfer across aligned bilingual texts, and the extraction of interpersonal relationships. A comparative analysis with a traditional NER framework shows that LLMs significantly outperform baseline models, particularly in recognizing complex person references, such as non-rigid designators and nested entities. For annotation projection, we demonstrate that LLMs can reliably transfer entity annotations between French and German texts without intermediate alignment layers, achieving over 97
As soon as they are in power, the US president must explain and justify their actions through speeches, the most important of which are the inaugural, State of the Union (SOTU), and farewell addresses. It has been admitted that these three text genres present distinct characteristics, and thus could form different categories. Based on the stylometric analysis of discourses made by Reagan, Clinton, G. W. Bush, Obama, Trump, and Biden, this study reveals the distinctive style and rhetoric of these types of speeches. Clearly, two well-defined speech genres appear: one regrouping the inaugurals, the other the SOTUs. Compared to the speeches of previous presidents, those of Biden and Trump have a more conversational tone and focus more on moral values in the former and on glorious terms (e.g., hero, incredible) for the latter.
In the digital age, the role of transcription in the editorial process is undergoing profound change. Thanks to AI and Automated Text Recognition (ATR), the transcription step has been drastically accelerated, multiplying the amount of text accessible to medievalists, reshaping our relationship to sources, and opening new possibilities for textual availability and analysis. Because transcriptions are decisive and often determine the entire digital pipeline of text acquisition and enrichment, they can no longer be regarded as a merely preliminary, mechanical step preceding the ”real” work of editing. Instead, transcription emerges as a locus of interpretation, where machine output, scholarly modeling, and editorial judgment converge. This article argues for a concept of stratified editing, in which successive layers—raw ATR output, corrected transcriptions, normalized forms, and enriched annotations—coexist and retain distinct epistemic value. Rather than being effaced beneath the final edition, each layer becomes reusable, supporting diverse research goals from computational analysis to philological study. Such stratification requires solid modeling of transcription at the intersection of scientific objectives, pragmatic practices, and technical constraints. Ultimately, transcription should be recognized as a scholarly object in its own right, at the heart of a new editorial ecology that reshapes how we edit, interpret, and engage with medieval texts.
The development of word frequencies over time is the subject of research in different branches of the humanities. Large temporal n-gram corpora have been created for this purpose, most notably the Google Books Ngram Corpus. While the concrete research questions vary between the different research works, there are similarities in the more abstract underlying information requirements, i.e., the structure of queries against a potential database system. Based on a systematic literature review, we extract these information requirements, leading to a categorization of existing articles into macro-areas of information requirements. Furthermore, we collect existing query systems for temporal n-gram corpora and evaluate their expressiveness regarding the information requirements we found.
Transitioning from research proposal to project management in Automatic Text Recognition (ATR) for cultural heritage text collections demands meticulous planning, realistic budgeting, and efficient team coordination. Insights from the Transkribus User Conference 2024 offer a roadmap for success, emphasising clear objectives, ethical considerations, and practical outcomes. Effective budgeting should include digitization, software, staffing, and long-term storage, while a comprehensive Data Management Plan (DMP) ensures efficient data handling. Key project management tasks include team formation, workflow documentation, and document layout optimization. High-quality training and validation data in ATR require clear guidelines, iterative testing, and agreed licensing. Crowdsourcing can enhance ATR projects, and sharing models and datasets through platforms like Transkribus Sites promotes collaboration and visibility.
Accelerating developments in information digitalization have induced a growing interest in harnessing digital tools and processes to investigate subjects traditionally classified as humanities. The advent of Digital Humanities (DH) has entailed not only a shift in research methodologies and approaches in disciplines such as history, literature, cultural studies, arts, and media, but it has also ignited debates about the ability of this scholarly movement to generate mature insights about the human essence of our lived experience. This article provides a meta-analytical analysis of how Digital Humanities are manifested in media studies in the Arab World from 2002 to 2022. The article addresses three facets of Digital Humanities in Arab media studies: research topics, methods, and media education curricula. The study concludes that while Digital Humanities in Arab media studies has not received adequate attention in academic discussions, a strong scholarly interest in the emerging digital landscapes and cyber communications suggests the region is gradually embracing DH features in both published research and media education curricula.