
Catalan crime fiction has grown rapidly in the last 20 years, becoming one of contemporary Catalan literature's most dynamic genres. This article examines its evolution from a geocritical perspective, focusing on the diversification of narrative spaces between 1954 and 2023. While Barcelona as a location has traditionally dominated Catalan noir, the genre has increasingly shifted towards other locations. Using quantitative and qualitative methods-including text mining, geoparsing, GIS visualization and close textual analysis-this study analyses 92 works to trace these shifts. Interrogating prior scholarship highlighting the role of Pep Coll's L'abominable crim de l'Alsina Graells (1999) in such a shift, this article demystifies its importance and identifies two major periods of decentralization instead: the 1980s, with Jaume Fuster and Maria Ant & ograve;nia Oliver highlighting the Pais Valenci & agrave; and Illes Balears, and the 2010s, marked by broader and spatially diversifying production by authors like Margarita Aritzeta, Teresa Solana and Jordi de Manuel. The study demonstrates GIS's value in complementing traditional literary analysis, showing how macro-level spatial trends refine geocritical theories. By integrating distant and close reading, it offers a methodological blueprint for future research, advancing understanding of Catalan noir's geographical evolution and the interplay of space, literature and culture.
This article presents a computational analysis of thematic and methodological developments in the field of digital humanities (DH) over the past two decades. Drawing on a corpus of 711 abstracts from Digital Humanities Quarterly (2007-25), the study employs text mining and topic modelling to map intellectual trends, research methods and conceptual priorities that have shaped the discipline's evolution. Using Latent Dirichlet Allocation (LDA) and co-occurrence analysis implemented in R, the research identifies 20 dominant topics, ranging from digital text analysis and pedagogy to cultural heritage and software infrastructure. The longitudinal distribution of these topics reveals both the persistence of core themes and the emergence of new methodological orientations, particularly the integration of machine learning, data visualization and multimodal research. The results indicate an increasing convergence between computational and interpretative approaches, reflecting DH's gradual transformation from tool-oriented experimentation to theory-informed analysis. By offering a reproducible framework for large-scale bibliometric text analysis, this study contributes to a more systematic understanding of how DH research agendas evolve across time and across modes of scholarly communication.
This article presents a bibliometric analysis of digital humanities (DH) research in China by examining its growth, thematic focus, institutional contributions and intellectual structures. Using 289 publications indexed from 2004 to 2025, data were analysed using Biblioshiny to assess authorship patterns, citation impact, collaboration networks and thematic clusters. The results show a rapid annual growth rate of 20.92 per cent, reflecting the field's expansion. Leading institutions include Peking University, Zhejiang University and Nanjing University, with scholars such as Jun Wan, Hao Wang and Xiaoguang Wang identified as highly influential. The keywords-digital humanities visualization and deep learning-highlight computational approaches and the preservation of cultural heritage. Although international collaboration is increasing, research is primarily driven by domestic partnerships. Key features include a concentration in elite universities, interdisciplinary integration, a domestically centred but internationally emerging collaboration network, and limited attention to ethical and methodological issues such as data privacy and censorship. This study provides one of the first systematic overviews of China's DH landscape and its growing significance.
We evaluate a Stanza-based pipeline for Old English that combines character-level processing with language-specific word embeddings derived from the Dictionary of Old English Corpus. On a 25,000-word dataset annotated with Universal Dependencies, character-level models yield consistent gains over token-based variants for lemmatization and dependency parsing (test UAS 88.92 per cent; LAS 79.65 per cent). Relative to a multilingual baseline for Old English, our best parser improves LAS by approximately 20 per cent; compared with a recent monolingual spaCy baseline on similar data sizes, gains are approximately 5 per cent on our split. We complement scores with qualitative analyses to illustrate strengths and limitations, and we provide implementation details to support replication. The main conclusion of this work is that, under low-resource conditions, language-specific embeddings and character-level modelling are more effective for Old English processing than cross-linguistic transfer learning.
Based on my opening keynote at DH2025, this article examines the impact that transformer-based machine learning brings to the interpretive work of historians. The discussion begins by revisiting an earlier phase of my research, which focused on digitally mediated, multiscale analysis-or what I refer to as 'digital (re)reading' - using database queries to detect structural patterns in digitized historical sources. Building on this foundation, I turn to the present, where my team and I investigate how large language models (LLMs) and vision-language models (VLMs) can support algorithmic reading across heterogeneous and semantically complex corpora. This new phase of inquiry explores how such models enable semantic, stylistic, sentiment and multimodal analysis, moving decisively beyond the constraints of keyword search and frequentist approaches. The article also introduces the DeepPast initiative, a modular artificial intelligence (AI) framework designed to promote the use of pluggable, task-specific components that run on low-power hardware rather than hyperscale, monolithic systems. The DeepPast architecture supports multiple interpretive modes, allowing historians to engage in a structured and purposeful dialogue with an AI assistant that functions as a genuine research partner. The discussion concludes with reflections on how to keep historians in the loop and to produce AI-assisted historical research marked by greater interpretive nuance and sophistication.
This article analyses the 'Following the Archives to View Shanghai' digital humanities platform, examining how it employs digital narratives to construct urban memory. Moving beyond a simple digital repository, the platform's innovation lies in its 'narrative intelligence' - its architectural design that integrates Geographic Information Systems, 3D modelling and knowledge graphs to create a dynamic, spatial-temporal-thematic narrative lattice. This structure transforms multimodal archival resources into an interactive 'possible world', enabling users to explore Shanghai's history through non-linear, associative pathways rather than a fixed chronology. The study finds that the platform successfully balances institutional curation with participatory engagement, fostering a user-driven and immersive exploration of the past. By synthesizing macro-historical perspectives with micro-narratives, the platform enhances public engagement with cultural heritage. It serves as a valuable blueprint for future digital archive initiatives, demonstrating that the strategic application of narrative theory and digital technologies can effectively bridge the gap between archival preservation and the dynamic construction of collective urban memory.
This research investigates the geographic distribution of the place origins and modes of travel of Blacks escaping slavery from the US South on the Philadelphia, Pennsylvania branch of the Underground Railroad between 1853 and 1861. We use digital data abstracted from nearly 1,000 narrative accounts of escape as originally recorded, and ultimately published as a book, by Black abolitionist William Still. We employ geographic information systems (GIS) to map the historic towns and counties from which individuals escaped to Philadelphia. Notable counties include Norfolk and Henrico Counties in Virginia (122 and 61 escapes, respectively), Dorchester and Baltimore Counties in Maryland (84 and 79 escapes, respectively), and Sussex County in Delaware (48 escapes). Travel by steamship and schooner were most common for escapes originating from the southern Chesapeake Bay. Travel by road and on foot were more common from origins nearer to Philadelphia. The total number of escapes increased from 1853 through 1857 before declining precipitously thereafter. Norfolk County emerged as a key origin of escape in the mid-1850s before rapidly declining, likely due to subsequent political organizing by enslavers. Dorchester County emerged as a key escape origin in 1857, likely due to the activities of Harriet Tubman.
This study investigates the conceptual relationship between the Arca Musarithmica - a seventeenth-century combinatorial music-generating apparatus devised by Athanasius Kircher (1602-80) - and contemporary artificial intelligence (AI)-based music generation systems. This apparatus is often recognized as a milestone in the partial mechanization of musical composition. However, comprehensive scholarly research examining its historical and conceptual significance as an antecedent to contemporary computational and AI-driven musical practices remains limited. The study investigates how the early use of fragmentation, mathematical representation, statistical quantification, randomization and specialization of music by Kircher aligns conceptually with core principles employed in modern AI music creation technologies. Through a systematic analysis of parallels and divergences between these historical and contemporary approaches, the study demonstrates that the apparatus constitutes a significant conceptual precursor to modern methods of computational creativity in music. Clear distinctions emerge regarding technological contexts, operational complexity and the degree of human intervention. This comparative historiographical analysis contributes to ongoing academic discourse on the origins and evolution of computational creativity, providing a nuanced historical perspective on contemporary AI music-generation practices.
Language models represent word meanings as vectors in a multidimensional space. Building on this property, this study offers a geometric perspective on parallelism in classical Chinese poetry, complementing traditional symbolic interpretations. To automatically detect parallelism in poetic verse, the authors trained a BERT-based classifier on a dataset of over 140,000 regulated poems (l & uuml;shi (sic)(sic)), achieving performance on par with state-of-the-art generative models such as GPT-4.1 and DeepSeek R1. Unlike general-purpose models, the custom classifier yields unique insights into how poetic meaning is encoded geometrically. The analysis shows that parallel lines exhibit alignment in the model's attention patterns: the 'key' vectors of corresponding characters point in the same direction, while this alignment disappears in non-parallel lines. This finding is interpreted through Peter Gardenfors's theory of cognitive semantics, which posits that humans make sense of the world by organizing experience into distinct conceptual regions. The authors argue that parallelism functions as a bridging mechanism that temporarily unites these disparate domains of meaning, suggesting a deeper, geometric order that underlies language itself.
Many digital humanities projects involve tedious and repetitive tasks that take time away from higher-level tasks further down the pipeline that require intelligent decision making. When funding is available, tedious and repetitive tasks are often assigned to research assistants, but when funding is scarce, those tasks tend to create bottlenecks that either impede progress or halt it altogether. This article argues that artificial intelligence and machine learning tools and techniques are worth exploring as cost-effective, accessible solutions to these problems. The Digital Latin Library project provides a case study through its experiments with fine-tuning pretrained transformer language models to process noisy bibliographic metadata. The results show that the models have potential for accelerating this tedious task, but that the experiments also had an unexpected, if positive, outcome: the models revealed gaps in the catalogue's coverage, helping to focus the efforts of the human experts working on the project.
This comparative study locates Frederick Douglass's (1818-95) religiosity within his published autobiographies as well as his 1886-7 unpublished travel diary-taken as part of the Library of Congress's Frederick Douglass Papers (1841-64).R-scripted text analysis of Douglass's expressed faith is foregrounded through the generation of clean digital transcriptions using handwritten text recognition (HTR), the process of converting non-consistent images into computer-readable text. In training a HTR model to 90.60 per cent character accuracy and manually correcting its outputs, this article forms a practical study in foregrounding automated transcription to benefit archival research and leverages multiple technologies to uncover lexical patterns across disparate autobiographical works. In showing that the abolitionist's introduction of humanistic themes holds little dependence to instances of religious expression, we challenge scholarly interpretations that Douglass committed to a hard pivot away from religious language and placed God 'off-the-stage' in his later works. In doing so, our approach demonstrates the range of techniques available to researchers auditing differences between private handwritten archives and public published accounts. This article therefore holds significance in documenting a scalable text extraction method on handwritten materials, digitized from smaller analogue library collections, while also advancing historical knowledge of Douglass's use of religious language.
Multimodal large language models (MLLMs) have sparked interest among scholars in digital humanities as they present a viable method for largescale analysis of multimodal data. For the study of multimodal communication, one potential application is corpus annotation, addressing the challenge of creating richly annotated corpora which demands significant time and expertise. This study explored GPT-4o, a state-of-the-art MLLM, for two foundational annotation tasks in building multimodal corpora of primary school science diagrams: (1) identifying individual modes of expression, such as written language, illustrations and arrows, and detecting their locations in the diagram layout; and (2) classifying the type of each diagram. Our quantitative evaluation against 1,000 human-annotated diagrams revealed that, while GPT-4o can identify the presence of expression, it cannot detect their location in the diagram layout. In classifying entire diagrams, GPT-4o can classify certain diagrams with moderate accuracy but struggles with others. Our results show that MLLMs cannot be used for annotating multimodal data as an out-of-the-box solution, due to their inability to perceive the structure of multimodal artefacts through pixel representations.
Corpus linguistics is an essential tool in digital humanities, and multilingual corpora are valuable resources in cross-linguistic studies. In this article we address the multilingual layout of the TenTen corpus family, questioning the rationale to call it a family, and advancing the idea of different degrees of kinship for its language members. The analysis focuses on the performance of the Sketch Engine Word Sketch tool in the English Web 2020 corpus (enTenTen20) in comparison with the latest release of the arTenTen, Arabic Web 2018 corpus (arTenTen18), which has been processed by CAMeL tools, an Arabic-specific software, and its previous version, the arTenTen12, tagged with Stanford CoreNLP. The study shows the challenges posed by the platform tools and the tagged corpora regarding the dissimilarities between the available data and the reliability of the results of these tools for both languages, as well as the efforts made to tackle the challenges. The concluding remarks point to the need for a better definition of multilingualism in the TenTen corpora and, by extension, in the digital humanities as a whole, based on the structural design of the resources and tools meant for such theoretical aspirations.
This article explores how geocomputational processes can be utilised to re-create historical spatial boundaries using Irish historical data. For this study, a series of spatial objects representing Irish Republican Army (IRA) brigades in County Limerick during 1921 were re-created in R using the Military Service Pension Collection (MSPC), one of the largest archival projects released by the Military Archives of Ireland. Using density equalising projections via cartograms, a series of analyses of the demography of the IRA was undertaken using these internal boundaries. County Limerick along the western seaboard of Ireland was chosen as a case study, as it represents a mixed urban/rural environment outside of the three major population centres on the island of Ireland: Cork, Dublin and Belfast. By deploying a geocomputational approach via the programming language R, this article provides a roadmap for historians and humanities scholars in general to reproduce these findings and spatial boundaries for future research.
Provenance research studies the origin and ownership history of objects including the conditions under which a change in ownership took place. In order to reconstruct the provenance of objects, provenance researchers utilise sources such as archival records, literature and online databases as well as examining the object itself. One essential online database for researchers is German Sales, which contains digitised auction catalogues from mostly German-speaking countries in the period from 1901 to 1945. Several catalogues contain handwritten annotations recording buyers, consignors and prices. However, in contrast to the printed text, the handwritten annotations are currently not searchable. To aid provenance researchers, it is imperative that the annotations be made machine-readable and searchable, which requires transcription, standardisation and enrichment of the handwritten notes. This article tests whether existing platforms, namely Transkribus, eScriptorium and ChatGPT, can be facilitated for the recognition and transcription of handwritten notes. While the experiments show the great potential of these methods, they also emphasise the need to train new models on these auction catalogues’ data.
This article draws upon recent developments in cognitive neuroscience and natural language processing to contribute a techno-cognitive perspective into the ‘deep reading’ versus ‘surface reading’ debate in literary studies. Research at the intersection of humanities and sciences suggests that narrative experience, including both production (decoding) and reception (encoding) of stories, constitutes a sequentially and hierarchically complex process shaped simultaneously by socio-cultural contexts, sensory-emotional dynamics, and cognitive integration across multiple levels of complexity. This interdisciplinary view contrasts with traditional humanities methodologies such as area studies, which privileges identity-based accounts of literary phenomena, or Marxist genealogy, which neglects extra-political sources of meaning. The article surveys relevant research findings across multiple domains and discusses the hermeneutic implications of the techno-cognitive approach for literary studies, exemplified in a reading of Zhang Xianliang's 1985 novel Half of Man is Woman.
This article examines whether authors change their styles when writing in different genres, like poetry and prose, particularly under the metrical constraints of the former. Using a corpus of works by five modern Egyptian authors, we analyse the effectiveness of N most frequent words and N most frequent morphological segments in distinguishing between poetry and prose written by the same author. Through supervised and unsupervised machine learning techniques, we demonstrate that authors exhibit distinct stylistic fingerprints in each genre, influenced by the unique conventions and constraints of each form. Results show that each author uses two different stylistic prints when writing prose and poetry. However, when mixing poetry and prose together, all genre-related texts cluster separately, potentially causing author obfuscation. Findings also show that the frequent word method is sufficient for accurately attributing authorship when it comes to mixed-genre texts. In short, the tested linguistic standard features prove resilient across different genres and even survive the constraints of formal poetic meter.
Inventar Colombia: una mirada desde el Orinoco is a digital research project that seeks to reconstruct aspects of the geographic imagination, forms of relationship with nature and its resources, of at least 15 Indigenous groups that inhabited the Orinoco region between the fifteenth and eighteenth centuries. Responding to the limits to what we can know in the archive about the ways in which various Indigenous groups inhabited and conceived the territory, we carried out a process of curation and extraction of geospatial data in historical and archaeological sources. We identified some patterns of territorial occupation and exchange networks of the multiplicity of Indigenous communities that have inhabited the Orinoco region. From iterations for the construction of maps with QGIS and typical GIS conventions, we explored other forms of visualization and narrative construction to deploy the arguments of the research on the web. We proposed a CSM-free web interface developed with JavaScript and open-source libraries and developed a set of cartographic narratives.