
Abstract This study aims to investigate the evolutionary patterns of the Panhu myth among the She people in southeastern China and to explore how mythological variation can provide insights into historical migration and cultural transmission. The research adopts a phylogenetic approach informed by digital humanities and cultural computation. A dataset of thirty-six homologous variants of the Panhu myth is analyzed to reconstruct their lineage relationships. Methods of distant reading and narrative data mining are employed to identify structural patterns and evolutionary trajectories across variants. The analysis identifies three major mythological systems centered in southeastern Guangdong, the Fujian–Guangdong corridor, and the Fujian–Zhejiang corridor. The latter two emerged later and exhibit further internal differentiation. Despite this divergence, continuous interaction among the three systems is evident. Combined with historical and ethnological evidence, the findings suggest that since the Song dynasty, the primary migration route of the She people followed a trajectory from Guangdong through Fujian to Zhejiang, accompanied by sustained cultural exchange. This study integrates phylogenetic modeling with mythological analysis, demonstrating the applicability of computational methods to the study of narrative traditions. It moves beyond traditional qualitative approaches by quantitatively reconstructing myth evolution and linking it to population movement. By applying phylogenetic analysis to mythological corpora, this research highlights the potential of digital humanities methods for uncovering large-scale cultural patterns. It provides a methodological framework for using narrative data to infer migration histories, offering new perspectives for the study of early populations.
Abstract Eugene O’Neill’s obsession with death is well-known. Many critics have written about the death motif in his plays. However, most studies are based on intuitive knowledge and are qualitative in nature. With a self-developed corpus, a quantitative study of the representation of the death motif in the play The Iceman Cometh is conducted, in combination with turn-taking and semantic domain analyses. It not only maps out the trajectory of the development of the death motif in the play but also confirms the correlation between the Iceman and death. This method opens up a new avenue for digital humanities in O’Neill criticism.
Abstract When debunking Serpa Pinto’s exaggerations about the perils encountered in Africa, Hermenegildo Capelo and Roberto Ivens strongly stated he had been accompanied by ‘the best men of the expedition’. But Serpa Pinto was not the only one who had company. An analysis of the two volumes of Capelo and Ivens’ From Benguella to the territory of Yacca reveals the duo was assisted by a vast network of local agents that included guides, porters, translators, and many other natives who contributed to the construction of scientific knowledge and the formation of the collections gathered during their journey. In this article, we use the network visualization software Gephi to map the relationships between the travellers and those mentioned in the travel book to demonstrate the importance of sociability to nineteenth-century scientific expeditions. While kept invisible for most of the historiography on these expeditions, the contributions of local agents have recently been the focus of work centred on circulation. With dedicated software, we can now visualize these networks, bringing into light the social nature of fieldwork.
Abstract This article examines historical relations between gay and transgender communities by analyzing hyperlink structures among Dutch LGBT websites archived in 2020. While the acronym LGBT suggests a cohesive “rainbow coalition,” historical relations between gay and transgender groups have been marked by tensions and uneven forms of integration. Drawing on a corpus of 192 websites archived by the Dutch National Library, this study uses hyperlink network analysis to investigate patterns of affiliation, boundary maintenance, and cross-identity engagement at scale. Hyperlinks are treated as traces of social and organizational orientation, making it possible to reconstruct relational structures within and between online communities. A combination of cluster analysis and in- and out-linking analysis yields three main findings. First, a distinct and internally cohesive transgender cluster emerges alongside a less cohesive but still pronounced gay cluster, complicating narratives of a post-gay era. Second, transgender websites display a significantly stronger internal orientation, receiving a higher proportion of in-links from within their own community, reflecting the enduringly different needs and experiences of trans people. Third, cross-linking across identities was sporadic, suggesting the persistence of relatively stable identity-based orientations rather than an integrated LGBT web sphere. Methodologically, the article demonstrates the value of archived hyperlink networks for historical research on digital, and particularly marginalized, communities. Substantively, it contributes to scholarship on LGBT history by providing the first systematic, web-historical analysis of gay–transgender relations. More broadly, it shows how network-based approaches can help digital humanities scholars study community formation, separation, and coalition-building in historical web environments.
Abstract This article introduces interpretive computing, a humanistic design framework for artificial intelligence systems that support collaborative meaning-making, narrative construction, and cultural responsiveness in archival contexts. Moving beyond explainability and automation, interpretive computing positions AI as a scaffold for situated interpretation—one that centres context, dialogue, and plural epistemologies. We develop this concept through a case study involving the Smithsonian Institution and U.S. Library of Congress Civil Rights History Project—oral history collection, where large language models are used to generate thematic metadata, segment interviews, and construct dynamic, user-curated story paths. Our prototype interface features include editable metadata annotations and thematic playlists, allowing users—whether educators, students, or community members—to engage in co-authorship of historical understanding. Grounded in traditions from hermeneutics, narratology, and digital humanities, this work reimagines metadata and interface design as interpretive acts, offering a model for AI systems that support dialogic, reflexive, and culturally situated engagements with the past.
This study examines how traditional Chinese painting (guohua) is framed and diffused through two major Chinese short-form video platforms: Douyin and Red Note. Drawing on framing theory and diffusion of innovations theory, the research examines how narrative and visual framing strategies shape user engagement and content diffusion in digital heritage communication. A dataset of 1,642 videos was collected and coded, and statistical analyses were conducted to test four hypotheses related to framing type, engagement, diffusion, and platform moderation. Results reveal that Douyin predominantly uses entertainment-oriented frames, while Red Note emphasizes educational and esthetic framings. In addition, educational and culturally symbolic content generates significantly higher levels of engagement and broader diffusion than purely performative content, with these effects being especially pronounced on Red Note. These findings highlight the role of platform-specific affordances in shaping the communication of traditional art and provide practical insights for creators and cultural institutions aiming to balance authenticity with algorithmic visibility.
Digital archives function as critical infrastructures for the preservation, circulation, and reinterpretation of Shakespearean texts and performances. Despite substantial practical development, the interdisciplinary research landscape of Shakespeare digital archiving has not yet been systematically synthesized. A Preferred Reporting Items for Systematic reviews and Meta-Analyses-guided systematic review is therefore conducted, drawing on thirty-two studies published between 1990 and 2025 in the Web of Science and Scopus databases. By integrating cross-analysis of thematic and methodological development, this review provides the first structured synthesis of the field's evolution over thirty-five years. Although participation from non-Western regions has increased, institutional leadership and scholarly authority remain largely concentrated within Western academic contexts. Research themes have broadened from early technical implementation to include pedagogy and interculturality, but conventional infrastructural and editorial standards continue to shape practice. Methodologically, qualitative case-based inquiry persists as the dominant paradigm, while quantitative and user-centered approaches remain comparatively marginal. This review positions Shakespeare digital archiving as a critical site where technological standards and institutional authority intersect, contributing to broader debates in digital humanities regarding openness, participation, and democratization.
The relationship between local elite groups and regional cultural evolution across temporal and spatial dimensions has long been critical for understanding how civilizations change. This study employed interdisciplinary methodologies from digital humanities and historical geography to conduct an empirical analysis of 32,207 biographies recorded in local chronicles from the Zhejiang region, spanning the Eastern Zhou Dynasty to the modern era (i.e., 770 BCE-1945 CE). The social composition, achievements, spatial concentration, and diffusion patterns of local elites within the context of different historical periods were of particular interest, alongside their interactions with regional cultural development. It was found that the diversification of local elite types in the Zhejiang region was pronounced, with distinct temporal characteristics, encompassing roles in science and education, politics, literature and art, and medicine. Elites in the field of science and education had the highest probability distribution, accounting for 30.8 per cent, while those in medicine and the military had the lowest, at 13.6 per cent and 12.2 per cent, respectively. The criteria for selecting and recording elite figures changed over time but consistently reflected the trajectory of alterations in mainstream societal values. The importance of cultural heritage and regional characteristics was particularly evident for literature and the arts but was less apparent in science, technology, and education, perhaps due to rapid technological progress and frequent shifts in policy orientation.
Amateur book reviews published on Digital Social Reading (DSR) platforms play a crucial role in sharing reading experiences and capturing reader preferences. However, little attention has been given to the linguistic features that influence the way reading experiences are conveyed and shape readers' perceptions of the aspects discussed in reviews. This study addresses this gap by combining Computational Stylometry and Machine Learning techniques to examine how the linguistic features of book reviews combined with the demographic characteristics of review readers impact the reception of book reviews, focusing in particular on three key aspects that might make a book review compelling, namely informativeness, writing style, and enjoyment. To this aim, we relied on a corpus of Italian Goodreads reviews to investigate how amateur reviewers communicate their reading experience and on a survey to collect human judgments about review reception. Additionally, we investigated the extent to which the review style and the demographic characteristics of review readers can predict the different aspects of review perception. Our findings revealed that linguistic characteristics play a crucial role in shaping reader perceptions and are equally or even more predictive than demographic information in automatic classification models. Nevertheless, while demographic factors such as gender and birth year offer limited utility in forming homogeneous reader groups, reading habits emerged as a relevant factor in identifying shared trends among readers. These insights contribute to a deeper understanding of reader involvement in DSR communities and offer valuable insights for both publishers and review platforms in defining book recommender systems.
This paper addresses digital humanities and digital humanism by defining, characterizing, and distinguishing them, while also making a critical reflection about the limits of the digital humanities regarding the analogical humanities and quantum humanities. Also, as digital humanities transcend humanities and are practiced by social scientists, we characterize them from the Science and Technology Studies and based on its practice as technosciences. This seek to resolve different discussions that have taken place and those that might arise regarding digital humanities, which must be aware of fashionable nonsenses, which can occur in interdisciplinary projects.
Multimodal deep learning (MDL) has potential for digital history (DH) research, from simple tasks such as image or text classification, to more sophisticated ones such as information retrieval (IR). We propose a novel IR method based on vision language models (VLMs), a class of MDL models, to extract embeddings from images and text to find connections among unrelated image and text datasets that are challenging to find using conventional approaches such as keyword search. We evaluate our method with a proof-of-concept case study on debates surrounding the civil usage of nuclear power in Europe, based on poster images from the Landelijk Kernenergie Archief website and German newspaper articles found in the Impresso corpus, where both datasets are assumed to share some common themes. We find that poster images and newspaper articles can be automatically matched from independent sources based on representations extracted by pre-trained VLMs. We further find that VLMs achieve better performance when translating newspaper articles to English or summarizing them in English. Ultimately, our work bridges the gap of cross-modal and cross-lingual IR for DH, showing real-world applicability of MDL and opening new avenues in DH research. Our code and data are available at https://github.com/StevenXuf/MatchMeIfYouCan.
This study aims to systematically examine discourse polarization on Chinese social media through an integrated cognitive and computational lens. It proposes a three-dimensional semantic-structural-strategic framework to capture how linguistic meaning, network structure, and discursive strategy interact to produce polarization in digital discourse. Two highly contentious cases, the Delayed Retirement policy and Yang Li's endorsement of JD.com, are analyzed across Weibo and Little Red Book. The study focuses on posts containing adversarial language, combining BERTopic clustering, semantic network construction, Louvain community detection, and critical discourse analysis. Structural indicators quantify polarization, while cohesion and fragmentation are assessed through community purity, intra-/inter-community similarity, and cross-community edges. Ten discourse strategies are coded, and sentiment analysis reveals affective feedback mechanisms. Results show that Weibo functions as a mass polarization accelerator, characterized by rapid isolate-integrate cycles and high emotional contagion, whereas Little Red Book exhibits low-intensity, niche polarization grounded in personal narratives. The findings demonstrate a feedback loop between extreme negative emotions and intra-community interaction, reinforcing group boundaries and antagonistic framing. This study develops a novel multidimensional framework that bridges micro-level discursive cognition and macro-level network structure, advancing the methodological integration of computational linguistics and cognitive discourse analysis. By operationalizing semantic, structural, and affective dimensions of online discourse, this research expands digital humanities approaches to include networked cognition and emotion-driven polarization. It provides empirical and conceptual tools for understanding collective meaning-making in digital environments and offers actionable insights for mitigating online hostility through platform-level discourse design.
This study analyzes the symbolic recurrence and thematic plasticity in Federico Garc & iacute;a Lorca's poetry through an interdisciplinary approach combining literary analysis and natural language processing. An extensive corpus-19 poetry collections, 390 poems, and nearly 9,000 verses-was analyzed using lexical metrics, frequency analysis, topic modeling, and LLM-assisted semantic interpretation to identify characteristic isotopies. The findings reveal sustained lexical richness and four major thematic axes that, while persistent, shift in prominence across Lorca's creative periods, reflecting both aesthetic coherence and expressive reinvention. Notably, significant semantic plasticity emerges at the verse level, especially in youthful and traditional phases, where verses flow between multiple thematic spaces. The presence of frontier verses-marked by thematic ambiguity and polyvalence-offers a quantifiable indicator of semantic density and stylistic evolution, and reveals the poet's capacity to activate multiple symbolic resonances. This research demonstrates how digital humanities and artificial intelligence can illuminate complex artistic phenomena, offering quantitative support to traditional philological interpretations of Lorca's imagery.
This study investigates how human participants evaluate the human likeness of Chat Generative Pretrained Transformer (ChatGPT)-generated texts compared to human-authored texts in terms of register alignment and functional appropriateness. It involved eighty-three participants who evaluated four texts (two by ChatGPT and two by humans) representing university textbook and blog registers. Participants answered questions about text authorship, register identification and alignment, and purpose identification and functional appropriateness. Findings showed that participants struggled with identifying ChatGPT-generated texts accurately, with only one ChatGPT text correctly identified by the majority, while they performed better in identifying human-authored texts. Participants were also challenged in identifying the correct register, with no notable differences between human-authored and ChatGPT-generated texts. However, participants generally rated all texts as aligning with their chosen registers, with texts identified as human-authored perceived as more aligned. Moreover, although participants had difficulty distinguishing between human and artificial intelligence (AI)-generated texts and identifying registers, they could successfully recognize the communicative purposes of texts regardless of authorship, which suggests that ChatGPT can convey the communicative purpose of a text. Nonetheless, texts identified as human-authored were perceived as more functionally appropriate. This study provides insights into the current state of AI-generated text quality. The findings suggest caution when considering the integration of ChatGPT-generated texts into language education. While ChatGPT shows potential in providing language practice and development, the perceived lower quality in terms of register alignment and functional appropriateness may affect its effectiveness. Educators should critically evaluate the use of AI-generated texts to ensure they meet the linguistic standards required for educational purposes.