
This review analyzes Fabrizio Venerandi’s Manuale di letteratura elettronica (2024), a work that aims to introduce a non-specialist audience to the world of born-digital literature, with a particular focus on text adventures and narrative video games. The volume highlights the centrality of video games as "new literature" and their ability to redefine the traditional categories of author, reader and text, combining a historical perspective with an eminently practical and educational approach. Through an extensive mapping of works and tools, the manual emphasizes its value as a catalogue raisonné and as an invitation to recognize video games as one of the main contemporary laboratories of electronic literature, raising questions of preservation, authorship, and critical use of digital medium that are central to the Digital Humanities as well.
The Voynich Manuscript (VM) is a medieval manuscript likely written in the 15th century (Yale Univ., Beinecke Rare Book & Manuscript Library MS 408).[1] The manuscript is written in an unknown language or code using an unidentified set of symbols that has yet to be made legible. Additionally, the codex contains many strange and fantastical images of plants, people, and cosmological/zodiac illustrations, the meaning of which are also unknown. One of the main research avenues into the VM is to examine its textual content to understand how it behaves relative to known texts; this can provide insight as to whether the mysterious writings contain decipherable text or not. In this paper, we explore the coherence and flow of the manuscript using Latent Semantic Analysis (LSA). LSA is a technique that may help ascertain whether the behavior of the text within the VM shows evidence of a coherent flow of topical content, by comparative analysis of text samples that are near each other, farther away from each other, at section breaks, or even page breaks. The advantage of this strategy is that LSA analysis can be undertaken without actually knowing the meaning of the text. We expect portions of text that are near to each other to have a relatively high similarity score, that is, to be potentially semantically related. We also expect that at anticipated topic breaks (pages or sections), the similarity score between adjacent text blocks would be smaller, as the breaks seem to represent a change in topic. Both of these patterns are observed in the control manuscript studied as proof-of-concept experiments. Patterns then observed in several sections of the VM suggest that there may be an overall coherence to the text.
In this essay we explore the relationship between code and the physical world through an in-depth analysis of the analog video game Volleyball, one of the first games for the first video game console ever produced, the Magnavox Odyssey, in 1972. We address the relative lack of critical humanities scholarship on analog programming by exploring the many ways that this code differs from digital program code, and the implications of the relations it establishes between user, digital logic, and the surrounding environment. We represent this program in four levels of abstraction: as hardware (circuitry), as mathematical models, as block program notation (historical analog programming notation) revealing the program’s logic flow, and as game, when combined with player input and human language instructions. We conclude that Volleyball is an example of code that enrolls its end users as fundamental constituents in its primary program loop, and thereby “plays in the gap” between the symbolic and the real, the analog and the digital, the inside and outside of code. We explore this model as multi-level code, arguing that it suggests new perspectives and modes of analysis for digital code as well.
This study examines stylometric consistency patterns in professional human translations of spoken Chinese to written English, using the CoVoST corpus. Moving beyond traditional translation quality assessment, the investigation studies how human translators navigate the speech-to-text modality shift and whether this process produces domain-invariant stylistic regularities. Through a framework combining stylometric analysis, machine learning, and digital humanities critique, I identify a "consistency bias" - manifested as standardized lexical diversity, flattened syntactic structures, and repetitive discourse patterns - that reveals how professional constraints and cognitive processing shape translation output. My analysis of 16,899 human-translated English sentences, with detailed statistical comparison of 300 sentences against contemporary spoken English baselines, demonstrates that speech-to-text translation exhibits significantly reduced lexical diversity (Cohen's d=0.34, p<0.001), shallower syntactic structures (d=0.33, p<0.001), and narrower modal verb usage (d=0.43, p<0.001) compared to original spoken English. These findings illuminate the cognitive and professional constraints shaping human translation practice; future research might consider how such patterns subsequently inform machine translation systems trained on human-translated corpora.
Code is an epistemic system predicated on the repression of state, but with the rise of global optimization and machine learning algorithms, code functions just as much to obscure knowledge as to reveal it. Code is constructed in response to two characteristics of the twentieth century episteme. First, knowledge is represented as a process. Second, this representation must be sufficient, such that its meaning is constituted by the representational form itself. In attempting to meet these requirements, process is separated into an essential part, code, and an inessential part, state. Although code has a relationship with state, in order to construct code as an epistemic object, state is limited and suppressed. This construction begins with the first formation of code in the 1940s and reaches its modern form in the structured programming movement of the later 1960s. But now, with the increasing dominance of global optimization and machine learning algorithms in computing, it has become apparent that state is vitally important, and yet our tools for understanding state are inadequate. This epistemic inadequacy nevertheless serves those who would act dangerously and shun responsibility for the consequences.
The rhetorical significance of naming practices is widely understood, but it — and many other rhetorical dimensions of language — are often overlooked in the domain of software development, especially in regards to code languages and relevant practices (as demonstrated in file names, functions, variables, and so on). While naming conventions in code are typically recognized as inherently arbitrary, they are also tangled up in numerous networks of community expectations, constraints, and mores, whether organizational or interpersonally social in nature. Given Kenneth Burke's argument for the revealing and concealing influences of terministic screens upon our engagement with the world (by establishing ways of seeing and not seeing), naming conventions in code play an important role in how meaningful invention occurs for human developers and readers of code files. Despite the apparent triviality of such a component of software projects, naming practices shine a light on the goals and values of a programmer in addition to the functional intentions that they might have for the use of a given body of code.
Multilingualism has been gaining importance in the digital humanities and scholarly communication, but the infrastructure used to disseminate scholarship has mostly remained in English. Monolingual research infrastructure creates a language barrier for non-Anglophone speakers to access scholarly outputs and reinforces the idea that English is the only legitimate language for disseminating scholarship. Drawing from the debates on multilingual DH and scholarly publishing, we argue that any digital research infrastructure purporting to support knowledge diversity across disciplinary and national contexts must actively work to provide tools to facilitate, publish, and promote research in languages other than English. To show how multilingualism can guide infrastructure development and foster connections with diverse audiences, we describe the translation process of the interface of a research infrastructure, the Humanities and Social Sciences Commons, into four languages: French, Spanish, Bangla, and Portuguese.
This article examines how authorship is approached in story generation research publications. Story generation research forms its own, distinctive context in which computer-generated literary texts are being produced. Through text analysis and comparisons to other forms of computer-generated literature, the article examines what kind of rhetoric is used when the authorship of computer-generated texts is described in this context, and how the roles of the programmers and the program are characterised. The findings suggest that the approach to authorship in story generation research is mainly technical, referring strictly to the production process of the text, and leaving out the meaning of authorship as responsibility and accountability of the work as an aesthetic whole. This technical view affects how human-computer relations are discussed in the research, as "human-generated" and "computer-generated" texts are contrasted with each other. Furthermore, this dichotomy of the human and the machine affects how the produced stories are evaluated.
Most industrial programming languages leverage the English language for reserved keywords-words which a program compiler recognizes as specific execution commands. The divide between expert and novice programmers showcases an intriguing middle-ground by which the polysemy of many keywords becomes revealed. This essay explores a sampling of the multitude of keyword interpretations that a novice programmer may derive from the Java language's syntactic style and keywords specifically, and how the polysemy of both "English" and "code" meanings to these terms affects the novice-expert programmer transition.The transition from novice to expert, and the mapping of the career of metaphor to these keywords as a part of that process, can have implications for both teaching and learning programming.
Computational literary studies has developed sophisticated methods for detecting textual relationships-similarity networks, text reuse detection, semantic comparison-yet these approaches remain largely constrained to pairwise analysis. They reveal which texts resemble one another but cannot capture how individual works contribute to the collective semantic architecture of a corpus. This paper introduces a methodology that addresses this gap by shifting focus from text-to-text relationships to field formation. Treating a corpus as an integrated semantic space constructed through GloVe word embeddings, the method employs ablation testing-the systematic removal of individual texts followed by retraining-to measure each text's structural contribution to that space. Procrustes analysis quantifies the geometric difference between full and ablated embeddings, and an iterative protocol with consistency-based exponential decay filtering distinguishes genuine semantic disruption from algorithmic noise. The resulting "disruption scores" identify texts that are structurally necessary to a corpus's semantic field, capturing a dimension of textual significance irreducible to similarity, centrality, or explicit citation. Drawing on T.S. Eliot's concept of the "ideal order" as a heuristic for the relational and achronological dynamics the model quantifies, the paper demonstrates the methodology on a corpus of 142 twentieth-century Nigerian novels in English. Results validate the approach against established literary-historical scholarship while also surfacing critically overlooked texts whose structural significance warrants further attention. The methodology is portable across corpora and offers scholars a data-driven means of engaging intertextuality's foundational insight: that literary significance is relational, systemic, and collectively constituted.
Daniel Huws's Repertory of Welsh Manuscripts and Scribes c.800-c.1800 was published in 2022 to widespread acclaim - a "giant of scholarship" [Russell 2023, 151] which "... handed the world the keys to unlock the great riches of a thousand years of Welsh manuscript culture" [Guy 2024, 95]. Containing descriptions of c.3,300 manuscripts from sixty-eight repositories, along with detailed indexes and plates illustrating scribal hands, the Repertory is "... the kind of work upon which all others build" [Guy 2024, 89]. In his introduction to the Repertory, Huws immediately points to the future, writing that it "... belongs to that class of publication which can only reach maturity in a second edition" [Huws 2022, xxxiii]. This case study describes work towards the production of that edition, not as an expanded or revised set of printed volumes but as a dataset. First, we set the Repertory in the context of catalogues of Welsh manuscripts, and in the wider context of digital resources derived from manuscript descriptions. Then we describe the process of converting the Repertory into a dataset, and how this process, while technical in nature, served as a close reading of its scope and purpose. We offer some practical guidance on the conversion of print volumes to datasets and outline some initial research findings. Finally, we suggest some possible next steps for the Repertory, covering issues such as access and sustainability, integration with other resources and minimal computing approaches to analysing and surfacing the data.
Learning computing can often be challenging to Digital Humanities (DH) students because its knowledge-production practices and values do not necessarily align with those of humanistic inquiry. Further, the breadth of the technologies used in the field can make it difficult to identify which computational ideas and skills are most essential to teach in DH. This paper argues that, in order to teach students to enact a paradigm of 'humanistic computing' which is reflective of humanistic interests and ways of knowing, it is necessary to centre the epistemological friction that exists between computing and the humanities. It proposes a four-faceted framework to aid in the identification and exploration of that friction, with the intention that it be used by DH educators to structure their computing curriculum-regardless of the technologies or level of abstraction it engages with-and by students or even practitioners to scaffold a critical approach to computational DH work.
Podcasting has become a central infrastructure for the circulation and normalization of misogynistic and extremist ideologies, yet the scale, length, and affective density of audio content pose significant challenges for critical qualitative research. This paper introduces ManoWhisper, a feminist computational research infrastructure designed to support the large-scale analysis of misogynistic podcast ecosystems while preserving the contextual depth required for interpretive and ethical engagement. ManoWhisper combines automated audio acquisition, transcription, sentence-level classification, indexing, and visualization within a searchable web-based interface that enables researchers to move between computational pattern detection and close qualitative reading. Grounded in a feminist methodology of dwelling, the tool is designed to slow analysis, foreground emotional labour, and support collaborative research across varying levels of technical expertise. It allows for an in-depth consideration of more extremist media ecosystems across a variety of key factors. This paper documents ManoWhisper’s end-to-end pipeline, from content collection and transcription to classification, indexing, and interface design. Further we demonstrate its application across multiple peer-reviewed and public-facing research projects examining misogyny, masculinity, and gender-based extremism in podcasting, as well as how it is being used in policy and government institutions. We position ManoWhisper as methodological infrastructure that redistributes analytic capacity, makes repetition and scale visible, and enables ethically grounded engagement with harmful media. We conclude by reflecting on the tool’s limitations, its implications for feminist digital methods, and its relevance for understanding how misogynistic content circulates not only across media platforms but into emerging domains such as AI training data.
It is sometimes said that all politicians sound the same with their speeches mired in political jargon full of clich & eacute;s and false promises. To investigate how distinct the plenary speeches of political parties truly are and what linguistic features make them distinct, we trained a BERT classifier to predict the party affiliation of Finnish members of parliament from their plenary speeches. We contrasted and compared model performance to human responses to see how humans and the model differ in their ability to distinguish between the parties. We used the model explainability method SHAP to identify the linguistic cues that the model most relies on. We show that a deep learning model can distinguish between parties much more accurately than the respondents to the questionnaire. The SHAP explanations and questionnaire responses reveal that whereas humans tend to rely on mostly topical cues, the model has learned to recognize other cues as well, such as personal style and rhetoric.
As digital platforms continually evolve, rapid changes to platform affordances quickly render digital tools and data collection methods obsolete. Researchers of digital culture therefore must proactively adapt to the ephemerality of data. This paper examines these challenges within the context of Twitter (X) following its 2022 acquisition by Elon Musk, and the subsequent limiting of access to the API for data collection. Using a combination of manual data-collection practices and Zeeschuimer [Peeters 2025], a browser extension that collects social media data while browsing, researchers developed a novel data collection method, and model methodological adaptability within shifting digital terrains.
This article examines human-machine co-authorship as it represented by Ada Lovelace in her famous translation of and appendices to L. F. Menabrea's "Sketch of The Analytical Engine Invented by Charles Babbage." Lovelace's translation notes and correspondence with Charles Babbage are read alongside the Engine itself, as a platform, and the histories of calculating engines, software development, and nineteenth-century clerical labor. Inspired by critical code studies, the article performs a close reading of what is today referred to as the "first computer program," a sequence of steps that Lovelace adds to her translation's final "Note G" as an example of something the Engine can do. Ultimately, the article argues that the gendered power structures of collaborative work in Lovelace's time - and the challenges women authors faced in the nineteenth century more broadly - influenced understandings of machine programming, and that Lovelace's representation of the human-machine relationship in the first programmable calculating machine complicates the organizational structures of both clerical labor and software design.
In this article, I present reading computer source code aloud, and reflect on it together in dialogue, as a method which offers new forms of engagement with code and programming. With my peer discussion partner, we experimented with this methodology by exchanging recorded audio messages in a research project involving computational data-analysis. We discovered that abandoning the screen allowed us to de-explain, de-familiarize and de-contextualize code we had programmed, and negotiate new, shared insight of it as a cultural and situated object through re-explanation, re-familiarization, and re-contextualization of our experience of code.
Accurate dating of historical texts is essential for understanding cultural and historical narratives. However, traditional methods, such as paleographic and physical examination, can be subjective, costly, and potentially damaging to manuscripts. This paper introduces a machine learning approach to predicting the authorship dates of historical texts by using named entities — specifically, person and place names — as temporal markers. Using a dataset from Trismegistos, which includes metadata on the earliest and latest possible writing dates, we apply regression models to estimate text origins. While linear models like Lasso and Ridge Regression showed limited success, nonlinear models, including Random Forest, XGBoost, and Neural Networks, performed significantly better, with ensemble methods delivering the best results. The top-performing ensemble model achieved a mean absolute error of 45.7 years, surpassing traditional techniques. This study demonstrates the potential of named entities as temporal indicators and the effectiveness of ensemble learning in capturing complex historical patterns, offering a scalable, non-destructive alternative to traditional methods.
This paper introduces the conceptual framework for open and community-curated tool registries, posing that such registries provide fundamental value to any field of research by acting as curated knowledge bases about a community's past and current methodological practices as well as authority files for individual tools. The modular framework of a basic data model, SPARQL queries, bash scripts, and a prototypical web interface builds upon the well-established and open infrastructures of Wikimedia, GitLab, and Zenodo for creating, maintaining, sharing, curating, and archiving linked open data. We demonstrate the feasibility of this framework by introducing our concrete implementation of a tool registry for digital humanities, initially repurposing data from existing silos, such as TAPoR and the SSH Open Marketplace, and retaining the established TaDiRAH classification scheme while being open to communal editing in every aspect.
This article introduces the second Digital Humanities Quarterly special issue on Critical Code Studies. It provides an overview of the contributions to the issue while situating them within recent transformations in programming and code reading brought about by generative AI and LLM-based coding assistants. Against narratives that suggest automation diminishes the cultural significance of source code, the essay argues that code remains a meaningful cultural text even when generated or mediated by AI systems and that critical reading is therefore more urgent, not less, in the age of coding assistants. Serving as a state-of-the-field report since the first special issue, the article also surveys recent and forthcoming work in Critical Code Studies, including Alan Blackwell’s Moral Codes, Daniel Temkin’s 44 Esolangs, Nick Montfort’s Narcissystem, Tom Boellstorff and Braxton Soderman’s Intellivision, and the forthcoming Inventing ELIZA by Sarah Ciston et al.