The first three iterations of the ACM Hypertext conference were highly popular, drawing a diverse audience of researchers and practitioners experimenting with hypertext systems before the arrival of the World Wide Web. While the rise of the Web has often been seen as a disruptive event in the history of hypertext, in this paper I argue that if we focus on the intersection between digital publishing and hypertext, there are stronger continuities. Through archival research into early participants of the Hypertext conferences including Ben Shneiderman and Michael Joyce, I recover some of these lost connections.
Generative AI has become a buzzword within the publishing industry over the last two years, with responses often falling into either high optimism or an overall sense of doom. We are now at a suitable distance from the launch of ChatGPT to appraise these developments from a more nuanced perspective and begin to explore their connections to the longer history of AI. In this article, I offer some suggestions for how publishers might approach this topic. With the public release of ChatGPT in November 2022, ‘Generative AI’ has been heralded as one of the most significant technological breakthroughs since the printing press. Cutting through the hyperbole and the numerous possible counterexamples, the comparison is useful. The transition from manuscript to print culture was an on-going process rather than a sharp shift and we have not stopped prizing forms of manuscript writing centuries later. One mode of communication does not completely displace another; there’s little signs that generative AI will eradicate our need and desire for human creativity in fields such as publishing.
Controversies around AI companies' use of pirated book collections, including Books3 and Library Genesis, to train Large Language Models (LLMs) has led to increased scrutiny of books as AI training data. In this article, I contextualize these controversies in relation to the perceived value of books as a training data source compared to other textual data sources. Books are a liminal source for LLMs as they provide edited and curated long-form content while simultaneously presenting substantial legal risks and not aligning with the most popular genres of writing outputted by Generative AI services. I propose using the technical concept of 'epochs' in machine learning as a proxy for the perceived value of a data source, and using this metric to understand how AI companies value books in the training mix.
The digitization of the US Patent and Trademark Office’s (USPTO) backfile of six million patents undertaken between 1951 and 2001 was a five-decade struggle, featuring several media transitions from print and microfilm to CD-ROMs and, finally, the Web. This mass digitization project is on a similar scale to Google Books and the Internet Archive, but it is rarely discussed within critical digitization scholarship or for its significance as a tool for knowledge production. In this article, I focus on the USPTO’s patent document’s digital and physical material form and how the current paradigm of access and storage of the digital backfile emerged. Through this case study, I build upon Ian Milligan’s distinction between the ‘text’ and ‘platform’ layers of a digitization project to demonstrate how historical decisions regarding format and metadata continue to influence how users retrieve and interpret documents, such as patents, online.
Hypertext professionals have been writing the history of hypertext since Ted Nelson coined the term in the 1960s and claimed Vannevar Bush's Memex as a precursor to his Xanadu system. Despite the abundance of papers celebrating important figures and anniversaries in hypertext history, there has been less critical reflection on the methods for conducting this analysis. In this paper, I outline the dominant methods of writing histories of hypertext within the community. Through tracing the overlaps and gaps within this literature, I argue for a greater focus on regular users of these technologies and comparative analyses of hypertext in relation to broader trends. The paper concludes with a brief demonstration of how to apply this historical work through a case study of reading on-screen and hypertext.
While popular histories of the ebook start in the 1990s, inventors were working on the form since at least the 1940s. In this article, I offer a media archaeological analysis of digital publishing patents to develop the ebook imagination, or the desires of readers and inventors for the future of reading on screen. Through an analysis of a corpus of 98 patents relating to ebooks, I demonstrate how the ebook imagination focused on the aesthetics of the book over focusing on replicating paper via a screen, which would later lead to the success of Amazon's Kindle in 2007.
The use of computational methods to develop innovative forms of storytelling and poetry has gained traction since the late 1980s. At the same time, legacy publishing has largely migrated to using digital workflows. Despite this possible convergence, the electronic literature community has generally defined their practice in opposition to print and traditional publishing practices more generally. Not only does this ignore a range of hybrid forms, but it also limits non-digital literature to print, rather than considering a range of physical literatures. In this article, I argue that it is more productive to consider physical and digital literature as convergent forms as both a historicizing process, and a way of identifying innovations. Case studies of William Gibson et al.’s Agrippa ( A Book of the Dead) and Christian Bök’s The Xenotext Project’s playful use of innovations in genetics demonstrate the productive tensions in the convergence between digital and physical literature.
In order to explore monograph peer review in the arts and humanities, this article introduces and discusses an applied example, examining the route to publication of Danielle Fuller and DeNel Rehberg Sedo's Reading Beyond the Book: The Social Practices of Contemporary Literary Culture (2013). The book's co-authors supplemented the traditional blind' peer-review system with a range of practices including the informal, DIY review of colleagues and clever friends', as well as using the feedback derived from grant applications, journal articles and book chapters. The article explodes' the book into a series of documents and non-linear processes to demonstrate the significance of the various forms of feedback to the development of Fuller and Rehberg Sedo's monograph. The analysis reveals substantial differences between book and article peer-review processes, including an emphasis on marketing in review forms and the pressures to publish, which the co-authors navigated through the introduction of clever friends' to the review processes. These findings, drawing on science and technology studies, demonstrate how such a research methodology can identify how knowledge is constructed in the arts and humanities and potential implications for the valuation of research processes and collaborations.
Since the mid-2000s, the ebook has stabilized into an ontologically distinct form, separate from PDFs and other representations of the book on the screen. The current article delineates the ebook from other emerging digital genres with recourse to the methodologies of platform studies and book history. The ebook is modelled as three concentric circles representing its technological, textual and service infrastructure innovations. This analysis reveals two distinct properties of the ebook: a simulation of the services of the book trade and an emphasis on user textual manipulation. The proposed model is tested with reference to comparative studies of several ebooks published since 2007 and defended against common claims of ebookness about other digital textual genres.
The article was previously published in beta version (see https://www.stir.ac.uk/research/hub/publication/22749).
Amazon leads the market in ebooks with the Kindle brand, which encompasses a range of dedicated e-reader devices and a large ebook store. Kindle users are able to share the experience of reading ebooks purchased from Amazon by selecting passages of text for upload to the Kindle Popular Highlights website. In this article, I propose that the Kindle Popular Highlights database contains evidence that readers are re-appropriating commonplacing – the act of selecting important passages from a text and recording them in a separate location for later re-use – while reading public domain titles on the Kindle. An analysis of keyness in a corpus of 34,044 shared highlights from public domain titles suggests that readers focus on words relating to philosophy and values to draw an understanding of contemporary society from these classic works. This form of highlighting takes precedence over understanding and sharing key narrative moments. An examination of the top ten most popular authors in the corpus, and case studies of Jane Austen’s Pride and Prejudice and William Shakespeare’s Hamlet, demonstrate variation in highlighting practice as readers are choosing to shorten famous commonplaces in order to change their context for an audience that extends beyond the original reader. Through this analysis, I propose that Kindle users’ highlighting patterns are shaped by the behaviour of other readers and reflect a shared understanding of an audience beyond the initial highlighter.
Digital media presents several challenges to the index, but this ignores the fact that the index has played an important role in the development of the computer. Hypertext, or links between chunks of text, is a vital concept in computation, and one which can be traced back to the index. The author explores the link between indexes and hypertext through three case studies of novels with indexes: Vladimir Nabokov’s Pale fire, Mark Z. Danielewski’s House of leaves and Steven Hall’s The raw shark texts. This analysis reveals how indexes can be used as a subversive part of experimental fiction that authors employ to encourage the reader to move beyond superficial forms of reading.