Nous n’avons guère à rappeler que, depuis son fondement il y a quelque vingt-cinq ans, la Sator cherche à conjuguer la recherche littéraire avec les nouvelles technologies aptes à aider à réaliser le travail à la hauteur des ambitions du projet, c’est-à-dire à mieux comprendre et étudier la récurrence narrative dans les textes français du Moyen Âge à la fin du XVIIIe siècle. Diverses tentatives se sont succédé, que l’on pense à Toposator, SatorBase, TopoScan, PBLit, et d’autres initiatives connexes, chacune laissant entrevoir un nouveau monde de possibilités tout en nous laissant sur notre faim. Qu’est-ce qui distingue alors ce nouveau projet portant le nom Toucher (Textes, Outils, chercheur en réseau) ? Très peu, à certains égards : il s’agit toujours d’un groupe de chercheurs à la fois convaincus et circonspects au sujet du potentiel de l’informatique pour enrichir et multiplier les façons d’étudier les phénomènes littéraires. Cela dit, quelques différences essentielles existent, motivées par l’expérience et une réflexion honnête sur l’appui réel que peut apporter l’informatique aux chercheurs individuels et à l’équipe. Le concept du réseau nous a paru propice pour structurer nos objectifs, tant pour les composantes individuelles (les textes, les outils, les chercheurs), que pour le réseau qui peut se dessiner entre les composantes (un véritable réseau de réseau, ou internet). Dès lors, il ne s’agit plus simplement de créer un outil en isolation qu’on espère saura convaincre un utilisateur éventuel, mais plutôt de prendre conscience de l’ensemble des textes disponibles et de réfléchir à quelles sortes d’outils s’insérerait le plus naturellement dans les pratiques actuelles des chercheurs. Nous présenterons dans cet article les premières balbutiements du projet Toucher. En particulier, nous décrirons notre utilisation de Zotero, un outil bibliographique collaboratif, pour compiler un grand nombre de textes déjà numérisés (en différents formats et états). De là, nous présenterons un outil de balisage topique développé pour le projet qui permet de classifier des occurrences de mots-clés ainsi que les termes contribuant à la classification. Le but de cet outil est de permettre à la machine d’apprendre à reconnaître elle-même des occurrences possibles, ce qui permettrait de constituer une ressource extrêmement puissante : un utilisateur pourrait chercher dans un grand corpus des exemples possibles d’occurrences topiques afin d’enrichir la compréhension de phénomènes locaux ou d’alimenter une réflexion sur des phénomènes plus larges. Ainsi, textes, outils et chercheurs fonctionnent en réseau.
In the summer of 2017, Quinn Dombrowski, an IT staff member in UC Berkeley’s Research IT group, approached Geoffrey Rockwell about the possibility of merging the DiRT Directory with TAPoR, both popular tool discovery portals. Dombrowski could no longer offer the time commitment required to maintain the organizational structure of the volunteer-run tool directory (2018). This decommissioning of DiRT illustrates a set of problems in the digital humanities around tool directories and the tools within as academic contributions. Tool development, in general, is not considered sufficiently scholarly and often suffers from a lack of ongoing support (Ramsay & Rockwell, 2012). When tool discovery portals are no longer maintained due to a lack of ongoing funding, this leads to a loss of digital humanities knowledge and history. While volunteer-based directories require less outright funding, managing and motivating those volunteers to ensure that they remain actively involved in directory upkeep requires a vast amount work to ensure long-term sustainability (Dombrowski, 2018). This paper will explore the difficult history of tool discovery catalogues and portals and the steps being taken to save the DiRT Directory by integrating it into TAPoR. In particular, we will: – Provide a brief history of the attempts to catalogue tools for digital humanists starting with the first software catalogues, such as those circulated through societies, and ending with digital discovery portals, including DiRT Directory and TAPoR. – Discuss the challenges around the maintenance of discovery portals – Consider the design and metadata decisions made in the merging of DiRT Directory with TAPoR.
This paper looks at innovations in Busa's Index Thomisticus project through the tension between mechanical and human labour. We propose that Busa and the IBM engineer Paul Tasman introduced two major innovations that allowed computers to process unstructured text. The first innovation was figuring out how represent unstructured text on punched cards, the way data was encoded and handled at that time. The second innovation was figuring out how to tokenize unstructured phrases on cards into words for further counting, sorting and concording. We think-through these innovations using replication as a form of media archaeology practice that can help us understand the innovations as they were thought through at the time. All this is framed by a letter found in the Busa archives criticizing the project as a "tremendous mechanical labour... of no great utility." We use this criticism to draw attention to the very different meshing of human and mechanical labour developed at Busa's concording factory.
Digital technologies have profoundly changed our research practices, our editing and publication practices, and how we exchange information. The means have certainly changed, but these striking changes have also influenced the methods and meaning of writing, the concepts of writer and reader, and the methods of access, the tools and methods of distributing texts, images, sounds, etc. Literature has played a major role in the development of cultural domains; it still holds a prominent place, perhaps even more so today, as our current context is greatly associated to the predominance of the textual form. The explosion of data and the transformation of knowledge, both in terms of the quality of available information and the manner of representing it, have furthermore had a considerable impact on human and social sciences research.This special issue, which brings texts together from presentations made at the 2015 Digital Humanities colloquium at Concordia University, offers an overview of current trends in theory and practice observed in the digital humanities domain in French, more specifically in the sphere of production and distribution of knowledge in the human and social sciences. RésuméLes technologies numériques ont profondément modifié nos pratiques de recherche, d’édition, de publication et d’échanges d’informations. Les supports ont certes changé, mais les bouleversements ont aussi touché les modalités et le sens de l’écriture, les concepts d’auteur et de lecteur, les modalités d’accès, les supports et les modes de diffusion de textes, d’images, de sons, etc. Si la littérature a toujours joué un rôle capital dans la formation des catégories culturelles, elle s’approprie de manière peut-être encore plus marquée aujourd’hui, une place fondamentale, puisque le contexte actuel est largement associé à une prédominance de la forme textuelle. L’explosion des données et la transformation des savoirs, tant au plan de la quantité d’information disponible que de la manière de la représenter, a en outre un impact considérable sur la recherche en sciences humaines et sociales.Ce dossier spécial, qui rassemble des textes découlant de communications présentées lors du colloque Humanités numériques 2015 à l’Université Concordia, propose un aperçu des tendances théoriques et pratiques observées actuellement dans le domaine des humanités numériques en français, plus précisément dans la sphère de production et de diffusion du savoir dans les sciences humaines et sociales. Mots clés: Humanités numériques; outils de recherche; tendances théoriques et pratiques; édition électronique; pédagogie numérique; visualisation
Digital and Environmental HumanitiesStrong Networks, Innovative Tools, Interactive Objects Stephanie Posthumus (bio), Stéfan Sinclair (bio), and Veronica Poplawski (bio) Since the early 1990s, various humanities disciplines have been developing specific branches to respond ethically, historically, creatively, and critically to issues related to humans and the environment. Environmental history, philosophy, ethics, literary theory, education, and art, all share the belief that the humanities play a key role in understanding the ways in which environmental problems are socially and politically driven. Moreover, these new studies and approaches are keenly aware of the need for interdisciplinary scholarship when attending to complex environmental issues and concerns.1 Gaining momentum since around the 1950s, the digital humanities (previously known as humanities computing) have been responding to the increasing use of computer technology in contemporary culture. Inclusive in nature, the digital humanities include media studies, digital text analysis, big data, and visualization studies, to name a few.2 Collaborative and interdisciplinary in nature, the digital humanities have much in common with the environmental humanities. And yet these two fields have evolved largely independently. In the present article, we will describe the work we have been doing to bring the digital and the environmental humanities together by way of a set of timely projects. We begin by offering a rapid overview of the parallels between these two fields. We then outline an initiative in the digital environmental humanities that we have been leading for the last six years. What began as a networking workshop held at McGill [End Page 156] University in September 2013 (funded by the Social Sciences and Humanities Research Council of Canada) has subsequently developed in two different directions: (1) the application of topic modeling and visualization, to study a collection of scholarly texts and understand the emergence of the environmental humanities; and (2) the creation of interactive digital exhibits, to disseminate research in the environmental humanities. We will conclude by proposing some further ways in which the digital environmental humanities can continue to build strong networks, innovative tools, and interactive objects. Parallel Paths: The Digital and the Environmental Humanities At first glance, the digital and the environmental humanities appear to represent opposing forces. How can a field that embraces environmentally unfriendly computer technology help to further understand environmental issues? Should we not be reducing our use of high-energy cloud computing and discouraging the production of yet more e-waste? While these questions have merit, they remain quite narrow in scope. Rather than dismissing outright any association with computers (are electronic gadgets not just as prevalent in academic scholarship in the environmental humanities?), the environmental humanities can learn much from critical engagement with technology that has characterized the digital humanities since its inception.3 At the same time, digital humanities scholars can take away from the environmental humanities a more critical look at how the Internet, cloud services (Google, Facebook, etc.), and the electronic devices we use to access them have a real impact on the environment. A closer look reveals that the digital humanities are very much steeped in a humanities culture like that of the environmental humanities, a culture that promotes critical thinking and public engagement.4 Moreover, the digital humanities' long historical view on the emergence of new technologies is helpful in contextualizing the changes that contemporary culture is undergoing. In other words, the digital humanities are not simply embracing quantification and big data in the humanities (though critical perspectives are not always at the forefront).5 While introducing new methodologies that would have been foreign to the humanities in the past due to limits of time and scale, the digital humanities are also illustrating what makes the humanities [End Page 157] distinct from purely quantitative approaches. They underscore the role of interpretation, experimentation, play, reflection, and critical thinking when developing tools for humanist scholarship.6 Moreover, the hands-on approach within the digital humanities has called for more inclusivity in terms of details such as who learns to code, what tools they have access to, and where research centers are established. The visible and ongoing debate about coding and diversity illustrates that the digital humanities are necessarily bound up in the questions that identity politics have been...
This paper takes a media archaeology look at the development of the Keyword-in-Context (KWIC) display by Peter Luhn and how the KWIC helped automate ways of disseminating information about information. The paper takes the development of the KWIC as an example of the development of a knowledge technology that frames knowledge in a certain way. The KWIC and other information technologies transform knowledge into information that can be quantified and processed. Developments like the KWIC are the beginning of language engineeringa new way of conceiving of text as information to be manipulated. Finally, the paper proposes a way of reflecting on developments like the KWIC by replicating these early technologies. Replications can take the form of demonstration devices or knowledge things that expose the processes in our infrastructure.
We provide a comprehensive introduction to DREaM (Distant Reading Early Modernity), a hybrid text analysis and text archive project that opens up new possibilities for working with the collection of early modern texts in the EEBO-TCP collection (Phases I & II). Key functionalities of DREaM include i) management of orthographic variance; ii) the ability to create specially-tailored subsets of the EEBO-TCP corpus based on criteria such as date, title keyword, or author; and iii) direct export of subsets to Voyant Tools, a multi-purpose environment for textual visualization and analysis.
Wikipedia: The Free Encyclopedia was launched in January 2001, and its articles now represent a major resource for understanding the world. Many of these articles have been negotiated and edited for a decade or more, and the history of that editing can provide insight into the recent history of ideas. This paper describes the development of a tool called WIScker that works with the Wikipedia Application Programming Interface (API) to scrape, or "wisck," the revision history of any Wikipedia article, in order to build a corpus for subsequent text analysis and visualization. As an example, we examine a fourteen-year revision history of the article "Terrorism," first introduced into Wikipedia in October 2001, the month after 9/11, and subsequently expanded to provide a more historically informed, though still politically motivated, entry.
Background: The history of reading, writing, and the dissemination of technology is one of epochal change, and each transition – indeed the history of the book – is marked by hybridity. In the mature years of print, publishers, librarians, and scholars had clearly defined and segregated roles. In the digital realm, the boundaries have broken down. Just now we have hybridity of form and of roles in the implementation of new reading environments.Analysis: This article provides: 1) an overview of e-reading environments; 2) a survey of the Dynamic Table of Contexts interface; and 3) a report on the hybrid production process of a particular online text, Regenerations.Conclusion and implications: Regenerations could only have emerged from a collaboration among a digital infrastructure project, research project, university press, and digital humanities tool suite.
What can you actually do with text analysis? Chapter 4 is the first Interlude or example of text analysis in action with Voyant. This Interlude asks about computing in the humanities and shows how one can study the evolution of discourse about a discipline like the digital humanities though the Humanist discussion list that has been a central place for discussion and announcements since 1987. Studying 21 years of discourse shows how the emergence of the web as a platform for electronic resources was a turning point. Attention shifted from hardware and software to services and social media.
Interactive text analysis panels are showing up on all sorts of web pages. Word clouds, for example, can be easily inserted into blogs. Online newspapers commission more complex interactives so that readers can explore the speeches of politicians. This chapter surveys the history of such supports for interpretation. The chapter goes back to the development of the first concordances in the 13th century and looks at how a concordance is a tool for interpretation that brings together passages with the same word. Computers allowed the generation of concordances to be automated and the chapter surveys important projects starting with Father Busa’s Index Thomisticus project with IBM that started in the late 1940s, to John B. Smith’s interactive ARRAS, right up to web based concording tools like TACT and now Voyant.
What sort of thing is a text analysis tool and how does it bear theory? The ninth chapter looks at how things like models or demonstration devices can be used to share theories. This helps us understand theories about text tools from concordances to interactive visualizations. We return to John B. Smith and his ideas about computer criticism and visualization. We also look at Stephen Ramsay to present a model theory of text analysis as play. Tools bear interpretative baggage or rules that constrain and open how you can enter into a playful dialogue with a text. We then look at how tools like models and instruments bear theory we can speculate about how they can be better designed to be open themselves to interpretation. We close with the principles we followed building Voyant so that it is open to use and interpretation.
Browsing for information is a significant part of most research activity, but many online collections hamper browsing with interfaces that are variants on a search box. Research shows that rich-prospect interfaces can offer an intuitive and highly flexible alternative environment for information browsing, assisting hypothesis formation and pattern-finding. This unique book offers a clear discussion of this form of interface design, including a theoretical basis for why it is important, and examples of how it can be done. It will be of interest to those working in the fields of library and information science, human-computer interaction, visual communication design, and the digital humanities as well as those interested in new theories and practices for designing web interfaces for library collections, digitized cultural heritage materials, and other types of digital collections.
How can dialogue be structured and how can use analytics to ask about dialogue? Chapter 10 is the last Interlude or example of text analysis using Voyant Tools. In this Interlude we look at Hume’s Dialogues Concerning Natural Religion as an example of written dialogue and an example of an author showing humanistic dialogue. We tell the story of an ideal text analysis experiment that follows the drama of the dialogue and asks how themes are presented when the author doesn’t speak (only his characters do.) The theme of scepticism is shown as a way of questioning that has a long history in the humanities. We tell a story of questioning a dialogue in dialogue that is extended with analytical tools. The main character Philo’s scepticism is challenged as unrealistic. We believe that Philo’s scepticism is ultimately shown to be productive of understanding just as text analysis can show us new views that we can respond to in productive ways.