This article introduces the ‘Environmental Scan’ as answer to the question of hidden biases in digital heritage collections. Its substantive focus is digitized nineteenth-century British provincial newspapers, and in particular the JISC corpus, a popular, publicly-funded resource for scholars. While multiple papers have meticulously investigated the genesis of such newspaper collections, in the process highlighting the often unacknowledged politics of collection, preservation and dissemination via microfilm and now digitization, our aim is to explore questions of representativeness and bias in new ways by enriching computational analysis of digital corpora with the historical insights that can be derived from a contemporaneous reference source: namely, the Victorian newspaper press directories.
This chapter discusses the open access digitisation programme undertaken by Living with Machines, exploring the range of constraints that inform digitisation strategies and selection priorities. Because the landscape of digitised newspaper collections is so complex, and research and digitisation processes operate on different timelines, we have focused on opportunities to make digitisation choices both transparent and pragmatic. Working towards solutions that reflect collaborations between library staff and scholars, we introduce: a) Press Picker, our custom visualisation tool designed to support decision making about digitisation; and b) the Environmental Scan, a process of automatic metadata generation from the Newspaper Press Directories, a contemporaneous record of British newspapers.
We present a new dataset for the task of toponym resolution in digitized historical newspapers in English. It consists of 343 annotated articles from newspapers based in four different locations in England (Manchester, Ashton-under-Lyne, Poole and Dorchester), published between 1780 and 1870. The articles have been manually annotated with mentions of places, which are linked—whenever possible—to their corresponding entry on Wikipedia. The dataset consists of 3,364 annotated toponyms, of which 2,784 have been provided with a link to Wikipedia. The dataset is published in the British Library shared research repository, and is especially of interest to researchers working on improving semantic access to historical newspaper content.
This work presents defoe, a new scalable and portable digital eScience toolbox that enables historical research. It allows for running text mining queries across large datasets, such as historical newspapers and books in parallel via Apache Spark. It handles queries against collections that comprise several XML schemas and physical representations. The proposed tool has been successfully evaluated using five different large-scale historical text datasets and two HPC environments, as well as on desktops. Results shows that defoe allows researchers to query multiple datasets in parallel from a single command-line interface and in a consistent way, without any HPC environment-specific requirements.
We present a policy and process framework for secure environments for productive data science research projects at scale, by combining prevailing data security threat and risk profiles into five sensitivity tiers, and, at each tier, specifying recommended policies for data classification, data ingress, software ingress, data egress, user access, user device control, and analysis environments. By presenting design patterns for security choices for each tier, and using software defined infrastructure so that a different, independent, secure research environment can be instantiated for each project appropriate to its classification, we hope to maximise researcher productivity and minimise risk, allowing research organisations to operate with confidence.
Although there has been a drive in the cultural heritage sector to provide large-scale, open data sets for researchers, we have not seen a commensurate rise in humanities researchers undertaking complex analysis of these data sets for their own research purposes. This article reports on a pilot project at University College London, working in collaboration with the British Library, to scope out how best high-performance computing facilities can be used to facilitate the needs of researchers in the humanities. Using institutional data-processing frameworks routinely used to support scientific research, we assisted four humanities researchers in analysing 60,000 digitized books, and we present two resulting case studies here. This research allowed us to identify infrastructural and procedural barriers and make recommendations on resource allocation to best support non-computational researchers in undertaking ‘big data’ research. We recommend that research software engineer capacity can be most efficiently deployed in maintaining and supporting data sets, while librarians can provide an essential service in running initial, routine queries for humanities scholars. At present there are too many technical hurdles for most individuals in the humanities to consider analysing at scale these increasingly available open data sets, and by building on existing frameworks of support from research computing and library services, we can best support humanities scholars in developing methods and approaches to take advantage of these research opportunities.
In this article, we introduce recently released, publicly available resources, which allow users to watch videos of hidden articulators (e.g. the tongue) during the production of various types of sounds found in the world's languages. The articulation videos on these resources are linked to a clickable International Phonetic Alphabet chart ([International Phonetic Association. 1999. Handbook of the International Phonetic Association: A Guide to the Use of the International Phonetic Alphabet. Cambridge: Cambridge University Press]), so that the user can study the articulations of different types of speech sounds systematically. We discuss the utility of these resources for teaching the pronunciation of contrastive sounds in a foreign language that are absent in the learner's native language.
Dynamic Dialects is the product of a collaboration between researchers at the University of Glasgow, Queen Margaret University Edinburgh, University College London and Napier University, Edinburgh. Dynamic Dialects is an accent database, containing an articulatory video-based corpus of speech samples from world-wide accents of English. Videos in this corpus contain synchronised audio, ultrasound-tongue-imaging video and video of the moving lips.
A pitch from UCLDH and the British Library on 14th July 2015 as part of the Jisc Research Data Spring program, phase II: reporting on our pilot project, and where we see the project going forward: Lots of money has been spent digitising heritage collections. Digitised heritage collections are data. But non-computationally trained scholars don't know what to ask of large quantities of data. Often they do not have access to high performance computing facilities and they don’t know how to use them. We have addressed this fundamental problem by extending research data management processes in order to enable novel research in the arts, humanities, and social and historical sciences and a deeper understanding of emerging research needs. In our first phase, we have successfully implemented large scale, complex search of a digitised collection: now we scale up…
Seeing Speech (www.seeingspeech.ac.uk) is a web-based audiovisual resource which provides teachers and students of Practical Phonetics with ultrasound tongue imaging (UTI) video of speech, magnetic resonance imaging (MRI) video of speech and 2D midsagittal head animations based on MRI and UTI data. The model speakers are Dr Janet Beck of Queen Margaret University (Scotland) and Dr John Esling of University of Victoria (Canada). The first phase of this resource began in July 2011 and was completed in September 2013. Further funding was obtained in 2014 to improve and augment this resource (this version) and to develop its sister site Dynamic Dialects. The website contains two main resources: An introduction to UTI, MRI vocal tract imaging techniques and information about the production of the articulatory animations. Clickable International Phonetic Association charts links to UTI, MRI and animated speech articulator video. This online resource is a product of the collaboration between researchers at six Scottish Universities: The University of Glasgow, Queen Margaret University, Napier University, the University of Strathclyde, the University of Edinburgh and the University of Aberdeen; as well as scholars from University College London and Cardiff University. For examples of various dialects of English, please go to the sister site http://www.dynamicdialects.ac.uk
This chapter discusses the computational challenges and innovations encountered in the development of the Scottish corpora (the Scottish Corpus of Texts & Speech and the Corpus of Modern Scottish Writing), considers how tools for corpus analysis can encourage new audiences and complement existing resources, and explores possible future technological advances for corpus creation and exploitation.
This paper will introduce and demonstrate DiaView1, a new tool to investigate and visualise word usage in diachronic corpora. DiaView highlights cultural change over time by exposing salient lexical items from each decade or year, and providing them to the user in an effortless visualisation. This is made possible by examining large quantities of diachronic textual data, in this case the Google Books corpus (Michel et al., 2010) of one million English books. This paper will introduce the methods and technologies at its core, perform a demonstration of the tool and discuss further possibilities.
This paper will demonstrate ComPair, a new tool to investigate and compare word usage, encouraging new ways to explore language variation. While remaining focussed on the usability and the promotion of navigation, this tool represents an evolutionary step forward from the author’s previous award winning visualisation applications. This paper will introduce the methods and technologies at its core, perform a demonstration of the tool and discuss opportunities for further collaboration.