This special issue of JOCCH aims to present a wide-ranging review over the current state-of-the-art of research in computational archival science. The large-scale digitisation of analogue archives, the emerging diverse forms of born-digital archive, and the new ways in which researchers across disciplines (as well as the public) wish to engage with archival material are disrupting to traditional archival theories and practices and are presenting challenges for practitioners and researchers who work with archival material. They also offer enhanced possibilities for scholarship, through the application of computational methods and tools to the archival problem space, and, more fundamentally, through the integration of “computational thinking” with “archival thinking.” This potential led the collaborators in this special issue to identify Computational Archival Science (CAS) as a new field of study, and our working definition is:
Since the first ESFRI roadmap in 2006, multiple humanities Research Infrastructures (RIs) have been set up all over the European continent, supporting archaeologists (ARIADNE), linguists (CLARIN-ERIC), Holocaust researchers (EHRI), cultural heritage specialists (IPERION-CH) and others. These examples only scratch the surface of the breadth of research communities that have benefited from close cooperation in the European Research Area. While each field developed discipline-specific services over the years, common themes can also be distinguished. All humanities RIs address, in varying degrees, questions around research data management, the use of standards and the desired interoperability of data across disciplinary boundaries. This article sheds light on how cluster project PARTHENOS developed pooled services and shared solutions for its audience of humanities researchers, RI managers and policymakers. In a time where the convergence of existing infrastructure is becoming ever more important – with the construction of a European Open Science Cloud as an audacious, ultimate goal – we hope that our experiences inform future work and provide inspiration on how to exploit synergies in interdisciplinary, transnational, scientific cooperation.
The purpose of this article is to offer scholars and practitioners a more coherent and holistic starting point for asking questions about information and communication technologies for peacebuilding than has been available so far. A transdisciplinary proposal is made that applies critical pedagogy of peace education to the way that digital media can be used to build peace in communities and societies. This argument is further underpinned by insights from cognitive science and social psychology. The concept of sociotechnical consciousness is developed, which describes what it is like to be experiencing a sociotechnical system. We conclude that, to deploy digital media as part of peacebuilding initiatives, the media’s impact on individuals and groups deserve as much consideration as the content that is delivered via these media. This has important implications for how to design and use media in peacebuilding contexts.
Research and experimentation are underway in libraries, archives, and research institutions on various digital strategies, including computational methods and tools, to manage "Collections as Data" [1].This involves new ways for librarians and archivists to manage, preserve, and provide access to their digital collections.A major component in this ongoing process is the education and training needed by information professionals to function effectively in the 21st century.Accessible and transferable infrastructure is a key requirement in creating a network of collaboration for information professionals to fully realize the full potential of managing "Collections as Data."Elements needed include:1. Open source research and educational platforms to remove barriers to access to curation tools and resources.These are needed to deliver and share computational educational programs.2. Creation of a Cloud-based student-learning environment.3. Development of Open Source software architectures that use computational infrastructure.4. Exploration of new pedagogies for educating librarians and archivists in computational methods and tools. 5. Establishment of a community of practice for developing collaborative projects, and liaising with the wider international iSchool community and practitioners in the field.Our "Blue Sky" proposal seeks to explore a number of these challenges (infrastructure, computation, collaboration, learning) that stimulate the iSchool research community and have the potential to jumpstart international collaborative networks.A significant outcome would be the development of an international computational network for supporting librarians and archivists, akin to the existing Sloan Foundation funded "Data Curation Network", which over the next five years seeks to model a cross-institutional staffing approach for curating research data in digital repositories [2].
This workshop is the first step towards an emerging research community that problematises aspects related to collective consciousness. Collective consciousness sits at the intersection of - and goes beyond - research on Collective Intelligence, Collective Awareness and Behavioural Change. One of the goals of this workshop is to construct a working definition of collective consciousness which responds to a realised shortcoming in our understanding of digital technologies in mediating the transactions between people's capability to be aware and reflect, and their capacity to change behaviour and take action. This workshop aims to bring together researchers and practitioners who have an explicit intent to work critically with how digital technologies augment our collective experiences and behaviours.
The typology presented in Chapter 3 forms the basis for a set of thematic case studies of academic crowdsourcing in the humanities and cultural heritage. A set of cases are developed around the three key areas of digital content: spatial information, text and imagery. Here we trace how, in each case, academic crowdsourcing has impacted on the relationship between professional researchers and curators and public contributors. These case studies provide an essential link between the typology presented in the previous chapter and the conclusion that academic crowdsourcing has undergone a progression from a task-oriented business process to a research methodology.
Purpose – For decades, archivists have been appraising, preserving, and providing access to digital records by using archival theories and methods developed for paper records. However, production and consumption of digital records are informed by social and industrial trends and by computer and data methods that show little or no connection to archival methods. The purpose of this chapter is to reexamine the theories and methods that dominate records practices. The authors believe that this situation calls for a formal articulation of a new transdiscipline, which they call computational archival science (CAS).Design/Methodology/Approach – After making a case for CAS, the authors present motivating case studies: (1) evolutionary prototyping and computational linguistics; (2) graph analytics, digital humanities, and archival representation; (3) computational finding aids; (4) digital curation; (5) public engagement with (archival) content; (6) authenticity; (7) confluences between archival theory and computational methods: cyberinfrastructure and the records continuum; and (8) spatial and temporal analytics.Findings – Each case study includes suggestions for incorporating CAS into Master of Library Science (MLS) education in order to better address the needs of today’s MLS graduates looking to employ “traditional” archival principles in conjunction with computational methods. A CAS agenda will require transdisciplinary iSchools and extensive hands-on experience working with cyberinfrastructure to implement archival functions.Originality/Value – We expect that archival practice will benefit from the development of new tools and techniques that support records and archives professionals in managing and preserving records at scale and that, conversely, computational science will benefit from the consideration and application of archival principles.
This article addresses an important challenge in artificial intelligence research in the humanities, which has impeded progress with supervised methods. It introduces a novel method to creating test collections from smaller subsets. This method is based on what we will introduce as distant supervision' and will allow us to improve computational modelling in the digital humanities by including new methods of supervised learning. Using recurrent neural networks, we generated a training corpus and were able to train a highly accurate model that qualitatively and quantitatively improved a baseline model. To demonstrate our new approach experimentally, we employ a real-life research question based on existing humanities collections. We use neural network based sentiment analysis to decode Holocaust memories and present a methodology to combine supervised and unsupervised sentiment analysis to analyse the oral history interviews of the United States Holocaust Memorial Museum. Finally, we employed three advanced methods of computational semantics. These helped us decipher the decisions by the neural network and understand, for instance, the complex sentiments around family memories in the testimonies.
In the past, crowdsourcing has been seen as a transient activity, whose aims are, by definition, to get things done, without incurring unacceptable costs to the individuals contributing. Chapters 5 and 6Chapter 5Chapter 6 explore how and why a small minority of individuals transcend this and become ‘super-contributors’. This process is central to why academic crowdsourcing has developed in the way it has, and begs the question of how crowdsourcing activities create memory, and what kinds of memory are created. Drawing on existing descriptions of collective and individual memory, this chapter explores further some of the case studies presented earlier in the volume. It is argued that individualized narrative memories of the subject area of the asset concerned can be distinguished from ‘methodological narratives’ about the tasks and tools involved, which are more susceptible to be shared among the contributor community.
Most studies addressing the motivations for participating in humanities crowdsourcing have concluded that the majority of contributors do not have a single motivation but are moved to participate by multiple factors. This chapter examines these motivations, conceptualizing them in terms of an existing framework from motivation theory, and addressing both the reasons for participants' initial engagement with crowdsourcing and the ways in which these reasons evolve through their continuing involvement. In particular, we examine the roles of competition and gamification, the extent to which participants learn or develop new skills through their activities, and the connections between motivation and notions of community and belonging.
This chapter complements Chapter 1, examining the origins of crowdsourcing in the business world, and in the parallel developments of citizen science and online social engagement. These various forms of participatory activity are compared and distinguished, and in particular we examine how humanities crowdsourcing has facilitated the development of self-organizing communities and forms of peer production. Finally, a number of existing typologies for crowdsourcing are examined, as context to the typology introduced in Chapter 3.
This chapter delineates three phases in the evolution of humanities crowdsourcing: functional crowdsourcing, in which public involvement was used as a means of creating or enhancing digital resources for use primarily by the academic community; a second phase, in which more collaborative relationships and communities developed around and between volunteers and project organizers; and finally, a phase in which more independent research is beginning to emerge outside conventional academic institutions, and academic crowdsourcing is moving towards the co-production of knowledge. We also examine some possible near futures of crowdsourcing.
The key premise developed in Chapters 1 and 2Chapter 1Chapter 2 is that academic crowdsourcing forms a set of observable processes. In this chapter, we present a typology of these processes, with detailed explications of how and where they have been applied. First presented in Dunn and Hedges (2013), this typology identifies particular areas of activity and content for which academic crowdsourcing has proved significant. Putting this in the framework of a ‘methodological commons’, we describe four phases to the life cycle of academic crowdsourcing: asset, task type, process and output.
This chapter focuses on the roles of individual participants in academic crowdsourcing activities, and the communities they form. A key characteristic of academic crowdsourcing is the way in which it fosters and encourages the emergence of ‘super-contributors’, individuals who make disproportionately large or valuable contributions to a project's outputs. This chapter considers the general background of ways in which these roles have emerged, and compares them with more conventional academic research roles. The kinds of communities which academic crowdsourcing projects form are also considered, and similarly, compared to more conventional types of research community, notably, that of the modern university. Anticipating the book's more general conclusions about the evolution of academic crowdsourcing, it is argued that what underpins academic crowdsourcing in the humanities and cultural heritage is the kinds of relationships formed between super-contributor communities and professional researchers; and this may be seen as a natural, even inevitable, evolution of the ‘conventional’ university in the digital age.
Analyzing large quantities of real‐world textual data has the potential to provide new insights for researchers. However, such data present challenges for both human and computational methods, requiring a diverse range of specialist skills, often shared across a number of individuals. In this paper we use the analysis of a real‐world data set as our case study, and use this exploration as a demonstration of our “insight workflow,” which we present for use and adaptation by other researchers. The data we use are impact case study documents collected as part of the UK Research Excellence Framework (REF), consisting of 6,679 documents and 6.25 million words; the analysis was commissioned by the Higher Education Funding Council for England (published as report HEFCE 2015). In our exploration and analysis we used a variety of techniques, ranging from keyword in context and frequency information to more sophisticated methods (topic modeling), with these automated techniques providing an empirical point of entry for in‐depth and intensive human analysis. We present the 60 topics to demonstrate the output of our methods, and illustrate how the variety of analysis techniques can be combined to provide insights. We note potential limitations and propose future work.