Introduction. Ensuring access to digital visual cultural heritage (VCH) for people who are blind or low vision (BLV) requires dedicated effort. Many digitised collections lack alt text, transcriptions, or descriptions that would enable independent access. Method. We conducted participatory design sessions with BLV participants cultural heritage experts, and stakeholders at the Library of Congress (LOC) to explore how access to VCH collections might be improved. Mixed groups of BLV arld sighted participants completed three activities: 1) identifying categories of information valued in image descriptions, 2) evaluating descriptive sources through a Wizard-of-Oz method, and 3) brainstorming to envision future systems. Analysis. We analysed category rankings from Activity 1, evaluations of descriptive squrces from Activity 2, and qualitatively examined insights from one study session of Activity 3. Results. Participants emphasised the importance of descriptions that extend beyond metadata, including vivid visual details, human presence, and contextua! information. Human-provided descriptions were most valued, Al-generated descriptions showed potential but required oversight, and metadata-based descriptions were seen as accurate yet insufficient. Conclusion. Our findings underscore the need for layered descriptive strategies combining factual precision with interpretive richness, ensuring equitable access fot blind users.
Text transcription is one of the most common forms of participatory activity undertaken by staff, researchers, and community groups who use collections from galleries, libraries, archives and museums (GLAMs) in different disciplinary contexts including citizen history or humanities and citizen science. This article furthers participatory transcription theory and practice by providing a framework of four common challenges in the creation, design, and execution of participatory transcription projects: 1) variety; 2) units of transcription; 3) single-track transcription or multi-track transcription with aggregation; and 4) bias. For each challenge, we articulate guiding questions based on participatory transcription and citizen science literature, conversations with other practitioners, and our own contributions to several large transcription projects and platforms including Zooniverse, which we present as a case study. Through this case study, we illuminate the affordances and challenges of participatory transcription and data aggregation methods through reference to multiple projects, publications, and hitherto unpublished information. We find that different project or platform design choices interact with the baseline challenges of text as a data type. This framework draws on and can be applied to other transcription projects and platforms beyond Zooniverse, and will support deeper understanding of participatory transcription practice by those creating new projects and those working with existing data. We conclude that, despite the rise of AI and threats to democratic institutions, participatory projects are vital in bringing together GLAMs, researchers, and the public in uniquely powerful ways.
Libraries can be a lifeline for people who are incarcerated or detained, their families and communities, yet library and information provision in American carceral settings varies wildly from state to state, and institution type to institution type. In this Commentary piece we describe how the ALA (with support from the Mellon Foundation) supported the work of writing a new standard for carceral library provision in the United States that better meets the needs of a justice-impaceted people and their families. The new Standards for Library Services for the Incarcerated or Detained provides concise recommendations and longer "Where it Worked" (WIW) narratives, showcasing how carceral librarians can partner with a broad range of stakeholders to meet the literacy, learning, legal, and recreational needs of individuals held in jails, prisons, detention facilities, juvenile facilities, immigration facilities, or prison work camps, whether public or private, military or civilian, in the United States and its territories. The new Standards explicitly address the needs of women, LGBTQIA+ people, the aged, people with dementia, people with a range of disabilities, and people who speak primary languages other than English. Library funding is often at the discretion of administrators who are not trained librarians, and who may not be aware of the extensive literature and evidence that demonstrates the importance of privacy of information access for incarcerated people (Austin 2021; Finlay and Bates 2019; Vogel 1995). The effects of restricted access to libraries and information have life-long implications for people who are incarcerated or detained, both inside carceral facilities and after release.
ABSTRACTThis panel explores the unique challenges and opportunities presented by the expanded sociotechnical practices associated with the preservation, curation, and use of visual information and ongoing perceptions of the affordances and constraints associated with new and emerging visual information objects. The panelists and the respondent work with the curation and preservation of visual information across a variety of research areas and utilize both qualitative and quantitative methods to understand the parallel and divergent sociotechnical challenges of curating visual information across analog, digitized, and born‐digital contexts. Sites of analysis range from format‐specific identity‐representation issues to the cultural practices of media archivists, translational challenges when moving complex scientific data across various digital and analog formats, and the role of accessibility in the design and deployment of discovery systems within cultural heritage institutions. Responding to the unique and intersecting challenges produced by working with visual information across various archival and digital curation contexts, the panelist will reflect on the practical and theoretical outcomes from ongoing research projects and identify persistent and emergent issues within the digital curation of visual information.
Many crowdsourcing and citizen science projects are conducted collaboratively by galleries, libraries, archives and museums (GLAMs) and research teams, and result in transcription data that could be helpful to blind and low vision users. Through usability and accessibility testing with blind and low vision scholarly users, we identify changes to GLAM systems and data sharing that can increase crowdsourcing, citizen science, and GLAM data discovery and utility for many blind and low vision users.
Galleries, libraries, archives, and museums (GLAM) are distinct but interconnected institutions that play crucial roles in preserving, studying, and sharing knowledge, culture, and heritage (GLAM). Many hold and preserve unique artifacts that they make available to the public in various forms, on-site through exhibitions and displays, and remotely, through digital surrogates such as images, audio/video files, in digital exhibits and through various forms of description such as catalog, archival, and museum records management systems. Oftentimes these artifacts and resources need be delivered to the public in varying modes and for varying reasons: no one description suits all occasions and purpose. Short and accurate descriptions of. GLAM resources and artifacts are vital to the work of public engagement. The appearance of large language models and interfaces to interact with them such as Chatgpt4 opens new opportunities for automatic content creation, while posing also new challenges. In this position paper, we propose to use these new tools as an aid for the curator to create suggestions for content that may be used as descriptions of artifacts while enabling also the adaptation of the content for varied audiences and even to personal preferences.
Screen readers are important assistive technologies for blind people, but they are complex and can be challenging to use effectively. Over the course of several studies with screen reader users, the authors have found wide variations and sometimes surprising differences in people's skills, preferences, navigation, and troubleshooting approaches when using screen readers. These differences may not always be considered in research and development. To help address this shortcoming, we have developed five user personas describing a range of screen reader experiences.
ABSTRACTHundreds of Libraries, Archives, and Museums (LAMs) around the world run crowdsourced transcription projects in order to engage users with their collections. Some LAMs explicitly use crowdsourcing projects to make non‐machine‐readable images of documents, such as manuscripts, discoverable to people who are blind or have low vision. We present findings from Crowdsourced Data: Accuracy, Accessibility, Authority (CDAAA), a 3‐year Institute of Museum and Library Services (IMLS) grant project that investigates whether and how LAMs integrate crowdsourced transcriptions into their discovery systems, and whether these efforts result in accessible web‐content for blind people and those with low vision who use assistive technology to navigate the web. We share research findings as well as practical suggestions for those in charge of crowdsourcing projects, the resulting transcription data, or similar web‐based textual content such as scholarly editions. These research and practice‐oriented findings are relevant to any national or local context where inaccessible images are transcribed, and are especially timely in the US context given recent Federal rule‐making to ensure that all web and app‐based content provided by US State and local governments is accessible, including tools, resources, and content created in‐house, through contracts or by license (2024).
The Shakespeare's World Datasets derive from a crowdsourced transcription project hosted on the Zooniverse platform between 2015 and 2019 (Van Hyning et al, 2015-2019). Volunteers transcribed 14,330 digitized early modern manuscripts from the Folger Shakespeare Library. 3,926 registered volunteers contributed, and over 94,570 anonymous sessions representing an unknown number of individuals, were recorded. This paper presents a cleaned dataset of individual volunteer transcriptions (IVT), containing all 203,389 valid classifications, along with three supplementary social datasets for a discussion forum and blog. These datasets provide insights into the transcription process, and volunteer and project owner interactions during the project's lifespan. The datasets have significant reuse potential in digital humanities, historical linguistics, and handwritten text recognition.
The Shakespeare’s World Datasets derive from a crowdsourced transcription project hosted on the Zooniverse platform between 2015 and 2019 (Van Hyning et al, 2015–2019). Volunteers transcribed 14,330 digitized early modern manuscripts from the Folger Shakespeare Library. 3,926 registered volunteers contributed, and over 94,570 anonymous sessions representing an unknown number of individuals, were recorded. This paper presents a cleaned dataset of individual volunteer transcriptions (IVT), containing all 203,389 valid classifications, along with three supplementary social datasets for a discussion forum and blog. These datasets provide insights into the transcription process, and volunteer and project owner interactions during the project’s lifespan. The datasets have significant reuse potential in digital humanities, historical linguistics, and handwritten text recognition.
ABSTRACT We are a team of citizen science volunteers and academics presenting a case study about how long‐serving Old Weather ( OW ) project volunteers left the leading citizen science platform, Zooniverse.org , and created their own opensource transcription tool to capture meteorological data from historic ship logbooks, for climate science. This project, built in LibreOffice Calc (hereafter LOC‐OW ) marks a transition from a hierarchical model of crowdsourcing to a co‐productive model in which the roles of the volunteers and the original project owners shifted, the volunteers gained expertise and developed a sense of ownership over the data production tools and process.
Abstract Current and future user expectations are not being met by the hierarchical structure of archival description and the emphasis on creating context that often comes at the expense of digitizing and releasing content in ways that are legible to both humans and machines. We propose a model of community and volunteer engagement that focuses limited staff time on the creation of web-searchable content through volunteer-driven ‘full text finding aids’ and scanning projects. Drawing on our individual experiences at the Folger and Zooniverse, and our collaboration on Shakespeare’s World, a crowdsourced transcription project, we advocate for community engagement via onsite and virtual crowdsourcing. Creating access to otherwise non-machine-readable content such as manuscript images, can expose archival silences, fill gaps in the historical record, better serve and include more diverse audiences, and create new pathways of discovery.
Historically, libraries, archives, and museums—or LAM institutions—have been complicit in enacting state power by surveilling and policing communities. This article broadens previous scholars’ critiques about individual institutions to LAM institutions writ large, drawing connections between these sites and ongoing racist, classist, and oppressive designs. We do so by dialing in on the ethical premise that justifies panoptic systems, utilitarianism, and how the glorification of pragmatism reifies systems of control and oppression. First, we revisit LIS applications of Benthamian and Foucauldian ideas of panoptic power to examine the role of LAM institutions as sites of social enmity. We then describe examples of surveillance and state power as they manifest in contemporary data infrastructure and information practices, showing how LAM institutional fixations with utilitarianism reify the U.S. carceral state through norms such as the aggregation and weaponization of user data and the overreliance on metrics. We argue that such practices are akin to widespread systems of surveillance and criminalization. Finally, we reflect on how LAM workers can combat structures that rely on oppressive assumptions and claims to information authority. Pre-print first published online February 10, 2023
The 'By the People' ('BTP') datasets comprise text of selected collections of the Library of Congress (LOC) created by volunteers in the 'By the People' crowdsourced transcription program, which invites public transcription of historical documents. All transcriptions are created and reviewed by volunteers in a consensus-based model in which two or more volunteers must agree on a transcription for it to be considered complete. Resulting transcriptions are added to the digital collections alongside the images to enable search and accessibility of the collections. Additionally, completed transcription “campaigns” are published as freely downloadable datasets of .CSV files containing all campaign transcriptions, as well as minimal metadata. The datasets can support a multitude of purposes including computational research in fields such as history, linguistics, economics, and political science.
Jaime Goodrich, Writing Habits: Historicism, Philosophy, and English Benedictine Convents, 1600 –1800, Tuscaloosa: University of Alabama Press, 2021, pp. 240, $59.95, ISBN: 978-0-8173-2103-1. - Volume 36 Issue 2