We present an ethnographic study of an alternative approach to data work, developed by a civic-tech initiative that builds datasets for training and benchmarking online safety systems. They aim to respond to online safety concerns from a feminist perspective, by building safety datasets collaboratively with those most impacted by online harms. In this paper, we examine how this approach aims to reorient data work as a site for repair and redress, and trace the struggles they encounter in the process. Specifically, we draw attention to the challenges and tensions involved in advancing just reward for data work and collective governance of AI datasets. Examining these challenges through an STS-informed lens of reparative justice and repair, we argue that the work of repairing data work (and AI) lies, fundamentally, in resetting the ties of accountability. At a time heightened emphasis on efforts like safety evaluations and red teaming to make AI more responsible, we highlight the need to confront foundational questions about how the humans involved in these efforts relate to the datasets and systems they help produce. A reparative lens demands that we interrupt prevailing norms of data work and place at their centre, not AI or datasets, but those most harmed by the neglect, oversight and exclusion animated in the current modes of dataset production. This, we argue, offers a bold vision for responsibility and contributes towards a critical agenda for building alternative futures of data and AI practice.
The Project of AI is a world-building endeavor, wherein those who fund and develop AI systems both operate through and seek to sustain networks of power and wealth. As they expand their access to resources and configure our sociotechnical conditions, they benefit from the ways in which a suite of decoys animate scholars, critics, policymakers, journalists, and the public into co-constructing industry-empowering AI futures. Regardless of who constructs or nurtures them, these decoys often create the illusion of accountability while both masking the emerging political economies that the Project of AI has set into motion, and also contributing to the network-making power that is at the heart of the Project's extraction and exploitation. Drawing on literature at the intersection of communication, science and technology studies, and economic sociology, we examine how the Project of AI is constructed. We then explore five decoys that seemingly critique - but in actuality co-constitute - AI's emergent power relations and material political economy. We argue that advancing meaningful fairness or accountability in AI requires: 1) recognizing when and how decoys serve as a distraction, and 2) grappling directly with the material political economy of the Project of AI. Doing so will enable us to attend to the networks of power that make 'AI' possible, spurring new visions for how to realize a more just technologically entangled world.
This one-day workshop aims to map data and its inherent connections to work (of all kinds) across a landscape of ongoing crises. The workshop brings together researchers and practitioners with an interest in data work that underpins automation, algorithmic systems and organizational and societal strives toward datafication. The workshop provides a forum for interdisciplinary discussions around controversies related to data and work - and data work in particular - with the aim to expand the toolbox for working with data by proposing and developing critical approaches, drawing on the rich contributions of the growing body of literature on data work and datafication. Through spatial and temporal mapping exercises, the workshop intends to both trace paths through past crises into a contemporary moment, and towards more hopeful futures.
This paper considers snooker's rise to popularity, and its relative decline, through the frame of recent British social history. The paper situates an ostensible decline in snooker spectatorship and a demonstrable decline in participation across the UK, against a backdrop of shifts in economic activity, class structure, cultures of masculinity and urban space. Drawing on theories of gender, class, subculture, media and critical urbanism, the paper argues that a sociological frame lends a lot to understanding snooker in the UK. At the same time, it argues that the frame of snooker might also lend a lot to a sociological understanding of industrial, and later, post-industrial Britain.
Human annotation plays a core role in machine learning -- annotations for supervised models, safety guardrails for generative models, and human feedback for reinforcement learning, to cite a few avenues. However, the fact that many of these human annotations are inherently subjective is often overlooked. Recent work has demonstrated that ignoring rater subjectivity (typically resulting in rater disagreement) is problematic within specific tasks and for specific subgroups. Generalizable methods to harness rater disagreement and thus understand the socio-cultural leanings of subjective tasks remain elusive. In this paper, we propose GRASP, a comprehensive disagreement analysis framework to measure group association in perspectives among different rater sub-groups, and demonstrate its utility in assessing the extent of systematic disagreements in two datasets: (1) safety annotations of human-chatbot conversations, and (2) offensiveness annotations of social media posts, both annotated by diverse rater pools across different socio-demographic axes. Our framework (based on disagreement metrics) reveals specific rater groups that have significantly different perspectives than others on certain tasks, and helps identify demographic axes that are crucial to consider in specific task contexts.
Artificial Intelligence (AI) is increasingly used in mainstream applications to make decisions that affect a large number of people. While research has focused on involving machine learning and domain experts during the development of responsible AI systems, the input of lay users has too often been ignored. By exploring the involvement of lay users, our work seeks to advance human-centric responsible AI development processes. To reflect on lay users’ views, we conducted an online survey of 1121 people in the United Kingdom. We found that respondents had concerns about fairness and transparency of AI systems which requires more education around AI to underpin lay user involvement. They saw a need for having their views reflected at all stages of the AI development lifecycle. Lay users mainly charged internal stakeholders to oversee the development process but supported by an ethics committee and input from an external regulatory body. We also probed for possible techniques for involving lay users more directly. Our work has implications for creating processes that ensure the development of responsible AI systems that take lay user perspectives into account.
In this paper, we examine the work of data annotation. Specifically, we focus on the role of counting or quantification in organising annotation work. Based on an ethnographic study of data annotation in two outsourcing centres in India, we observe that counting practices and its associated logics are an integral part of day-to-day annotation activities. In particular, we call attention to the presumption of total countability observed in annotation - the notion that everything, from tasks, datasets and deliverables, to workers, work time, quality and performance, can be managed by applying the logics of counting. To examine this, we draw on sociological and socio-technical scholarship on quantification and develop the lens of a 'regime of counting' that makes explicit the specific counts, practices, actors and structures that underpin the pervasive counting in annotation. We find that within the AI supply chain and data work, counting regimes aid the assertion of authority by the AI clients (also called requesters) over annotation processes, constituting them as reductive, standardised, and homogenous. We illustrate how this has implications for i) how annotation work and workers get valued, ii) the role human discretion plays in annotation, and iii) broader efforts to introduce accountable and more just practices in AI. Through these implications, we illustrate the limits of operating within the logic of total countability. Instead, we argue for a view of counting as partial - located in distinct geographies, shaped by specific interests and accountable in only limited ways. This, we propose, sets the stage for a fundamentally different orientation to counting and what counts in data annotation.
Machine learning approaches often require training and evaluation datasets with a clear separation between positive and negative examples. This risks simplifying and even obscuring the inherent subjectivity present in many tasks. Preserving such variance in content and diversity in datasets is often expensive and laborious. This is especially troubling when building safety datasets for conversational AI systems, as safety is both socially and culturally situated. To demonstrate this crucial aspect of conversational AI safety, and to facilitate in-depth model performance analyses, we introduce the DICES (Diversity In Conversational AI Evaluation for Safety) dataset that contains fine-grained demographic information about raters, high replication of ratings per item to ensure statistical power for analyses, and encodes rater votes as distributions across different demographics to allow for in-depth explorations of different aggregation strategies. In short, the DICES dataset enables the observation and measurement of variance, ambiguity, and diversity in the context of conversational AI safety. We also illustrate how the dataset offers a basis for establishing metrics to show how raters' ratings can intersects with demographic categories such as racial/ethnic groups, age groups, and genders. The goal of DICES is to be used as a shared resource and benchmark that respects diverse perspectives during safety evaluation of conversational AI systems.
Large language models (LLMs) trained on real-world data can inadvertently reflect harmful societal biases, particularly toward historically marginalized communities. While previouswork has primarily focused on harms related to age and race, emerging research has shown that biases toward disabled communities exist. This study extends prior work exploring the existence of harms by identifying categories of LLM-perpetuated harms toward the disability community. We conducted 19 focus groups, during which 56 participants with disabilities probed a dialog model about disability and discussed and annotated its responses. Participants rarely characterized model outputs as blatantly offensive or toxic. Instead, participants used nuanced language to detail how the dialog model mirrored subtle yet harmful stereotypes they encountered in their lives and dominant media, e.g., inspiration porn and able-bodied saviors. Participants often implicated training data as a cause for these stereotypes and recommended training the model on diverse identities from disability-positive resources. Our discussion further explores representative data strategies to mitigate harm related to different communities through annotation co-design with ML researchers and developers.
Diversity in datasets is a key component to building responsible AI/ML. Despite this recognition, we know little about the diversity among the annotators involved in data production. We investigated the approaches to annotator diversity through 16 semi-structured interviews and a survey with 44 AI/ML practitioners. While practitioners described nuanced understandings of annotator diversity, they rarely designed dataset production to account for diversity in the annotation process. The lack of action was explained through operational barriers: from the lack of visibility in the annotator hiring process, to the conceptual difficulty in incorporating worker diversity. We argue that such operational barriers and the widespread resistance to accommodating annotator diversity surface a prevailing logic in data practices—where neutrality, objectivity and ‘representationalist thinking’ dominate. By understanding this logic to be part of a regime of existence, we explore alternative ways of accounting for annotator subjectivity and diversity in data practices.
Recent advancements in conversational AI have created an urgent need for safety guardrails that prevent users from being exposed to offensive and dangerous content. Much of this work relies on human ratings and feedback, but does not account for the fact that perceptions of offense and safety are inherently subjective and that there may be systematic disagreements between raters that align with their socio-demographic identities. Instead, current machine learning approaches largely ignore rater subjectivity and use gold standards that obscure disagreements (e.g., through majority voting). In order to better understand the socio-cultural leanings of such tasks, we propose a comprehensive disagreement analysis framework to measure systematic diversity in perspectives among different rater subgroups. We then demonstrate its utility by applying this framework to a dataset of human-chatbot conversations rated by a demographically diverse pool of raters. Our analysis reveals specific rater groups that have more diverse perspectives than the rest, and informs demographic axes that are crucial to consider for safety annotations.
In this paper, we present findings from an semi-experimental exploration of rater diversity and its influence on safety annotations of conversations generated by humans talking to a generative AI-chat bot. We find significant differences in judgments produced by raters from different geographic regions and annotation platforms, and correlate these perspectives with demographic sub-groups. Our work helps define best practices in model development -- specifically human evaluation of generative models -- on the backdrop of growing work on sociotechnical AI evaluations.
Conversational AI systems exhibit a level of human-like behavior that promises to have profound impacts on many aspects of daily life -- how people access information, create content, and seek social support. Yet these models have also shown a propensity for biases, offensive language, and conveying false information. Consequently, understanding and moderating safety risks in these models is a critical technical and social challenge. Perception of safety is intrinsically subjective, where many factors -- often intersecting -- could determine why one person may consider a conversation with a chatbot safe and another person could consider the same conversation unsafe. In this work, we focus on demographic factors that could influence such diverse perceptions. To this end, we contribute an analysis using Bayesian multilevel modeling to explore the connection between rater demographics and how raters report safety of conversational AI systems. We study a sample of 252 human raters stratified by gender, age group, race/ethnicity group, and locale. This rater pool provided safety labels for 1,340 human-chatbot conversations. Our results show that intersectional effects involving demographic characteristics such as race/ethnicity, gender, and age, as well as content characteristics, such as degree of harm, all play significant roles in determining the safety of conversational AI systems. For example, race/ethnicity and gender show strong intersectional effects, particularly among South Asian and East Asian women. We also find that conversational degree of harm impacts raters of all race/ethnicity groups, but that Indigenous and South Asian raters are particularly sensitive to this harm. Finally, we observe the effect of education is uniquely intersectional for Indigenous raters, highlighting the utility of multilevel frameworks for uncovering underrepresented social perspectives.
Under what technoscientific conditions might the scarcity of food be understood as contingent on heterogeneous actors? And how might the possibilities of food abundance be approached as a reparative project of valuing their manifold relations? Blockchain promises to be an infrastructure that presents both productive imaginaries and also challenges to such restorative and sustainable work. In a series of workshops, we critically experimented with these possibilities and challenges. Working with diverse participants including community growers, organizers, artists and technologists we used a variety of playful methods to act out fictional scenarios set in 2025, when all of London had been transformed into a city farm. For organizations and participants, reparation meant working in the aftermath of social and environmental collapse to bring into being more-than-human-value systems that radically decentred human knowledge and experience.
Digital technologies such as sensors, blockchain, and artificial intelligence are increasingly being used in global food production and consumption, including in urban contexts through the notion of the “smart city”. Food governance is addressed through logics of efficiency in supply chains, increasing profits for shareholders of large companies but doing little to address unsustainable social and environmental inequalities. Human–computer interaction (HCI) designers and researchers are increasingly interested in algorithmic governance of smart cities, raising concerns around issues of control, agency, access, and benefit. Against these concerns, some HCI researchers have also started to question a human-centred perspective to designing socio-technical systems, drawing on more-than-human perspectives to consider the interrelations and interdependencies between human and non-human others within the food web. In this chapter, these emerging perspectives within HCI are drawn together to consider the ways in which new technologies such as blockchain can be used in urban food governance. A case study of co-designing futures of algorithmic food governance with grassroots urban communities that account for multispecies actors, labours, and relationships is presented. The project surfaces new possibilities for computation to intervene in urban food governance in ways that are more sustainable and fairer.
Current global crises, including natural disasters, pandemics, and political cataclysms, are exacerbating intolerable living conditions for the majority. This chapter examines potentials and perils in the uneven terrains of the current food systems. It problematises the current design and maintenance of food systems driven by technocapitalist agendas and highlights the need to interject new structures and processes for governance. Drawing from emerging calls for open, commons-based urban food futures, the chapter discusses why urban food governance must focus on the process of relating multiple, always changing aspects of food. Here, questioning and imagining diverse futures is critical. This chapter suggests two ways to achieve this: (1) repairing and repurposing current closed systems to become more open, and (2) building capacity in learning to cope with complex relational conditions bringing with them uncertainties, differences, and emergence of the new. The chapter argues that transdisciplinary and creative practices must become an inherent part of building food futures, urban and otherwise, as a means of imagining and raising questions, rather than solving predefined or assumed problems, thereby creating spaces for new forms of governance to emerge, based on equity, justice, and pluralism.
Nicholas Villar合作论文数Sensors and Devices Group, part of the Computer Mediated Living Group of Microsoft Research13