While AI ethics interventions often focus on how researchers should navigate consequential choices, they may overlook a prior question: when do researchers recognize they are making a decision at all? This qualitative study examines how academic NLP teams confront “decision moments” – junctures where latent alternative paths could be considered. We propose a railyard problem analogy: where trolley problems presume a discrete choice between visible options, railyard problems concern whether alternative paths register as possibilities at all. Drawing on decision-tracing interviews across four NLP projects, we demonstrate how technical defaults, institutional structures, and tacit norms (infraethics) combine to organize research as a human-infrastructural process. Many consequential outcomes arise through "passive decisions", where alternatives exist but never become sufficiently visible, viable, or voiced (VVV) to warrant deliberation; "active decisions" only emerge when VVV conditions are met. Our analysis suggests ethics interventions should cultivate the collaborative conditions under which alternatives become recognizable.
This article challenges the view that war and interdependence are inherently incompatible by examining how combatants manage collective institutions during conflict. Using the internet as a case of such an institution, we show that belligerents selectively preserve or disrupt mutual access based on battlefield conditions. Disruption is more likely during mobile offensives, which offer greater operational freedom, while static or constrained operations incentivize maintaining interdependence for co-ordination, intelligence, or deception. Drawing on geolocated data from internet outages in the Russia-Ukraine war (2022-3) and qualitative evidence from this conflict and the Armenia-Azerbaijan conflicts (2020, 2023), we find that the disruption likelihood declines as battlefield constraints increase. These findings reveal how interdependence can serve as a tactical asset rather than merely a casualty of war. This has important implications for understanding the relationship between institutions and conflict, as wartime strategies shape not only battlefield outcomes but also prospects for post-war peace building.
Mobile broadband has long been central to increasing access to Internet connectivity [5] and enabling smarter and more connected communities [6]. However, hyperlocal knowledge about where mobile broadband is available and how well it performs is a persistent challenge [7]. The radio propagation models that are used to create mobile broadband coverage maps often overstate coverage [1, 2]. Under the data generated by these models, there exist underserved communities unable to access or effectively advocate for adequate broadband services. Within this critical gap between official datasets and the everyday experience of communities, the United States' Federal Communications Commission (FCC) launched the Mobile Availability Challenge (MAC)-a novel process that allows citizens to dispute the reported availability of mobile broadband service using on-the-ground measurements. This process presented an unprecedented opportunity for civic engagement to inform Internet policy and drive broadband deployment efforts. However, the tools used to collect these measurements have not been designed to center users and collective civic action, which poses a critical threat to meaningful adoption and sustainable impact. As part of our ongoing work to address digital inequalities and empower civic action around mobile broadband, we sought to design an open-source, mobile broadband measurement app that centers users and communities. CellWatch has been carefully co-designed with 69 community stakeholders representing municipal broadband efforts, tribal networks, and nonprofit organizations with the goal of incorporating a more community-centric perspective on mobile broadband measurement and mapping. CellWatch was designed through iterative user workshops using the FCC Speed Test app [3] as an initial model for compliance with the FCC MAC process. Our team worked with the Measurement Lab (M-Lab) team to use the Measurement Swiss Army Knife to integrate multi-stream TCP upload and download speed tests and latency measurements using distributed M-Lab servers as measurement endpoints. In contrast to the FCC Speed Test app, we focused on providing transparency into elements of cognizable challenges. Data collected from the Mobile App is submitted to the FCC Challenge Servers via official API. The CellWatch app is the first third-party application to be officially approved by the FCC for use in the broadband data collection mobile challenge [4]. A copy of the measurements are also stored in the CellWatch backend, implemented as a Supabase data repository with a REST API. The web-based researcher dashboard provides a large map view and comprehensive filtering and download capabilities for visualizing and analyzing measurement data. When a researcher clicks a measurement point on the map, they are able to see the download, upload and latency results collected at that location. The dashboard supports researchers in two ways. First, it allows researchers to explore data that has been collected to quickly make sense of spatial and temporal density of collected measurements. This can inform how data might be utilized for analysis or trace-driven modeling. Second, the dashboard also allows researchers to export data in various formats, including CSV and JSON. This supports researchers in conducting their own offline analysis. Link to demo: https://youtu.be/JUKG5pwoX3s
Scholars and practitioners alike have long been concerned with how design might contribute to social and political change. Much of that work has looked at the outcomes of design, the things made by designers. In this paper, we shift the view to look at how we—as design researchers—work with others. We offer the concept and practice of accompaniment to think about and do design research differently, with a focus on the character of the relationship between designers and those they work with.
In this paper, we provide the first comprehensive longitudinal analysis of government-ordered Internet shutdowns and spontaneous outages (i.e., disruptions not ordered by the government). We describe the available tools, data sources and methods to identify and analyze Internet shutdowns. We then merge manually curated datasets on known government-ordered shutdowns and large-scale Internet outages, further augmenting them with data on real-world events, macroeconomic and sociopolitical indicators, and network operator statistics. Our analysis confirms previous findings on the economic and political profiles of countries with government-ordered shutdowns. Extending this analysis, we find that countries with national-scale spontaneous outages often have profiles similar to countries with shutdowns, differing from countries that experience neither. However, we find that government-ordered shutdowns are many more times likely to occur on days of mobilization, coinciding with elections, protests, and coups. Our study also characterizes the temporal characteristics of Internet shutdowns and finds that they differ significantly in terms of duration, recurrence interval, and start times when compared to spontaneous outages.
We apply Lave & Wenger's construct of a community of practice to identify and position members of the data work community of practice, focusing on members on the periphery who have received less attention - as compared to full practitioners (e.g., data scientists). Reporting on results of interviews with 19 civic workers who perform data work as their main task, we identify an atypical relationship between subject-domain experts (such as our interviewees) and full members of the data work community. Our interviewees may have less computational skill in data work, but they have extensive and varied practices to engage in data contextualization that data scientists and other full community members could learn from. In identifying the attributes of data workers on the periphery, we also hope to call attention to the challenges they face in performing data work in low resources institutions (e.g., governmental, non-profit). Our findings contribute to the larger conversations in human-centered data science about who performs data work and how they go about it, in order to addresses questions of power, fairness, and bias in data-intensive systems.
Informed by critical data literacy efforts to promote social justice, this paper uses qualitative methods and data collected during two years of workplace ethnography to characterize the notion of critical novice data work. Specifically, we analyze everyday language used by novice data workers at DataWorks, an organization that trains and employs historically excluded populations to work with community data sets. We also characterize challenges faced by these workers in both cleaning and being critical of data during a project focused on police-community relations. Finally, we highlight novel approaches to visualizing data the workers developed during this project, derived from data cleaning and everyday experience. Findings and discussion highlight the generative power of everyday language and visualization for critical novice data work, as well as challenges and opportunities to foster critical data literacy with novice data workers in the workplace.
In response to widespread calls for computer scientists to better engage with the ethical dimensions of their work, there has been a surge of interest to embed ethics across the computer science (CS) curriculum. Yet one key set of barriers to doing so can be broadly described as scaling challenges -- in the number and breadth of courses in a curriculum and in the number of students in the CS major. Our paper describes and makes available a novel activity for teaching ethics using role-play that has advantages for scaling across different courses and in different delivery modes, including synchronous and asynchronous online course offerings. We describe our design process and early findings from developing the activity in a large first year seminar course, a senior-level computing and society class, and three different online graduate level courses. Further, we describe an evaluation survey that instructors can use to assess the short-term impact of the activity. We analyze survey results and our direct observations to reflect on the strengths and challenges of the activity. Our experiences suggest that role-play as a pedagogical tool can be particularly useful to broaden student perspectives and meaningfully incorporate ethics into CS courses.
Notions of the smart city look to mobilize information technology to increase organizational efficiency, and more recently, to support new forms of community engagement and involvement in addressing municipal issues. As cities turn to civic enterprise technology platforms, we need to better understand how that class of system might be positioned and used to collaborate with informal community-born coalitions. Beginning in 2019, we undertook an embedded collaborative research project in Albany Georgia, a small rural city, to understand three primary research questions: (1) How do community organizing practices take shape around joint initiatives with local government? (2) What data, tools, and process are needed to support those initiatives? (3) How do the affordances of City-run enterprise platforms support such community-born initiatives? To develop insight into these questions, we deployed a mixed-methods study that interwove participant observation, qualitative fieldwork, and participatory workshops. From this, we point to several mismatches that arose between the assumptions of a managed enterprise environment and the complex needs of establishing and supporting a multiparty community coalition.
Data has become central to the technologies and services that human-computer interaction (HCI) designers make, and the ethical use of data in and through these technologies should be given critical attention throughout the design process. However, there is little research on ethics education in computer science that explicitly addresses data ethics. We present and analyze Re-Shape, a method to teach students about the ethical implications of data collection and use. Re-Shape, as part of an educational environment, builds upon the idea of cultivating care and allows students to collect, process, and visualize their physical movement data in ways that support critical reflection and coordinated classroom activities about data, data privacy, and human-centered systems for data science. We also use a case study of Re-Shape in an undergraduate computer science course to explore prospects and limitations of instructional designs and educational technology such as Re-Shape that leverage personal data to teach data ethics.
We estimate the cost and impact of a proposed anti-displacement program in the Westside of Atlanta (GA) with data science and machine learning techniques. This program intends to fully subsidize property tax increases for eligible residents of neighborhoods where there are two major urban renewal projects underway, a stadium and a multi-use trail. We first estimate household-level income eligibility for the program with data science and machine learning approaches applied to publicly available household-level data. We then forecast future property appreciation due to urban renewal projects using random forests with historic tax assessment data. Combining these projections with household-level eligibility, we estimate the costs of the program for different eligibility scenarios. We find that our household-level data and machine learning techniques result in fewer eligible homeowners but significantly larger program costs, due to higher property appreciation rates than the original analysis, which was based on census and city-level data. Our methods have limitations, namely incomplete data sets, the accuracy of representative income samples, the availability of characteristic training set data for the property tax appreciation model, and challenges in validating the model results. The eligibility estimates and property appreciation forecasts we generated were also incorporated into an interactive tool for residents to determine program eligibility and view their expected increases in home values. Community residents have been involved with this work and provided greater transparency, accountability, and impact of the proposed program. Data collected from residents can also correct and update the information, which would increase the accuracy of the program estimates and validate the modeling, leading to a novel application of community-driven data science.
While there has been much anticipation that open government data (OGD) would increase the inclusion of marginalized groups in government decision-making processes, researchers have found little evidence of it. Such findings or lack of findings of social impact have led researchers to call for critical review of present notions of OGD's impact and also for better theoretical frameworks. In response to these calls, we develop a theoretical framework based on an ethnographic study of civic use of OGD in Hong Kong. We argue that constrained by the deliberative democracy models that focus on existing mechanisms of political participation, researchers have tended to overlook the use of OGD for protests, contestation, and other expressions of adversarial politics, which also produce a use of OGD for social impacts.
Researchers in human-centered computing have surfaced a feminist ethic of care in interaction with technologies, in data collection, and in data work. Drawing on two years of ethnographic fieldwork, we consider how democratic caring might be enacted and sustained through collaborative data work. We employ philosopher Joan Tronto's theory of caring democracy to structure our analysis of a resident-led initiative that uses data to organize and address issues of neglect and abandonment in their neighborhood. Adding to the CSCW literature on sociotechnical systems of care, we look particularly at Tronto's concept of caring democracy where caring needs and the ways in which they are met are an ongoing and inclusive process of assigning and reassigning caring responsibilities, characterized by both equality of voice and freedom from domination. This work develops grounded insight into the practice of democratic caring and how collaborative data work is relevant to this caring practice. We discuss opportunities and challenges for a data-supported caring democracy and address how caring democracy technologies are different from other modern civic technology practices. We conclude with a call to researchers to identify and enact democratic caring experiments in the small.
Data science is an interdisciplinary field that extracts insights from data through a multi-stage process of data collection, analysis and use. When data science is applied for social good, a variety of stakeholders are introduced to the process with an intention to inform policies or programs to improve well-being. Our goal in this paper is to propose an orientation to care in the practice of data science for social good. When applied to data science, a logic of care can improve the data science process and reveal outcomes of "good" throughout. Consideration of care in practice has its origins in Science and Technology Studies (STS) and has recently been applied by Human Computer Interaction (HCI) researchers to understand technology repair and use in under-served environments as well as care in remote health monitoring. We bring care to the practice of data science through a detailed examination of our engaged research with a community group that uses data as a strategy to advocate for permanently affordable housing. We identify opportunities and experiences of care throughout the stages of the data science process. We bring greater detail to the notion of human-centered systems for data science and begin to describe what these look like.
In this paper, we document the counter-data action and data activism of a grassroots affordable housing advocacy group in Atlanta. Our observation and insight into these data activities and strategies are achieved through ethnographic and engaged research and participatory design. We find that counter-data action through community-collected data is rooted in a legacy of Atlanta’s black activism and black scholarship; that this data activism enabled resource mobilization and critical conscious making; and that design and media production are essential post counter-data action activities in data activism. Based on these findings, we urge the field of open government data to broaden their concept of social impact of data to include the use data to mobilize resources within oppressed communities not to influence policy and government but to build capacities within community in order to transform, not join, political structures. We also advocate that scholars within the fields of open government data, critical data studies, and data activism recognize the legacy and historic practice of data activism by black communities working towards social change.
This paper reports on the use of web form, SMS, and chatbot for social election monitoring in the Dominican Republic for the May 15, 2016 General Elections. This case study provides evidence that bots can support the work of social election monitoring by effectively collecting election-relevant and actionable reports from voters on election irregularities. While this suggests bots are a promising tool, social media aggregators that enable crossmedia sourcing of reports are still needed as web form based crowdsourcing was the most effective for generating reports.
Cities across the United States are undergoing great transformation and urban growth. Data and data analysis has become an essential element of urban planning as cities use data to plan land use and development. One great challenge is to use the tools of data science to promote equity along with growth. The city of Atlanta is an example site of large-scale urban renewal that aims to engage in development without displacement. On the Westside of downtown Atlanta, the construction of the new Mercedes-Benz Stadium and the conversion of an underutilized rail-line into a multi-use trail may result in increased property values. In response to community residents' concerns and a commitment to development without displacement, the city and philanthropic partners announced an Anti-Displacement Tax Fund to subsidize future property tax increases of owner occupants for the next twenty years. To achieve greater transparency, accountability, and impact, residents expressed a desire for a tool that would help them determine eligibility and quantify this commitment. In support of this goal, we use machine learning techniques to analyze historical tax assessment and predict future tax assessments. We then apply eligibility estimates to our predictions to estimate the total cost for the first seven years of the program. These forecasts are also incorporated into an interactive tool for community residents to determine their eligibility for the fund and the expected increase in their home value over the next seven years.