Concerns about the environmental footprint of machine learning are increasing. While studies of energy use and emissions of ML models are a growing subfield, most ML researchers and developers still do not incorporate energy measurement as part of their work practices. While measuring energy is a crucial step towards reducing carbon footprint, it is also not straightforward. This paper introduces the main considerations necessary for making sound use of energy measurement tools and interpreting energy estimates, including the use of at-the-wall versus on-device measurements, sampling strategies and best practices, common sources of error, and proxy measures. It also contains practical tips and real-world scenarios that illustrate how these considerations come into play. It concludes with a call to action for improving the state of the art of measurement methods and standards for facilitating robust comparisons between diverse hardware and software environments.
Language models play a vital role in various natural language processing tasks, but their training can be computationally intensive and lead to significant carbon emissions. In this study, we explore the effectiveness of timeshifting strategies to mitigate the environmental impact of long-running large language models (LLMs). We develop a simulation tool that estimates carbon emissions for LLMs, enabling developers to make informed decisions prior to running their workloads. By leveraging historical carbon intensity data from WattTime, we investigate the potential benefits and limitations of timeshifting in different locations, considering diverse energy profiles. Our findings demonstrate that timeshifting can substantially reduce emissions, but it is highly dependent on the region’s carbon intensity and energy mix. We present insights into the trade-offs between emissions reduction and workload runtime, acknowledging the need for further advancements in carbon-aware computing practices. Our research contributes to the growing field of sustainable computing and encourages developers to adopt environmentally conscious strategies in language model training.
Abstract Language models play a vital role in various natural language processing tasks, but their training can be computationally intensive and lead to significant carbon emissions. In this study, we explore the effectiveness of timeshifting strategies to mitigate the environmental impact of long-running language models (LLMs). We develop a simulation tool that estimates carbon emissions for LLMs, enabling developers to make informed decisions before running their workloads. By leveraging historical carbon intensity data from WattTime, we investigate the potential benefits and limitations of time-shifting in different locations, considering diverse energy profiles. Our findings demonstrate that time-shifting can substantially reduce emissions for certain workloads, but it is highly dependent on the region's carbon intensity and energy mix. We present insights into the trade-offs between emissions reduction and workload runtime, acknowledging the need for further advancements in carbon-aware computing practices. Our research contributes to the growing field of sustainable computing and encourages developers to adopt environmentally conscious strategies in language model training.
What makes one dataset powerful for civic advocacy, and another fall flat? Drawing from a citizen science project on environmental health, I argue that there is an underacknowledged quality of datasets-their topology-that shapes the social, cultural, and political possibilities they can sustain or subvert. Data topologies are formal qualities of a dataset that connect data collectors' intentions with the types of calculations that can and cannot be performed. This configures how numerical arguments are made, and the sociotechnical imaginaries those arguments sustain or subvert. The citizen science project's data topology made any easy notion of shared exposure to pollutants, or singular health effects, unravel. The data appeared to tell a story of atypicality at scale, where each person suffers differently from different exposure. Lacking a central tendency, or pockets of tendency disproportionately carried by different subgroups, it became it harder, not easier, for citizen scientists to use data in regulatory contexts, where dominant sociotechnical imaginaries conceive of difference in epidemiological and toxicological terms.
Responsible artificial intelligence guidelines ask engineers to consider how their systems might harm. However, contemporary artificial intelligence systems are built by composing many preexisting software modules that pass through many hands before becoming a finished product or service. How does this shape responsible artificial intelligence practice? In interviews with 27 artificial intelligence engineers across industry, open source, and academia, our participants often did not see the questions posed in responsible artificial intelligence guidelines to be within their agency, capability, or responsibility to address. We use Suchman's “located accountability” to show how responsible artificial intelligence labor is currently organized and to explore how it could be done differently. We identify cross-cutting social logics, like modularizability, scale, reputation, and customer orientation, that organize which responsible artificial intelligence actions do take place and which are relegated to low status staff or believed to be the work of the next or previous person in the imagined “supply chain.” We argue that current responsible artificial intelligence interventions, like ethics checklists and guidelines that assume panoptical knowledge and control over systems, could be improved by taking a located accountability approach, recognizing where relations and obligations might intertwine inside and outside of this supply chain.
Critiques of data colonialism and surveillance capitalism focus on data collected from online behavior. We propose that analytical concepts from these critiques-namely, regimes of value and patterns of alienation and attunement-could be applied more widely to better understand the threats that datafication poses to equity and democracy in the social and environmental realms. Regimes of value, which include the institutions and technologies that make data meaningful and render them selectively available for appropriation, are relevant both to for-profit companies' data practices and to states' participation in the datafication of the environment; examining regimes of value raises questions about how data are exploited and how they are neglected. Patterns of alienation associated with datafication include the potential for alienation from the environment; however, at least in some value regimes, alienation may be accompanied by possibilities for attunement to natural and social phenomena that might otherwise have escaped notice.
Data transmission tends to be neglected when considering the carbon efficiency of systems, even though the electricity usage of data networks as a whole is as large, or larger, than that of data centers. Accounting for carbon cost of the movement of data is hard, and is often assumed to be the responsibility of the receiver or an intermediate provider. To be able to account for the carbon footprint of networks, mutually agreed metrics are required, covering the end-to-end environmental cost of data transmission and up-the-stack network software costs of data processing, rather than merely the independent network devices. Beyond discussing the considerations for defining these metrics, this paper suggests building upon existing practices, such as network telemetry, programmable network elements and cost-aware routing to enable carbon-intelligent networking, a concept that goes beyond network energy efficiency and considers the impact of energy decarbonization on the routing and scheduling of data transmission.
This paper explores the complex relationship between intellectual property (IP) and the transdisciplinary collaborative design (co-design) of new digital technologies for agriculture (AgTech). More specifically, it explores how prioritizing the capturing of IP as a central researcher responsibility can cause disruptions to research relationships and project outcomes. We argue that boundary-making processes associated with IP create a particular context through which responsibility can, and must, be located and cultivated by researchers working within transdisciplinary collaborations. We draw from interview data and situated IP practices from a transdisciplinary co-design project in Aotearoa New Zealand to illustrate how IP is a fluid boundary-requiring-and-producing object that impels researchers into its management, and produces tensions that need to be noticed and skillfully navigated within research relations. We propose located response-ability as a conceptual tool and practice to reposition IP within the relations that make up a transdisciplinary co-design project, as opposed to prioritizing IP by default without recognizing its possible impacts on collaborative relations and other project aims and accountabilities. This can support researchers practicing responsible innovation in making everyday decisions on how to protect potential IP without disrupting the collaborative relations that make the creation of potential IP possible, and the existence of protected IP relevant and beneficial to project collaborators and wider societal actors. This may help to ensure that societal benefits can be generated, and positive science–society relationships prioritized and preserved, in the design of new AgTech.
Open source software communities are a significant site of AI development, but "Ethical AI" discourses largely focus on the problems that arise in software produced by private companies. Design, policy and tooling interventions to encourage "Ethical AI" based on studies in private companies risk being ill-suited for an open source context, which operates under radically different organizational structures, cultural norms, and incentives.
HCI tends to treat the humble office computer as a solved problem, yet most office workers still experience frustration when IT helpdesks need to be called. Why does this apparently “solved” problem persist? The software/hardware stack on a standard enterprise computer involves an astounding variety of possible drivers and application versions that can conflict with one another, leading to greater opportunity for breakdown, regardless of skills or resources of IT organizations. This circumstance lends itself to the use of telemetry and artificial intelligence (AI) for problem diagnosis and stokes aspirations of fully automating enterprise PC maintenance. To explore the human and organizational factors at work in applying data and AI to this problem, we designed a series of exploratory studies at a large technology company in the United States: (1) remote diary study with semi-structured interviews (n = 30), (2) quasi-experimental study with pretest-posttest design (n = 11), and (3) ethnographic study with open-ended interviews (n = 8). The results show that user frustration with malfunctioning PCs persisted because of the sociotechnical dynamic between employees, PCs, and IT support. Feedback loops between employees and IT played a central role in dialing up or tamping down frustration that accumulated over the long term. The results also indicate that telemetry and AI could provide new opportunities to tamp down user frustration when data were treated as a communication medium between employees and IT support. The results suggest three major design recommendations for preventing frustration buildup: (1) Redesigning PC telemetry data and transparency mechanisms to support two-way communication between IT and users, including shared analysis of malfunctioning data; (2) considering users’ buildup of frustration, not just the quality of any single interaction, when designing any IT service solutions; (3) incorporating uses of technology that embrace human-AI collaboration technologies not to automate IT troubleshooting work but to support the human creativity necessary for troubleshooting. Utilizing the design principles we identified in this study, there is a need for further research and development to explore novel feedback systems between enterprise PC users and IT.
dAwn nAFus so it made sense to them that my ethnographic research had shown that self-trackers were good at contextualizing their data and did not always take the messages from tracking devices at face value.In fact, one self-tracker in our early research was so careful with his data, both mathematically and conceptually, that one of the computer scientists on the team exclaimed, "This guy's my hero!" Indeed, others on the development team were themselves self-trackers.These confirmations from multiple perspectives softened my anticipation of being told I was crazy yet again, which had come to take on a more stinging valence over the years, being both a woman and a nonengineer in a male-dominated tech company.In these circumstances, I felt a deep sense of responsibility to get it right.This first daylong team meeting turned to discussion of the types of data that would be ingested into the system-strings of text?Numbers?URLs?Enum (or answers to multiple-choice questions)?Something called a float?I didn't understand why we needed to anticipate all these things beforehand, but I took it on trust that we did and secretly wondered to myself what it could possibly mean for a number to float.It became important to know whether time series data would be the only kind of data that we would need to bring into this system."Dawn, do self-trackers collect any data that isn't time series?"I had no idea.I could tell them all about the politics of governmentality and measurement, but whether only time series data was in play, I had no idea.I could guess that time series was probably the most common data type-keeping track involved a notion of time, after all-but whether other data types were of note was a mystery.This prompted a debate about what we meant by data "types"-enum and float or heart rate and blood pressure?-andhow much definition we needed at this early stage.At some point, our newly hired full-stack 1 developer said that we didn't have to solve all the corner cases just yet.I had to ask about this one.If I didn't understand the design and development process from the beginning, there was no way I'd be of much use."What's a corner case?"I asked.Quizzical looks turned my way.It turns out that defining a corner case was something I was expected to be able to contribute to.It is a scenario of use that is rare but could be necessary to take into account when designing a system.Surely it was my job to define what those are, so exposing myself as not even having the language for it would have earned me the label of "crazy" in a less forgiving team.Later I would come to understand this forgiveness not only as a sign that I was finally working in an inclusive team but also as a sign that we were not building a data visualization tool, but rather an infrastructure where the unknowns would proliferate and put everyone on equally unsteady footing.Exploratory data analysis required us to contemplate design considerations that ranged from database issues (in tech industry speak, low in the stack 2 ) all the way through to considerations that involved no technology at all, up so high it was above the stack entirely.Exploratory data analysis complicates any naive understanding of what serves as infrastructure.It is a sociotechnical process
While extensive research has gone into demand response techniques in data centers, the energy consumed in edge computing systems and in network data transmission remains a significant part of the computing industry’s carbon footprint. The industry also has not fully leveraged the parallel trend of decentralized renewable energy generation, which creates new areas of opportunity for innovation in combined energy and computing systems. Through an interdisciplinary sociotechnical discussion of current energy, computer science and social studies of science and technology (STS) literature, we argue that a more comprehensive set of carbon response techniques needs to be developed that span the continuum of data centers, from the back-end cloud to the network edge. Such techniques need to address the combined needs of decentralized energy and computing systems, alongside the social power dynamics those combinations entail. We call this more comprehensive range “carbon-responsive computing,” and underscore that this continuum constitutes the beginnings of an interconnected infrastructure, elements of which are data-intensive and require the integration of social science disciplines to adequately address problems of inequality, governance, transparency, and definitions of “necessary” tasks in a climate crisis.
Data aggregations are an under-acknowledged site of social relations. The social and technical specifics of how data aggregate are arenas for rich debates about how knowledge ought to be produced, and who should produce it. Consumer goods such as fitness trackers create conditions where data scientists or professional researchers are no longer the only ones making decisions about how to aggregate data. Users of these products also rework their data to discover something medically significant to them. These practices call attention to a modality of ‘scaling up’ datasets about a single person that is different from, and until recently largely invisible to, clinical approaches to big data, which privilege the creation of a ‘bird’s eye’ view across as many people. Both technical questions how to build these aggregations, and social questions of who should be involved, betray broader epistemological issues about how new knowledge is created from electronic devices.
This chapter contributes to debates about the social life of methods by describing the co-evolution of a research practice and infrastructure. It offers straightforward description of an ethnographic practice where I collect sensor data with participants, such as heart rate or ambient temperature, rework it into interpretable form, and sit down together with them to co-produce its meaning. The chapter reflects on how that method evolved with the available technical infrastructure, and argues that taking the social life of methods seriously might also mean intervening directly in the tools that sustain cultures of big data.