Social and moral norms are a fabric for holding human societies together and helping them to function. As such they will also become a means of evaluating the performance of future human–machine systems. While machine ethics has offered various approaches to endowing machines with normative competence, from the more logic‐based to the more data‐based, none of the proposals so far have considered the challenge of capturing the “spirit of a norm,” which often eludes rigid interpretation and complicates doing the right thing. We present some paradigmatic scenarios across contexts to illustrate why the spirit of a norm can be critical to make explicit and why it exposes the inadequacies of mere data‐driven “value alignment” techniques such as reinforcement learning RL for interactive, real‐time human–robot interaction. Instead, we argue that norm learning, in particular, learning to capture the spirit of a norm, requires combining common‐sense inference‐based and data‐driven approaches.
Recent attention has been brought to robots that "disobey" or so-called "rebel" agents that might reject commands. However, any discussion of autonomous agents that "disobey" risks engaging in a potentially hazardous conflation of simply non-conforming behavior with true disobedience. The goal of this paper is to articulate a sense of what constitutes desirable and true disobedience from autonomous systems. To do this, we begin by discussing what it is not. First, we attempt to disentangle figurative uses of the term "disobedience" from those connotative of deeper senses of agency. We then situate true disobedience as being committed by an agent through an action that presupposes some understanding of the violated instruction or command.
Machine ethics has sought to establish how autonomous systems could make ethically appropriate decisions in the world. While mere statistical machine learning approaches have focused on learning human preferences from observations and attempted actions, hybrid approaches to machine ethics attempt to provide more explicit guidance for robots based on explicit norm representations. Neither approach, however, might be sufficient for real contexts of human-robot interaction, where reasoning and exchange of information may need to be distributed across automated processes and human improvisation, requiring real-time coordination within a dynamic environment (sharing information, trusting in other agents, and arriving at revised plans together). This paper builds on discussions of "extended minds" in philosophy to examine norms as "extended" systems supported by external cues and an agent's own applications of norms in concrete contexts. Instead of locating norms solely as discrete representations within the AI system, we argue that explicit normative guidance must be extended across human-machine collaborative activity as what does and does not constitute a normative context, and within a norm, might require negotiation of incompletely specified or derive principles that not be self-contained, but become accessible as a result of the agent's actions and interactions and thus representable by agents in social space.
Explainability has emerged as a critical AI research objective, but the breadth of proposed methods and application domains suggest that criteria for explanation vary greatly. In particular, what counts as a good explanation, and what kinds of explanation are computationally feasible, has become trickier in light of oqaque “black box” systems such as deep neural networks. Explanation in such cases has drifted from what many philosophers stipulated as having to involve deductive and causal principles to mere “interpretation,” which approximates what happened in the target system to varying degrees. However, such post hoc constructed rationalizations are highly problematic for social robots that operate interactively in spaces shared with humans. For in such social contexts, explanations of behavior, and, in particular, justifications for violations of expected behavior, should make reference to socially accepted principles and norms. In this article, we show how a social robot’s actions can face explanatory demands for how it came to act on its decision, what goals, tasks, or purposes its design had those actions pursue and what norms or social constraints the system recognizes in the course of its action. As a result, we argue that explanations for social robots will need to be accurate representations of the system’s operation along causal, purposive, and justificatory lines. These explanations will need to generate appropriate references to principles and norms—explanations based on mere “interpretability” will ultimately fail to connect the robot’s behaviors to its appropriate determinants. We then lay out the foundations for a cognitive robotic architecture for HRI, together with particular component algorithms, for generating explanations and engaging in justificatory dialogues with human interactants. Such explanations track the robot’s actual decision-making and behavior, which themselves are determined by normative principles the robot can describe and use for justifications.
HRI researchers have made major strides in developing robotic architectures that are capable of reading a limited set of social cues and producing behaviors that enhance their likeability and feeling of comfort amongst humans. However, the cues in these models are fairly direct and the interactions largely dyadic. To capture the normative qualities of interaction more robustly, we propose “consent” as a distinct, critical area for HRI research. Convening important insights in existing HRI work around topics like touch, proxemics, gaze, and moral norms, the notion of consent reveals key expectations that can shape how a robot acts in social spaces. Consent need not be limited to just an explicit permission given in ethically charged or normatively risky scenarios. Instead, it is a richer notion, one that covers even implicit acquiescence in scenarios that otherwise seem normatively neutral. By sorting various kinds of consent through social and legal doctrine, we delineate empirical and technical questions to meet consent challenges faced in major application domains and robotic roles. Attention to consent could show, for example, how extraordinary, norm-violating actions can be justified by agents and accepted by those around them. We argue that operationalizing ideas from legal scholarship can better guide how robotic systems might cultivate and sustain proper forms of consent.
This paper addresses ethical challenges posed by a robot acting as both a general type of system and a discrete, particular machine. Using the philosophical distinction between “type” and “token,” we locate type-token ambiguity within a larger field of indefinite robotic identity, which can include networked systems or multiple bodies under a single control system. The paper explores three specific areas where the type-token tension might affect human–robot interaction, including how a robot demonstrates the highly personalized recounting of information, how a robot makes moral appeals and justifies its decisions, and how the possible need for replacement of a particular robot shapes its ongoing role (including how its programming could transfer to a new body platform). We also consider how a robot might regard itself as a replaceable token of a general robotic type and take extraordinary actions on that basis. For human–robot interaction robotic type-token identity is not an ontological problem that has a single solution, but a range of possible interactions that responsible design must take into account, given how people stand to gain and lose from the shifting identities social robots will present.
In this paper we describe moral quasi-dilemmas (MQDs): situations similar to moral dilemmas, but in which an agent is unsure whether exploring the plan space or the world may reveal a course of action that satisfies all moral requirements. We argue that artificial moral agents (AMAs) should be built to handle MQDs (in particular, by exploring the plan space rather than immediately accepting the inevitability of the moral dilemma), and that MQDs may be useful for evaluating AMA architectures.
As a way to address both ominous and ordinary threats of artificial intelligence (AI), researchers have started proposing ways to stop an AI system before it has a chance to escape outside control and cause harm. A so-called “big red button” would enable human operators to interrupt or divert a system while preventing the system from learning that such an intervention is a threat. Though an emergency button for AI seems to make intuitive sense, that approach ultimately concentrates on the point when a system has already “gone rogue” and seeks to obstruct interference. A better approach would be to make ongoing self-evaluation and testing an integral part of a system’s operation, diagnose how the system is in error and to prevent chaos and risk before they start. In this paper, we describe the demands that recent big red button proposals have not addressed, and we offer a preliminary model of an approach that could better meet them. We argue for an ethical core (EC) that consists of a scenario-generation mechanism and a simulation environment that are used to test a system’s decisions in simulated worlds, rather than the real world. This EC would be kept opaque to the system itself: through careful design of memory and the character of the scenario, the system’s algorithms would be prevented from learning about its operation and its function, and ultimately its presence. By monitoring and checking for deviant behavior, we conclude, a continual testing approach will be far more effective, responsive, and vigilant toward a system’s learning and action in the world than an emergency button which one might not get to push in time.
The challenge of training AI systems to perform responsibly and beneficially has inspired different approaches for teaching a system what people want and how it is acceptable to attain that in the world. In this paper we compare work in reinforcement learning, in particular inverse reinforcement learning, with our norm inference approach. We test those two systems and present results. Using the idea of the "intentional stance", we explain how a norm inference approach can work even when another agent is acting strictly according to reward functions. In this way norm inference presents itself as a promising, more explicitly accountable approach with which to design AI systems from the start.
The complex role of touch is an increasingly appreciated horizon for HRI research. The explicit and implicit registers of touch, both human-to-robot and robot-to-human, have opened up pressing questions in design and HRI ethics about embodiment, communication, care, and human affection. In this paper we present results of an MTurk survey about robot-initiated touch in a social context. We examine how a positive or negative attitude from the robot, as well as whether the robot touches an interactant, affects how a robot is judged as a worker and teammate. Our findings confirm previous empirical support for the idea of touch as enhancing social appraisals of a robot, though the extent of that positive tactile role was complicated and tempered by the survey responses’ gender effects.
Machine learning’s advances have led to new ideas about the feasibility and importance of machine ethics keeping pace, with increasing emphasis on safety, containment, and align- ment. This paper addresses a recent suggestion that inverse reinforcement learning (IRL) could be a means to so-called “value alignment.” We critically consider how such an approach can engage the social, norm-infused nature of ethical action and outline several features of ethical appraisal that go beyond simple models of behavior, including unavoidably temporal dimensions of norms and counterfactuals. We propose that a hybrid approach for computational architectures still offers the most promising avenue for machines acting in an ethical fashion.
HRI research has yielded intriguing empirical results connected to ethics and how we act in social contexts with robots, even though much of this work has focused on task-based, one-on-one interaction. In this paper, we point to the need to investigate a wider range of ethically relevant dynamics that interaction with robots carries with it -- individually and in groups, with a single robot or more. We specifically examine three areas: 1) the primacy and implicit dynamics of bodily perception, 2) the competing interests at work in a single robot-human interaction, and 3) the social intricacy of multiple agents -- robots and humans -- communicating and making decisions. While these areas are not exhaustive by any means, we find they yield concrete directions for how HRI can contribute to a widening, intensifying set of ethical debates with critical empirical insight, starting to explore more of the ethical landscape in HRI.
Soft robots promise an exciting design trajectory in the field of robotics and human-robot interaction (HRI), promising more adaptive, resilient movement within environments as well as a safer, more sensitive interface for the objects or agents the robot encounters. In particular, tactile HRI is a critical dimension for designers to consider, especially given the onrush of assistive and companion robots into our society. In this article, we propose to surface an important set of ethical challenges for the field of soft robotics to meet. Tactile HRI strongly suggests that soft-bodied robots balance tactile engagement against emotional manipulation, model intimacy on the bonding with a tool not with a person, and deflect users from personally and socially destructive behavior the soft bodies and surfaces could normally entice.
Robots designed for sexual interaction present distinctive ethical challenges to received notions of physical intimacy, pleasure, social relationships, and social space. In this chapter, we build upon our recent survey on attitudes toward sex robots with the results from a second, expanded survey that broaches possible advantages and disadvantages of interacting with such robots, both individually and socially. We show that the first study’s results were replicated with respect to appropriate forms, contexts, and uses for sex robots; in addition, we find a systematic concern with how robots might risk harming human relationships. We conclude that ethical reflection on sex robots must include a wider consider-ation of the impact of social robots as a whole, with finer-grained examination of how intimacy and companionship define human relationships.
Collaborative human activities are grounded in social and moral norms, which humans consciously and subconsciously use to guide and constrain their decision-making and behavior, thereby strengthening their interactions and preventing emotional and physical harm. This type of norm-based processing is also critical for robots in many human-robot interaction scenarios (e.g., when helping elderly and disabled persons in assisted living facilities, or assisting humans in assembly tasks in factories or even the space station). In this position paper, we will briefly describe how several components in an integrated cognitive architecture can be used to implement processes that are required for normative human-robot interactions, especially in collaborative tasks where actions and situations could potentially be perceived as threatening and thus need a change in course of action to mitigate the perceived threats.
Much existing work examining the ethical behaviors of robots does not consider the impact and effects of long- term human-robot interactions. A robot teammate, col- laborator or helper is often expected to increase task performance, individually or of the team, but little dis- cussion is usually devoted to how such a robot should balance the task requirements with building and main- taining a “working relationship” with a human partner, much less appropriate social relations outside that team. We propose the “Relational Enhancement” framework for the design and evaluation of long-term interactions, which composed of interrelated concepts of efficiency, solidarity, and prosocial concern. We discuss how this framework can be used to evaluate common existing ap- proaches in cognitive architectures for robots and then examine how social norms and mental simulation may contribute to each of the components of the framework.
This paper argues against the moral Turing test (MTT) as a framework for evaluating the moral performance of autonomous systems. Though the term has been carefully introduced, considered, and cautioned about in previous discussions (Allen et al. in J Exp Theor Artif Intell 12(3):251–261, 2000; Allen and Wallach 2009), it has lingered on as a touchstone for developing computational approaches to moral reasoning (Gerdes and Øhrstrøm in J Inf Commun Ethics Soc 13(2):98–109, 2015). While these efforts have not led to the detailed development of an MTT, they nonetheless retain the idea to discuss what kinds of action and reasoning should be demanded of autonomous systems. We explore the flawed basis of an MTT in imitation, even one based on scenarios of morally accountable actions. MTT-based evaluations are vulnerable to deception, inadequate reasoning, and inferior moral performance vis a vis a system’s capabilities. We propose verification—which demands the design of transparent, accountable processes of reasoning that reliably prefigure the performance of autonomous systems—serves as a superior framework for both designer and system alike. As autonomous social robots in particular take on an increasing range of critical roles within society, we conclude that verification offers an essential, albeit challenging, moral measure of their design and performance.