A growing body of work in Ethical AI attempts to capture human moral judgments through simple computational models. The key question we address in this work is whether such simple AI models capture the critical nuances of moral decision-making by focusing on the use case of kidney allocation. We conducted twenty interviews where participants explained their rationale for their judgments about who should receive a kidney. We observe participants: (a) value patients' morally-relevant attributes to different degrees; (b) use diverse decision-making processes, citing heuristics to reduce decision complexity; (c) can change their opinions; (d) sometimes lack confidence in their decisions (e.g., due to incomplete information); and (e) express enthusiasm and concern regarding AI assisting humans in kidney allocation decisions. Based on these findings, we discuss challenges of computationally modeling moral judgments as a stand-in for human input, highlight drawbacks of current approaches, and suggest future directions to address these issues.
Wage theft—the underpayment or nonpayment of workers’ wages and benefits by employers—is pervasive in the US and abroad, adversely affecting the lives and livelihoods of millions of people annually. Although academics, advocacy groups, and investigative journalists have made advances in documenting the pervasiveness and severity of wage theft practices across states, nations, and industries, research has yet to identify and characterize the processes that make the public see such practices as legitimate. Across four well-powered studies (total N = 2291), we leverage theory and research on moral identity and on evolutionary approaches to moral values to investigate a possible underlying psychological antecedent of the judged legitimacy of wage theft. We propose that the value placed on a widely shared and moralized principle—loyalty—may (ironically) underlie the legitimization of wage theft practices. Although people often consider loyalty to be a positive moral principle or virtue that ought to be valued and exemplified in social and business relations, we find consistent evidence that placing more value on loyalty closely tracks stronger beliefs that wage theft practices are legitimate. Ultimately, differences in the valuation of loyalty may help to explain and predict judgments about the legitimacy of wage theft practices across individuals, groups, organizations, and cultures.
Prior work shows that people are often more sensitive to moral transgressions that target ingroup members than outgroup members. But does that depend on which groups are involved? We investigate how lifelong U.S. citizen participants make judgments about moral transgressions that target fellow lifelong citizens, compared with refugees or undocumented immigrants. Across five studies (N = 1,953), we find that participants overall judge moderate transgressions targeting refugees and undocumented immigrants to be more wrong than those targeting fellow lifelong citizens. This pattern emerges specifically for moderate-severity transgressions but occurs across physical harm, emotional harm, deception, fairness, and property violations. Responses are predicted by political orientation; more liberal participants show the pattern more than conservative participants. We find mediational and experimental evidence for perceived vulnerability/welfare and sympathy toward groups as partial mechanisms: People judge it to be worse to harm more victims they perceive to be more vulnerable. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Recent AI work trends towards incorporating human-centric objectives, with the explicit goal of aligning AI models to personal preferences and societal values. Using standard preference elicitation methods, researchers and practitioners build models of human decisions and judgments, which are then used to align AI behavior with that of humans. However, models commonly used in such elicitation processes often do not capture the true cognitive processes of human decision making, such as when people use heuristics to simplify information associated with a decision problem. As a result, models learned from people's decisions often do not align with their cognitive processes, and can not be used to validate the learning framework for generalization to other decision-making tasks. To address this limitation, we take an axiomatic approach to learning cognitively faithful decision processes from pairwise comparisons. Building on the vast literature characterizing the cognitive processes that contribute to human decision-making, and recent work characterizing such processes in pairwise comparison tasks, we define a class of models in which individual features are first processed and compared across alternatives, and then the processed features are then aggregated via a fixed rule, such as the Bradley-Terry rule. This structured processing of information ensures such models are realistic and feasible candidates to represent underlying human decision-making processes. We demonstrate the efficacy of this modeling approach in learning interpretable models of human decision making in a kidney allocation task, and show that our proposed models match or surpass the accuracy of prior models of human pairwise decision-making.
Theoretical debates have raged around whether conscious perception is necessary for responsibility. It is still unclear, however, what lay people think, and lay views can be important to legal and sociopolitical decision-making. To explore this issue, the current work conducted three online, vignette-based studies to test how lay third-party responsibility judgments varied with what agents unconsciously and consciously visually perceived when deciding how to act. The findings showed that, for both good and bad outcomes, people judge conscious perception not to be necessary for responsibility: an agent was still judged to be at least partially responsible without having consciously perceived pertinent information about how to act appropriately. However, conscious perception did modulate judgments about degrees of responsibility: insofar as the information was perceptually available and accurate, the agent was judged to be more responsible for the outcome when they had consciously perceived pertinent information compared to when they only unconsciously perceived it. For bad outcomes, this effect was mediated by judgments about whether the agent should and could have consciously perceived pertinent information. These findings are interpreted within current theories of consciousness and responsibility and provide insights into how the public may judge someone as responsible for real-world successes and wrongdoing.
How we should design and interact with social artificial intelligence depends on the socio-relational role the AI is meant to emulate or occupy. In human society, relationships such as teacher-student, parent-child, neighbors, siblings, or employer-employee are governed by specific norms that prescribe or proscribe cooperative functions including hierarchy, care, transaction, and mating. These norms shape our judgments of what is appropriate for each partner. For example, workplace norms may allow a boss to give orders to an employee, but not vice versa, reflecting hierarchical and transactional expectations. As AI agents and chatbots powered by large language models are increasingly designed to serve roles analogous to human positions - such as assistant, mental health provider, tutor, or romantic partner - it is imperative to examine whether and how human relational norms should extend to human-AI interactions. Our analysis explores how differences between AI systems and humans, such as the absence of conscious experience and immunity to fatigue, may affect an AI's capacity to fulfill relationship-specific functions and adhere to corresponding norms. This analysis, which is a collaborative effort by philosophers, psychologists, relationship scientists, ethicists, legal experts, and AI researchers, carries important implications for AI systems design, user behavior, and regulation. While we accept that AI systems can offer significant benefits such as increased availability and consistency in certain socio-relational roles, they also risk fostering unhealthy dependencies or unrealistic expectations that could spill over into human-human relationships. We propose that understanding and thoughtfully shaping (or implementing) suitable human-AI relational norms will be crucial for ensuring that human-AI interactions are ethical, trustworthy, and favorable to human well-being.
Preference elicitation frameworks feature heavily in the research on participatory ethical AI tools and provide a viable mechanism to enquire and incorporate the moral values of various stakeholders. As part of the elicitation process, surveys about moral preferences, opinions, and judgments are typically administered only once to each participant. This methodological practice is reasonable if participants' responses are stable over time such that, all other relevant factors being held constant, their responses today will be the same as their responses to the same questions at a later time. However, we do not know how often that is the case. It is possible that participants' true moral preferences change, are subject to temporary moods or whims, or are influenced by environmental factors we don't track. If participants' moral responses are unstable in such ways, it would raise important methodological and theoretical issues for how participants' true moral preferences, opinions, and judgments can be ascertained. We address this possibility here by asking the same survey participants the same moral questions about which patient should receive a kidney when only one is available ten times in ten different sessions over two weeks, varying only presentation order across sessions. We measured how often participants gave different responses to simple (Study One) and more complicated (Study Two) repeated scenarios. On average, the fraction of times participants changed their responses to controversial scenarios was around 10-18% across studies, and this instability is observed to have positive associations with response time and decision-making difficulty. We discuss the implications of these results for the efficacy of moral preference elicitation, highlighting the role of response instability in causing value misalignment between stakeholders and AI tools trained on their moral judgments.
Abstract When two people need a kidney transplant, but only one kidney is available, we need to decide who gets it. If one of the potential recipients needs the kidney because of their own voluntary behavior, but the other is not at all responsible for needing a kidney, then we need to decide whether this fault should be a consideration in favor of the other patient getting the kidney. While there has been considerable philosophical debate on this issue, there is far less research into the views of the public. To explore opinions on these issues, we first asked survey participants to ascribe or deny responsibility for the risky behavior for kidney disease, and for being deprived of a kidney in cases of drinking alcohol, drug abuse, smoking and unhealthy eating when the patient did or did not stop the risky behavior after being diagnosed with kidney disease. Next, we asked participants who should get the kidney when the patient who engaged in risky behavior did or did not know, or have easy access to, the information that the behavior was risky. We found that participants generally ascribed responsibility on the basis of knowledge but allocated the kidney on the basis of access to information. These findings have important implications for moral theories as well as medical policies.
Computational preference elicitation methods are tools used to learn people's preferences quantitatively in a given context. Recent works on preference elicitation advocate for active learning as an efficient method to iteratively construct queries (framed as comparisons between context-specific cases) that are likely to be most informative about an agent's underlying preferences. In this work, we argue that the use of active learning for moral preference elicitation relies on certain assumptions about the underlying moral preferences, which can be violated in practice. Specifically, we highlight the following common assumptions (a) preferences are stable over time and not sensitive to the sequence of presented queries, (b) the appropriate hypothesis class is chosen to model moral preferences, and (c) noise in the agent's responses is limited. While these assumptions can be appropriate for preference elicitation in certain domains, prior research on moral psychology suggests they may not be valid for moral judgments. Through a synthetic simulation of preferences that violate the above assumptions, we observe that active learning can have similar or worse performance than a basic random query selection method in certain settings. Yet, simulation results also demonstrate that active learning can still be viable if the degree of instability or noise is relatively small and when the agent's preferences can be approximately represented with the hypothesis class used for learning. Our study highlights the nuances associated with effective moral preference elicitation in practice and advocates for the cautious use of active learning as a methodology to learn moral preferences.
When making substituted judgments for incapacitated patients, surrogates often struggle to guess what the patient would want if they had capacity. Surrogates may also agonize over having the (sole) responsibility of making such a determination. To address such concerns, a Patient Preference Predictor (PPP) has been proposed that would use an algorithm to infer the treatment preferences of individual patients from population-level data about the known preferences of people with similar demographic characteristics. However, critics have suggested that even if such a PPP were more accurate, on average, than human surrogates in identifying patient preferences, the proposed algorithm would nevertheless fail to respect the patient's (former) autonomy since it draws on the 'wrong' kind of data: namely, data that are not specific to the individual patient and which therefore may not reflect their actual values, or their reasons for having the preferences they do. Taking such criticisms on board, we here propose a new approach: the Personalized Patient Preference Predictor (P4). The P4 is based on recent advances in machine learning, which allow technologies including large language models to be more cheaply and efficiently 'fine-tuned' on person-specific data. The P4, unlike the PPP, would be able to infer an individual patient's preferences from material (e.g., prior treatment decisions) that is in fact specific to them. Thus, we argue, in addition to being potentially more accurate at the individual level than the previously proposed PPP, the predictions of a P4 would also more directly reflect each patient's own reasons and values. In this article, we review recent discoveries in artificial intelligence research that suggest a P4 is technically feasible, and argue that, if it is developed and appropriately deployed, it should assuage some of the main autonomy-based concerns of critics of the original PPP. We then consider various objections to our proposal and offer some tentative replies.
The DSM–5 characterizes mental disorders as significant disturbances in cognition, emotion, or behavior. But what might unite the disturbances on this list? We hypothesize that mental disorders can all be meaningfully characterized as failures of attention. We understand these as failures to distribute attention in the way one has most reason to, and we include both failures of tendency and of ability. We discuss six examples of mental disorders and offer a preliminary gloss of how to recast each as centrally involving a failure of attention. We close by highlighting theoretical and practical upshots of our proposal.
In response to the pressing challenge of kidney allocation, characterized by growing demands for organs, this research sets out to develop a data-driven solution to this problem, which also incorporates stakeholder values. The primary objective of this study is to create a method for learning both individual and group-level preferences pertaining to kidney allocations. Drawing upon data from the 'Pairwise Kidney Patient Online Survey.' Leveraging two distinct datasets and evaluating across three levels - Individual, Group and Stability - we employ machine learning classifiers assessed through several metrics. The Individual level model predicts individual participant preferences, the Group level model aggregates preferences across participants, and the Stability level model, an extension of the Group level, evaluates the stability of these preferences over time. By incorporating stakeholder preferences into the kidney allocation process, we aspire to advance the ethical dimensions of organ transplantation, contributing to more transparent and equitable practices while promoting the integration of moral values into algorithmic decision-making.
This cross-cultural study compared judgments of moral wrongness for physical and emotional harm with varying combinations of in-group vs. out-group agents and victims across six countries: the United States of America (N = 937), the United Kingdom (N = 995), Romania (N = 782), Brazil (N = 856), South Korea (N = 1776), and China (N = 1008). Consistent with our hypothesis we found evidence of an insider agent effect, where moral violations committed by outsider agents are generally considered more morally wrong than the same violations done by insider agents. We also found support for an insider victim effect where moral violations that were committed against an insider victim generally were seen as more morally wrong than when the same violations were committed against an outsider, and this effect held across all countries. These findings provide evidence that the insider versus outsider status of agents and victims does affect moral judgments. However, the interactions of these identities with collectivism, psychological closeness, and type of harm (emotional or physical) are more complex than what is suggested by previous literature.
In his admirable review, Ballantyne characterizes intellectual humility (IH) as a personal way 'to manage evidence horizontal ellipsis in seeking truth.' However, not every way of managing truth is virtuous. Since IH is supposed to be an intellectual virtue, we propose that IH should be understood as a 'golden mean' or 'middle path' between extremes of intellectual arrogance and lack of self-confidence (or between dogmatism and gullibility). The golden mean should not be characterized descriptively by the statistical mean of a population but instead either epistemically by accuracy in intellectual assessments of oneself and others or pragmatically by the kinds of such assessments that enable or lead to successful inquiry. This comment explains and considers advantages and disadvantages of these two ways of locating the golden mean.
Click to increase image sizeClick to decrease image size Disclosure statementNo potential conflict of interest was reported by the author.