Some advances in reproductive technologies raise substantial ethical and psychological challenges, for example, the use of preimplantation genetic testing to select embryos based on non-medical traits. While studies have explored public willingness to use preimplantation genetic testing for medical or non-medical attributes, stated willingness may not reflect implantation decisions when such information is available. This large cross-national study examined public views on polygenic embryo testing. In a US sample (N = 1,467), participants were more willing to test for medical conditions (for example, heart disease) than non-medical traits (for example, antisocial behaviour or low intelligence), although over half supported testing for non-medical traits. In a forced-choice implantation task, participants used medical and non-medical information to a similar extent when making decisions. Similar patterns were observed in Chinese participants (N = 623). These findings suggest variability in choices across specific traits and conditions, rather than a uniform distinction between medical conditions and non-medical traits. The Stage 1 protocol for this Registered Report was accepted in principle on 7 October 2025. The protocol, as accepted by the journal, can be found at https://osf.io/vb9c2 .
The substantial literature on charitable giving–a prototypical pro-social behavior–commonly faces constraints on sample size, diversity, or experimental design, resulting ina fragmented understanding of factors that influence charitable giving and theirinteractions. We developed a game that co-varies multiple factors that influencegiving—akin to running thousands of experiments. Over 2.7 million decisions, someincentivized, from 257,000 participants in 203 countries were collected. Helping at least 3or 6 strangers, on average, surpassed even the strongest effects of giving to oneself or arelative, respectively. This finding holds cross-culturally. Further, we found substantialheterogeneity from experimental stimuli in all main factor effects (e.g., identifiable victimeffect), often reversing direction. Our findings paint a more comprehensive and promisingpicture of human prosociality.
Life expectancy, access to education, household income, perceived safety, and life satisfaction are all increasing globally. While these advancements are worth celebrating, progress is not distributed evenly or equitably. Moreover, human progress, according to other metrics, has also come at huge cost for the wider collective – our non-human collectives, ecosystems, and the planet. The current paper argues for more inclusive, comprehensive, and interdisciplinary approaches to promote mutually sustaining flourishing. We propose an ecological-collective-flourishing (e-co-flourishing) approach whereby one’s own human flourishing is pursued along with that of natural ecosystems and their collective and individual constituents. We offer a pragmatic framework that includes: a methodological approach; guiding principles; and proposed programme theory in terms of hypothesised outcomes, pathways of change (mechanisms), key contextual factors (moderators), and intervention characteristics. This framework aims to bridge theory to real-world applications of e-co-flourishing interventions that are effective, scalable, and promote change through target mechanisms of action. We present Mindfulness-Based Cognitive Therapy – Balanced Living for Us and the Earth’s Ocean (MBCT-BLUE) as a case study to help outline programme theory and how this framework can be used to identify, refine, develop, and evaluate interventions with e-co-flourishing potential. We also provide definitions of key conceptual terms and proxy measurements. Practitioners, researchers, and decision-makers working at the interface of mental health, public health, sustainability, and planetary health can utilise this framework and its resources to guide emerging research and innovation, with the aim of addressing interconnected crises impacting both human and non-human entities and collectives.
Cyberbullying (CB) has emerged as a growing concern among adolescents, with nearly 10% of European children affected monthly and almost half experiencing it at least once. Unlike traditional bullying, CB thrives in digital environments where anonymity and impunity are prevalent. Despite its increasing prevalence, understanding the causal mechanisms behind CB remains challenging due to the limitations of conventional statistical methods, which often rely on correlations and are prone to spurious associations. In this paper, we introduce a novel human-machine consensus framework for causal discovery, aimed at supporting social scientists in unraveling the complex dynamics of CB. We leverage recent advances in data-driven causal inference, particularly the use of Directed Acyclic Graphs (DAGs), to identify and interpret causal relationships from observational data. Our approach integrates automatic causal discovery algorithms with expert knowledge, addressing the limitations of both purely algorithmic and purely expert-driven methods, and allows for the creation of a model ensemble estimation of the causal effects. To enhance interpretability and usability, we advocate for the use of Probabilistic Graphical Causal Models (PGCMs), or Bayesian Networks, which combine probabilistic reasoning with graphical representation. This hybrid methodology not only mitigates cognitive biases and inconsistencies in expert input but also fosters transparency and critical reflection in model construction. Cyberbullying serves as a compelling case study where ethical constraints preclude experimental designs, highlighting the value of interpretable, expert-informed causal models for guiding policy and intervention strategies.
Artificial Intelligence (AI) has potential to address the increasing demand for capacity in Air Traffic Control (ATC). However, its integration poses several challenges and requires deep understanding of public perception. Insights from the context of Autonomous Vehicles (AVs), in which more studies have been done, can inform such understanding. In this article, we investigate how the public perceives the automated future of ATC in close comparison to AVs. We conducted two studies to examine public trust and blame attribution toward human and AI operators in different Human-AI Interaction (HAI) models, covering three levels of automation (Level 0: AI tool, Level 3: AI trainee, and Level 5: AI manager). We also explored their perceptions of ATC and vehicle driving (VD) by using ten task-related measures (Familiarity, Expertise, Tech Awareness, Openness, Media Discourse, Stake, two measures of Uncertainty, Positive Safety, and Negative Safety) and five agent-related characteristics (Capability, Robustness, Predictability, Honesty and Cooperativeness). The results showed greater trust and less blame attributed to humans in both ATC and VD, except in the Level 3 AI trainee model where humans were blamed more than AI. We also found both similarities and differences in people’s perceptions of the two contexts. Our findings provide evidence-based insights into how the public attribute trust and blame to the operators in ATC and VD. These results will inform industries on the development and implementation of AI integration in aviation and advise policymakers in evaluating public opinion on AI regulation.
How we should design and interact with social artificial intelligence depends on the socio-relational role the AI is meant to emulate or occupy. In human society, relationships such as teacher-student, parent-child, neighbors, siblings, or employer-employee are governed by specific norms that prescribe or proscribe cooperative functions including hierarchy, care, transaction, and mating. These norms shape our judgments of what is appropriate for each partner. For example, workplace norms may allow a boss to give orders to an employee, but not vice versa, reflecting hierarchical and transactional expectations. As AI agents and chatbots powered by large language models are increasingly designed to serve roles analogous to human positions - such as assistant, mental health provider, tutor, or romantic partner - it is imperative to examine whether and how human relational norms should extend to human-AI interactions. Our analysis explores how differences between AI systems and humans, such as the absence of conscious experience and immunity to fatigue, may affect an AI's capacity to fulfill relationship-specific functions and adhere to corresponding norms. This analysis, which is a collaborative effort by philosophers, psychologists, relationship scientists, ethicists, legal experts, and AI researchers, carries important implications for AI systems design, user behavior, and regulation. While we accept that AI systems can offer significant benefits such as increased availability and consistency in certain socio-relational roles, they also risk fostering unhealthy dependencies or unrealistic expectations that could spill over into human-human relationships. We propose that understanding and thoughtfully shaping (or implementing) suitable human-AI relational norms will be crucial for ensuring that human-AI interactions are ethical, trustworthy, and favorable to human well-being.
This paper explores how humans make contextual moral judgments to inform the development of AI systems capable of balancing rule-following with flexibility. We investigate the limitations of rigid constraints in AI, which can hinder morally acceptable actions in specific contexts, unlike humans who can override rules when appropriate. We propose a preference-based graphical model inspired by dual-process theories of moral judgment and conduct a study on human decisions about breaking the social norm of "no cutting in line." Our model outperforms standard machine learning methods in predicting human judgments and offers a generalizable framework for modeling moral decision-making across various contexts. This short paper summarizes the main findings of our paper published in the journal Autonomous Agents and Multi-Agent Systems. [2]
Artificial Intelligence (AI) advancements might deliver autonomous agents capable of human-like deception. Such capabilities have mostly been negatively perceived in HCI design, as they can have serious ethical implications. However, AI deception might be beneficial in some situations. Previous research has shown that machines designed with some level of dishonesty can elicit increased cooperation with humans. This raises several questions: Are there future-of-work situations where deception by machines can be an acceptable behaviour? Is this different from human deceptive behaviour? How does AI deception influence human trust and the adoption of deceptive machines? In this paper, we describe the results of a user study published in the proceedings of AAMAS 2023. The study answered these questions by considering different contexts and job roles. Here, we contextualise the results of the study by proposing ways forward to achieve a framework for developing Deceptive AI responsibly. We provide insights and lessons that will be crucial in understanding what factors shape the social attitudes and adoption of AI systems that may be required to exhibit dishonest behaviour as part of their jobs.
Education systems are dynamically changing to accommodate technological advances, industrial and societal needs, and to enhance students' learning journeys. Curriculum specialists and educators constantly revise taught subjects across educational grades to identify gaps, introduce new learning topics, and enhance the learning outcomes. This process is usually done within the same subjects (e.g. math) or across related subjects (e.g. math and physics) considering the same and different educational levels, leading to massive multi-layer comparisons. Having nuanced data about subjects, topics, and learning outcomes structured within a dataset, empowers us to leverage data science to better understand the progression of various learning topics. In this paper, Bidirectional Encoder Representations from Transformers (BERT) topic modeling was used to extract topics from the curriculum, which were then used to identify relationships between subjects, track their progression, and identify conceptual gaps. We found that grouping learning outcomes by common topics helped specialists reduce redundancy and introduce new concepts in the curriculum. We built a dashboard to avail the methodology to curriculum specials. Finally, we tested the validity of the approach with subject matter experts.
Constraining the actions of AI systems is one promising way to ensure that these systems behave in a way that is morally acceptable to humans. But constraints alone come with drawbacks as in many AI systems, they are not flexible. If these constraints are too rigid, they can preclude actions that are actually acceptable in certain, contextual situations. Humans, on the other hand, can often decide when a simple and seemingly inflexible rule should actually be overridden based on the context. In this paper, we empirically investigate the way humans make these contextual moral judgements, with the goal of building AI systems that understand when to follow and when to override constraints. We propose a novel and general preference-based graphical model that captures a modification of standard dual process theories of moral judgment. We then detail the design, implementation, and results of a study of human participants who judge whether it is acceptable to break a well-established rule: no cutting in line. We then develop an instance of our model and compare its performance to that of standard machine learning approaches on the task of predicting the behavior of human participants in the study, showing that our preference-based approach more accurately captures the judgments of human decision-makers. It also provides a flexible method to model the relationship between variables for moral decision-making tasks that can be generalized to other settings.
This study investigates how people assign blame to autonomous vehicles (AVs) when involved in an accident. Our experiment (N = 2647) revealed that people placed more blame on AVs than on human drivers when accident details were unspecified. To examine whether people assess major classes of blame-relevant information differently for AVs and humans, we developed a causal model and introduced a novel concept of prevention effort, which emerged as a crucial factor or blame judgement alongside intentionality. Finally, we addressed the “many hands” problem by exploring how people assign blame to entities associated with AVs and human drivers, such as the car company or an accident victim. Our findings showed that people assigned high blame to these entities in scenarios involving AVs, but not with human drivers. This necessitates adapting a model of blame for AVs to include other agents and thus allow for blame allocation “outside” of autonomous vehicles.
Synthetic data generation has been a growing area of research in recent years. However, its potential applications in serious games have not been thoroughly explored. Advances in this field could anticipate data modelling and analysis, as well as speed up the development process. To try to fill this gap in the literature, we propose a simulator architecture for generating probabilistic synthetic data for serious games based on interactive narratives. This architecture is designed to be generic and modular so that it can be used by other researchers on similar problems. To simulate the interaction of synthetic players with questions, we use a cognitive testing model based on the Item Response Theory framework. We also show how probabilistic graphical models (in particular Bayesian networks) can be used to introduce expert knowledge and external data into the simulation. Finally, we apply the proposed architecture and methods in a use case of a serious game focused on cyberbullying. We perform Bayesian inference experiments using a hierarchical model to demonstrate the identifiability and robustness of the generated data.
Introducing automated vehicles (AVs) on roads may challenge established norms as drivers of human-driven vehicles (HVs) interact with AVs. Our study explored drivers’ decisions in game-theoretical scenarios amid mixed traffic using an online survey study. We manipulated factors including interaction types (HV-HV vs. HV-AV), scenario types (chicken game vs. public goods game), vehicle driving styles (aggressive vs. conservative), and time constraints (high vs. low). The quantitative results showed that human drivers tended to “defect” more, that is, not cooperate, against vehicles with conservative driving styles. The effect of vehicle driving styles was pronounced when interacting with AVs and in chicken game scenarios. Drivers exhibited more “defection” in public goods game scenarios and the effect of scenario types was weakened under high time constraints. Only drivers with moderate driving styles “defected” more in HV-AV interaction. Our qualitative findings provide essential insights into how drivers perceived conditions and formulated strategies for decision-making.
Machines powered by artificial intelligence have the potential to replace or collaborate with human decision-makers in moral settings. In these roles, machines would face moral tradeoffs, such as automated vehicles (AVs) distributing inevitable risks among road users. Do people believe that machines should make moral decisions differently from humans? If so, why? To address these questions, we conducted six studies (N = 6805) to examine how people, as observers, believe human drivers and AVs should act in similar moral dilemmas and how they judge their moral decisions. In pedestrian-only dilemmas where the two agents had to sacrifice one pedestrian to save more pedestrians, participants held them to similar utilitarian norms (Study 1). In occupant dilemmas where the agents needed to weigh the in-vehicle occupant against more pedestrians, participants were less accepting of AVs sacrificing their passenger compared to human drivers sacrificing themselves (Studies 1-3) or another passenger (Studies 5-6). The difference was not driven by reduced occupant agency in AVs (Study 4) or by non-voluntary occupant sacrifice in AVs (Study 5), but rather by the perceived social relationship between AVs and their users (Study 6). Thus, even when people adopt an impartial stance as observers, they are more likely to believe that AVs should prioritize serving their users in moral dilemmas. We discuss the theoretical and practical implications for AV morality.
Introducing automated vehicles (AVs) on public roads may challenge established norms as drivers in human-driven vehicles (HVs) learn to interact with AVs. Our study utilizes a game theory framework to investigate how driver and vehicle driving styles (aggressive vs. conservative), interaction types (HV-HV vs. HV-AV), scenario types (chicken game vs. public goods game scenarios), and time constraints (high vs. low) influence human drivers’ decision-making in mixed-traffic environments. According to an online survey study, drivers with aggressive driving styles and high time constraints were more likely to take aggressive actions. More importantly, there were significant interaction effects between vehicle driving styles and scenario types, between scenario types and time constraints, and between driver driving styles and interaction types on driver decision-making. Our findings provide essential insights into the design of AVs and promote the development of related laws and policies to facilitate human-machine cooperation in mixed-traffic environments.
Cyberbullying among minors is a pressing concern in our digital society, necessitating effective prevention and intervention strategies. Traditional data collection methods often intrude on privacy and yield limited insights. This study explores an innovative approach, employing a serious game - designed with purposes beyond entertainment - as a non-intrusive tool for data collection and education. In contrast to traditional correlation-based analyses, we propose a causality-based approach using Bayesian Networks to unravel complex relationships in the collected data and quantify result uncertainties. This robust analytical tool yields interpretable outcomes, enhances transparency in assumptions, and fosters open scientific discourse. Preliminary pilot studies with the serious game show promising results, surpassing the informative capacity of traditional demographic and psychological questionnaires, suggesting its potential as an alternative methodology. Additionally, we demonstrate how our approach facilitates the examination of risk profiles and the identification of intervention strategies to mitigate this cybercrime. We also address research limitations and potential enhancements, considering the noise and variability of data in social studies and video games. This research advances our understanding of cyberbullying and showcase the potential of serious games and causality-based approaches in studying complex social issues.
The use of Artificial Intelligence (AI) in hiring is becoming increasingly popular, but little is known about the direct impact of its use on applicant decisions. In a series of novel economic experiments using a representative US sample (N=1,002), we provide some of the first comprehensive causal evidence on how the use of AI and debiasing affect the quality and gender diversity of applicants. We study application decisions for competitive jobs in two experiments where participants are faced with different evaluators: human, AI, debiased human, and debiased AI. Overall, we find that the use of AI does not affect the quality and gender diversity of applicants compared to human evaluators, whereas debiasing (whether human or AI) increases gender diversity without reducing the number of high quality applicants. Our findings suggest that firms with diversity goals wishing to use AI in hiring could do so without hindering such goals, as long as their algorithm is debiased.
Do people hold robots responsible for their actions? While Clark and Fischer present a useful framework for interpreting social robots, we argue that they fail to account for people's willingness to assign responsibility to robots in certain contexts, such as when a robot performs actions not predictable by its user or programmer.
Artificial Intelligence (AI) advancements might deliver autonomous agents capable of human-like deception. Such capabilities have mostly been negatively perceived in HCI design, as they can have serious ethical implications. However, AI deception might be beneficial in some situations. Previous research has shown that machines designed with some level of dishonesty can elicit increased cooperation with humans. This raises several questions: Are there future-of-work situations where deception by machines can be an acceptable behaviour? Is this different from human deceptive behaviour? How does AI deception influence human trust and the adoption of deceptive machines? In this paper, we describe a user study to answer these questions by considering different contexts and job roles. We report differences and similarities with the perception of humans behaving deceptively in the same roles. Our findings provide insights and lessons that will be crucial in understanding what factors shape the social attitudes and adoption of AI systems that may be required to exhibit dishonest behaviour as part of their jobs.
Ensuring safe and clean drinking water for communities is crucial, and necessitates effective tools to monitor and predict water quality due to challenges from population growth, industrial activities, and environmental pollution. This paper evaluates the performance of multiple linear regression (MLR) and nineteen machine learning (ML) models, including algorithms based on regression, decision tree, and boosting. Models include linear regression (LR), least angle regression (LAR), Bayesian ridge chain (BR), ridge regression (Ridge), k-nearest neighbor regression (K-NN), extra tree regression (ET), and extreme gradient boosting (XGBoost). The research’s objective is to estimate the surface water quality of Al-Seine Lake in Lattakia governorate using the MLR and ML models. We used water quality data from the drinking water lake of Lattakia City, Syria, during years 2021–2022 to determine the water quality index (WQI). The predictive performance of both the MLR and ML models was evaluated using statistical methods such as the coefficient of determination (R2) and the root mean square error (RMSE) to estimate their efficiency. The results indicated that the MLR model and three of the ML models, namely linear regression (LR), least angle regression (LAR), and Bayesian ridge chain (BR), performed well in predicting the WQI. The MLR model had an R2 of 0.999 and an RMSE of 0.149, while the three ML models had an R2 of 1.0 and an RMSE of approximately 0.0. These results support using both MLR and ML models for predicting the WQI with very high accuracy, which will contribute to improving water quality management.