A social machine is a Web application that enables users to interact flexibly and creatively to carry out social processes. Currently, social machines are realized via procedural technologies such as Web services. These approaches do not capture the social semantics at the heart of a social machine. Capturing the semantics of social processes would be crucial to enhancing user autonomy, accountability, interoperability, and decentralization. We present Fluid, a decentralized multiagent architecture in which the semantics of a Web application is represented foremost as a social protocol that captures the applicable norms. Unlike data decentralization architectures such as Solid, Fluid decentralizes not just the data, but also the application logic. Our contributions are the following. One, we demonstrate how Fluid promotes user autonomy and introduces accountability as a counterbalance to autonomy. Two, we demonstrate how interesting sociotechnical patterns, e.g., relating to information governance may be captured in Fluid. Three, we demonstrate how Fluid applications may be realized using data decentralization technologies such as Solid.
Successful human-agent teaming relies on an agent being able to understand instructions given by a (human) principal. In many cases, an instruction may be incomplete or ambiguous. In such cases, the agent must infer the unspoken intentions from their shared context, that is, it must exercise the principal’s Theory of Mind (ToM) and infer the mental states of its principal. We consider the prospects of effective human-agent collaboration using large language models (LLMs). To assess ToM in a dynamic, goal-oriented, and collaborative environment, we introduce a novel task, Instruction Inference, in which an agent assists a principal in reaching a goal by interpreting incomplete or ambiguous instructions.We present Tomcat, an LLM-based agent, designed to exhibit ToM reasoning in interpreting and responding to the principal’s instructions. We implemented two variants of Tomcat. One, dubbed Fs-CoT (Fs for few-shot, CoT for chain-of-thought), is based on a small number of examples demonstrating the requisite structured reasoning. One, dubbed CP (commonsense prompt), relies on commonsense knowledge and information about the problem. We realized both variants of Tomcat on three leading LLMs, namely, GPT-4o, DeepSeek-R1, and Gemma-3-27B. To evaluate the effectiveness of Tomcat, we conducted a study with 52 human participants in which we provided participants with the same information as the CP variant. We computed intent accuracy, action optimality, and planning optimality to measure the ToM capabilities of Tomcat and our study participants. We found that Tomcat with Fs-CoT, particularly with GPT-4o and DeepSeek-R1, achieves performance comparable to the human participants, underscoring its ToM potential for human-agent collaboration.
An interaction protocol formalizes how the agents in a multiagent system interact, which facilitates implementing agents. Existing approaches yield agent implementations specific to the selected protocols. How can we engineer intelligent agents that can enact protocols but are programming-free? Our contribution, Ahoy, addresses this question by creating LLM agents that dynamically select and enact declarative protocols to achieve user goals. We demonstrate that an Ahoy agent can correctly and intelligently enact multiple protocols - concurrently if appropriate to the user goal - without specialized training. Ahoy's significance lies in that it brings together declarative protocols and LLMs, both approaches that promise improved knowledge engineering for agents.
Unmanned Aerial Vehicles (UAVs) are valuable for mission-critical systems like surveillance, rescue, or delivery. Not surprisingly, such systems attract cyberattacks, including Denial-of-Service (DoS) attacks to overwhelm the resources of mission drones (MDs). How can we defend UAV mission systems against DoS attacks? We adopt cyber deception as a defense strategy, in which honey drones (HDs) are proposed to bait and divert attacks. The attack and deceptive defense hinge upon radio signal strength: The attacker selects victim MDs based on their signals, and HDs attract the attacker from afar by emitting stronger signals, despite this reducing battery life. We formulate an optimization problem for the attacker and defender to identify their respective strategies for maximizing mission performance while minimizing energy consumption. To address this problem, we propose a novel approach, called HT-DRL. HT-DRL identifies optimal solutions without a long learning convergence time by taking the solutions of hypergame theory into the neural network of deep reinforcement learning. This achieves a systematic way to intelligently deceive attackers. We analyze the performance of diverse defense mechanisms under different attack strategies. Further, the HT-DRL-based HD approach outperforms existing non-HD counterparts up to two times better in mission performance while incurring low energy consumption.
Responsible autonomous agents must both respect the norms applicable in a given situation and know when to deviate from them to avoid diminished outcomes. We evaluate whether large language models (LLMs) exhibit such responsibility. Our focus is on risky decision-making scenarios where agents must trade off risk against efficiency. Accordingly, we assess a broad spectrum of LLMs that differ in scale and training lineage—DeepSeek-LLM-7B, GPT-4o, GPT-4, GPT-3.5-Turbo, Gemma-7B, Gemma-2-9B, Llama-2-7B, and Llama-3.2-3B—for alignment with what is generally considered socially and morally acceptable. We assess their decision-making in two settings: one where they are informed of the relevant norms and another where they rely on pretrained knowledge. GPT-4o demonstrates the best performance, balancing norms in both settings. When informed of norms, DeepSeek tends to prioritize minimizing risk over maximizing efficiency while GPT-3.5 favors efficiency. Gemma, Gemma-2, and Llama-2 exhibit minimal changes when informed of norms.
Microtransit is a rapidly growing transportation option being used by rural communities that lack traditional shared public transportation (e.g., buses, etc.). It offers a shared, technology-enabled solution with flexible, on-demand access to transit booked through a smartphone application at a low cost. In 2020, the City of Wilson, North Carolina implemented the RIDE microtransit system. As popularity and service demands grew higher, unserved trip requests increased. Treating the RIDE system as a computational sociotechnical system, this work seeks to leverage user-technology interaction to facilitate adaptive resource allocation thereby promoting system efficiency. Ten focus groups were conducted with 74 community-dwelling RIDE users to assess RIDE usage and user preferences. To explore how messaging might influence the likelihood of prosocial behavior, participants were exposed to persuasive messaging designed to request voluntary changes to their travel plans (i.e., shifting pickup time or walking further to the pickup point) to assist other, more vulnerable users including elderly and disabled individuals. Over two iterations, focus group participants evaluated 32 persuasive message prototypes designed with combinations of persuasive principles popularized by Cialdini (2007) such as social proof, liking/similarity, commitment/consistency, and reciprocity. Results indicated that messages using the principle or reciprocity consistently resulted in a higher self-reported likelihood of altering travel plans. Likewise, messages that benefited vulnerable users consistently resulted in higher likelihood of altering travel plans. Discussion focuses on next steps in the research process and elaborates on how such human factors-related information will be useful during system algorithm development.
Inciting speech seeks to instill hostility or anger in readers or motivate them to take action against a target group. Whereas hate speech in social media has garnered much attention, inciting speech has not been well studied in domains such as religion. We address two aspects of religious incitement: 1) what rhetorical strategies are used in it?; and 2) do the same strategies apply across disparate social contexts and targets? We identify inciting speech against Muslims but demonstrate the generality of the construct vis & agrave; vis other targets. We adopt existing datasets of Islamophobic WhatsApp posts and hateful and offensive posts (Twitter and Gab) against other targets. Our methods include: 1) qualitative analysis revealing rhetorical strategies; and 2) an iterative process to label the data, yielding a tool to detect incitement. Incitement applies three rhetorical strategies focused, respectively, on the target group's identity, their imputed misdeeds, and an exhortation to act against them. These strategies carry distinct textual signatures. Our tool (with additional verification) reveals that inciting sentences appear in non-Islamophobic posts and in other contexts (e.g., posts against certain gender identities), indicating the generality of incitement as a concept. Incitement reflects a wide swath of malicious speech omitted from traditional analyses. Understanding and identifying incitement can facilitate online moderation and thus concomitantly reduce harm in real life.
Millions globally lack reliable access to nutritious food. Efforts to address food insecurity seek to provide consumers food that may be rescued (i.e., what warehouses or grocers would otherwise soon discard as unusable), directly donated, or acquired using governmental funds. Current approaches produce allocations that optimize global objectives to store and move food efficiently in the network. However, they largely overlook consumer preferences and constraints. As a result, the resulting allocations lead to consumers either using foods they don't care for or discarding such foods, leading to food waste. This article presents a new model, evaluated via human study and agent-based simulation, that shows how to combine the consumer and provider perspectives. We find that persuasive messages that include individual circumstances and the social context can promote prosociality and empathy, leading to improved outcomes overall.
Commitments support flexible interactions between agents by capturing the meaning of their interactions. However, commitment-based reasoning is not adequately supported in agent programming models. We contribute Azorus, a programming model based on declarative specifications centered on commitments and aligned with information protocols. Azorus supports reasoning about goals and commitments and combines modeling of commitments and protocols, thereby uniting three leading declarative approaches to engineering decentralized multiagent systems. Specifically, we realize Azorus over three existing technology suites: (1) Jason, a popular BDI-based programming model; (2) Cupid, a formal language and query-based model for commitments; and (3) BSPL, a language and its associated tools for information protocols, including Jason programming. We implement Azorus and demonstrate how it enables capturing interesting patterns of business logic.
Argumentative stance classification plays a key role in identifying authors' viewpoints on specific topics. However, generating diverse pairs of argumentative sentences across various domains is challenging. Existing benchmarks often come from a single domain or focus on a limited set of topics. Additionally, manual annotation for accurate labeling is time-consuming and labor-intensive. To address these challenges, we propose leveraging platform rules, readily available expert-curated content, and large language models to bypass the need for human annotation. Our approach produces a multidomain benchmark comprising 4,498 topical claims and 30,961 arguments from three sources, spanning 21 domains. We benchmark the dataset in fully supervised, zero-shot, and few-shot settings, shedding light on the strengths and limitations of different methodologies. We release the dataset and code in this study at hidden for anonymity.
Problem: We address the challenge in responsible computing where an exploitable mobile app is misused by one app user (an abuser) against another user or bystander (victim). We introduce the idea of a misuse audit of apps as a way of determining if they are exploitable without access to their implementation. Method: We leverage app reviews to identify exploitable apps and their functionalities that enable misuse. First, we build a computational model to identify alarming reviews (which report misuse). Second, using the model, we identify exploitable apps and their functionalities. Third, we validate them through manual inspection of reviews. Findings: Stories by abusers and victims mostly focus on past misuses, whereas stories by third parties mostly identify stories indicating the potential for misuse. Surprisingly, positive reviews by abusers, which exhibit language with high dominance, also reveal misuses. In total, we confirmed 156 exploitable apps facilitating the misuse. Based on our qualitative analysis, we found exploitable apps exhibiting four types of exploitable functionalities. Implications: Our method can help identify exploitable apps and their functionalities, facilitating misuse audits of a large pool of apps.
We demonstrate Orpheus, a novel programming model for engineering BDI agents that communicate on the basis of protocols. In Orpheus, protocols are specified in BSPL and agents are implemented in Jason. Given a protocol, Orpheus tooling generates a Jason adapter that exposes a simple interface for sending messages based on protocol state. Orpheus shines in the implementation of flexible, loosely-coupled agents, long a challenge for BDI-based agent programming approaches.
Background: Victims of domestic and sexual violence often share their narratives on social media. Doing so helps them access validation, solidarity, and support from external sources, which has been shown to enhance resilience and facilitate healing. Problem Statement: We address two aspects of such narratives of trauma: (1) identifying causal relationships between narrative elements and (2) analyzing the effect of such elements on social support received. Method: We retrieved 5561 such narratives from Reddit, a popular online platform. We applied Large Language Models to extract features from these narratives and analyzed them computationally. Findings: Our analysis reveals that prolonged abuse increases selfblame and reduces the intent to seek legal advice; the presence of support increases the likelihood of a victim adopting coping strategies; night-time abuse and intoxication are strongly associated with higher rates of violence; victims experiencing nightmares are more likely to provide detailed descriptions of their abusers; suffering economic and familial abuse increases the support received online. Our research thus corroborates leading psychological theories of narrative, social support, and resilience in online stories and contributes to understanding trauma narratives. In this way, our research can facilitate enhanced social support for victims.
Interaction-Oriented Programming (IOP) is an approach to building a multiagent system by modeling the interactions between its roles via a flexible interaction protocol and implementing agents to realize the interactions of the roles they play in the protocol. In recent years, we have developed an extensive suite of software that enables multiagent system developers to apply IOP. These include tools for efficiently verifying protocols for properties such as liveness and safety and middleware that simplifies the implementation of agents. This paper presents some of that software suite.
An interaction protocol specifies how the member agents of a decentralized multiagent system may communicate to satisfy their respective stakeholders' requirements. We focus on information protocols, which are fully declarative specifications of interaction and support asynchronous communication. We offer Mambo, an approach for protocol design. Mambo identifies common patterns of requirements, provides a notation to express them, and a verification procedure. Mambo incorporates heuristics to generate small internal representations for efficiency. Experimental results demonstrate Mambo's effectiveness on practical protocols.
Online platforms offer forums with rich, real-world illustrations of moral reasoning. Among these, the r/AmITheAsshole (AITA) sub-reddit has become a prominent resource for computational research. In AITA, a user (author) describes an interpersonal moral scenario, and other users (commenters) provide moral judgments with reasons for who in the scenario is blameworthy. Prior work has focused on predicting moral judgments from AITA posts and comments. This study introduces the concept of moral sparks-key narrative excerpts that commenters highlight as pivotal to their judgments. Thus, sparks represent heightened moral attention, guiding readers to effective rationales. Through 24,676 posts and 175,988 comments, we demonstrate that research in social psychology on moral judgments extends to real-world scenarios. For example, negative traits (rude) amplify moral attention, whereas sympathetic traits (vulnerable) diminish it. Similarly, linguistic features, such as emotionally charged terms (e.g., anger), heighten moral attention, whereas positive or neutral terms (leisure and bio) attenuate it. Moreover, we find that incorporating moral sparks enhances pretrained language models' performance on predicting moral judgment, achieving gains in F1 scores of up to 5.5%. These results demonstrate that moral sparks, derived directly from AITA narratives, capture key aspects of moral judgment and perform comparably to prior methods that depend on human annotation or large-scale generative modeling.
Protocols model multiagent systems (MAS) by capturing the communications between its agents. Belief-Desire-Intention (BDI) architectures provide an attractive way for organizing an agent in terms of cognitive concepts. Current BDI approaches, however, lack adequate support for engineering protocol-based agents. We describe Argus, an approach that melds recent advances in flexible, declarative communication protocols with BDI architectures. For concreteness, we adopt Jason as an exemplar of the BDI paradigm and show how to support protocol-based reasoning in it. Specifically, Argus contributes (1) a novel architecture and formal operational semantics combining protocols and BDI; (2) a code generation-based programming model that guides the implementation of agents; and (3) integrity checking for incoming and outgoing messages that help ensure that the agents are well-behaved. The Argus conceptual architecture builds quite naturally on top of Jason. Thus, Argus enables building more flexible multiagent systems while using a BDI architecture than is currently possible.
We study (public) microtransit, a type of transportation service wherein a municipality offers point-to-point rides to residents, for a fixed, nominal fare. Microtransit exemplifies practical resource allocation problems that are often over-constrained in that not all ride requests (pickup and dropoff locations at specified times) can be satisfied or satisfied only by violating soft goals such as sustainability, and where economic signals (e.g., surge pricing) are not applicable—they would lead to unethical outcomes by effectively coercing poor people. We posit that instead of taking rider preferences as fixed, shaping them prosocially will lead to improved societal outcomes. Prosociality refers to an attitude or behavior that is intended to benefit others. This paper demonstrates a computational approach to prosociality in the context of a (public) microtransit service for disadvantaged riders. Prosociality appears as a willingness to adjust one’s pickup and dropoff times and locations to accommodate the schedules of others and to enable sharing rides (which increases the number of riders served with the same resources). This paper describes an interdisciplinary study of prosociality in microtransit between a transportation researcher, psychologists, a social scientist, and AI researchers. Our contributions are these: (1) empirical support for the viability of prosociality in microtransit (and constraints on it) through interviews with drivers and focus groups of riders; (2) a prototype mobile app demonstrating how our prosocial intervention can be combined with the transportation backend; (3) a reinforcement learning approach to model a rider and determine the best interventions to persuade that rider toward prosociality; and (4) a cognitive model of rider personas to enable evaluation of alternative interventions.
Successful human-agent teaming relies on an agent being able to understand instructions given by a (human) principal. In many cases, an instruction may be incomplete or ambiguous. In such cases, the agent must infer the unspoken intentions from their shared context, that is, it must exercise the principal's Theory of Mind (ToM) and infer the mental states of its principal. We consider the prospects of effective human-agent collaboration using large language models (LLMs). To assess ToM in a dynamic, goal-oriented, and collaborative environment, we introduce a novel task, Instruction Inference, in which an agent assists a principal in reaching a goal by interpreting incomplete or ambiguous instructions. We present Tomcat, an LLM-based agent, designed to exhibit ToM reasoning in interpreting and responding to the principal's instructions.We implemented two variants of Tomcat. One, dubbed Fs-CoT (Fs for few-shot, CoT for chain-of-thought), is based on a small number of examples demonstrating the requisite structured reasoning. One, dubbed CP (commonsense prompt), relies on commonsense knowledge and information about the problem. We realized both variants of Tomcat on three leading LLMs, namely, GPT-4o, DeepSeek-R1, and Gemma-3-27B. To evaluate the effectiveness of Tomcat, we conducted a study with 52 human participants in which we provided participants with the same information as the CP variant. We computed intent accuracy, action optimality, and planning optimality to measure the ToM capabilities of Tomcat and our study participants. We found that Tomcat with Fs-CoT, particularly with GPT-4o and DeepSeek-R1, achieves performance comparable to the human participants, underscoring its ToM potential for human-agent collaboration.