
Transfer learning leverages knowledge gained from previous tasks to accelerate learning in related target tasks. In robotics, and especially in Human-Robot Interaction, this capability is crucial due to the scarcity and high cost of collecting social interaction data. Using transfer learning, a robot can learn new tasks faster and with less data. Successor Representations (SR) have traditionally been used to transfer knowledge between tasks with shared environment dynamics but differing reward functions. In this work, we propose a novel decomposition of SR, Modular Successor Representations (MSR) that facilitates transfer between tasks where only a subset of the environment dynamics changes, while others remain invariant. We evaluate MSR in a multi-agent Social Navigation scenario in simulation and show that it reduces the amount of social data required for training. Finally, we discuss remaining challenges, including scaling to high-dimensional continuous state spaces and handling dynamic social behaviors.
As cybersecurity threats have become increasingly complex, conventional rule-based detection systems struggle to keep pace. Existing cybersecurity training simulators rely on manually scripted scenarios and rigid operational flows, resulting in limited scalability and low robustness against novel attack strategies. To overcome these challenges, this study proposes a cyber-attack scenario generation framework based on an agent-based system. The framework is composed of two attack agents that construct threat scenarios and an evaluation agent that analyzes the effectiveness of the generated scenarios. Our fine-tuned language model agent generates technically coherent scenarios rooted in domain knowledge, whereas our prompt-engineered large language model attack agent produces flexible and diverse scenarios through structured reasoning. The evaluation agent supports objective and multi-perspective validation, ensuring both consistency and feasibility. The agent-based framework offers a scalable foundation for generating practical threat simulations that are aligned with real-world complexities.
The Influence Maximization (IM) problem is a fundamental problem on social networks where you are required to choose a set of few seeds from which to start an information campaign aiming to reach as many nodes as possible in the network. In this work, we consider the IM problem in a setting where neither network nodes nor their relationships are known, except for very few samples. Thus, you have to orchestrate the campaign while learning information about the network. This problem has been recently showed to have applications in public health: e.g., to maximize the diffusion of HIV prevention information among marginalized people, such as homeless. In this work we propose a two-level bandit approach to address the IM problem with partially observed networks: the lower layer implements a contextual bandit that selects nodes to query based on the current observed subgraph, available nodes, and edge discovery rewards; the upper meta-layer dynamically chooses between two exploration strategies: a global approach maximizing immediate edge discovery, and a component-focused strategy targeting the least-explored connected component. This dual approach prevents local over-exploitation while maintaining efficient global exploration. The proposed method outperforms the state-of-the-art method and shows robustness across diverse network topologies.
This paper proposes a Stackelberg game model to simulate the strategic interaction between electric vehicle (EV) users and charging stations in an EV charging market. In our model, charging stations act as leaders by setting prices first, and EV users act as followers, selecting stations to minimize their total cost in response to these prices. We show that the competition among EV users admits a unique symmetric mixed-strategy Nash equilibrium for any price profile, and that the EV charging game admits a unique Stackelberg equilibrium. In our experiments, we find that charging prices and waiting times are the primary factors influencing EV users’ choices, while travel distance tends to be less decisive when stations are evenly distributed.
This study constructs an artificial market simulation incorporating large language models (LLMs) as decision-making agents and evaluates their ability to reproduce market stylized facts. LLMs have been widely applied across various domains, including finance, where they can process texts and outputs in a manner similar to humans. Through training on extensive corpora, these models are capable of generating data that accords with average thought patterns. In this research, we propose utilizing LLMs for decision-making processes in artificial market simulations. Previous research has employed simulations that manually constructed traders’ algorithms incorporating fundamental, chartist (trend), and noise factors. However, our simulation design replaces the fundamental and chartist factors with LLM decision-making, while retaining noise factors. For our LLM agent prompts, we created eight prompt patterns, which include three optional elements ( 2^3 patterns): prospect theory prompt, position prompt, and fundamental prompt. We then analyzed the performance effects of each optional prompt. Our experimental results demonstrate that appropriate prompt design can successfully reproduce market stylized facts. Notably, all three prompt elements–fundamentals, position, and prospect theory prompts–were necessary to reproduce the stylized facts of financial markets. It means that effective prompt engineering plays a crucial role in enhancing the realism of artificial markets. This study not only showcases the potential applications of LLMs in financial market research but also provides important insights into actual LLM usages for finance.
Argumentation games, which model reasoning as adversarial dialogue, offer intuitive and explainable mechanisms for decision-making in AI. However, their implementation has lagged behind inference-focused approaches, particularly in structured argumentation frameworks like assumption-based argumentation (ABA). This work presents, to our knowledge, the first application of multi-shot answer set programming (ASP) for implementing argument games, focusing on ABA dispute derivations. Leveraging a recent rule-based representation of ABA disputes, our method combines a declarative program with lightweight script-based control of multi-shot aspects, yielding a modular and adaptable system. We extend this core approach to support alternative games and show how it can also be used to implement argument games for Dung’s abstract argumentation formalism. Empirical results show that our implementation outperforms existing ABA dispute systems. We also introduce an approximate variant that further improves efficiency – reaching the level of the best inference-focused ABA system – while maintaining perfect specificity (true negative rate), demonstrating the practical value of multi-shot ASP, especially in interactive, explainable settings.
Designing effective and compact state representations is a key challenge in applying Deep Reinforcement Learning (DRL) to combinatorial problems such as the Electric Vehicle Routing Problem (EVRP). Current approaches often include extensive handcrafted features to ensure constraint coverage, but this can lead to high-dimensional inputs that slow down learning, reduce generalization, and obscure policy behavior. In this work, we propose a novel methodology that leverages Explainable Artificial Intelligence (XAI), specifically SHapley Additive exPlanations (SHAP), to guide feature pruning in DRL-based EVRP agents. We begin by benchmarking commonly used features in state representations for EVRP and apply SHAP to evaluate their relative importance throughout training. Correlation analysis is used alongside SHAP scores to identify redundant or low-impact features. The pruned state representations are then retrained and compared against baseline models using the full feature set. Results show that agents trained on XAI-pruned features achieve comparable or improved performance in terms of average travel distance, model convergence, training stability, and loss, all while reducing computational cost and enhancing interpretability. This study demonstrates that explainability can go beyond post-hoc analysis to actively inform and optimize the design of DRL agents. While evaluated in the context of EVRP, the proposed methodology is generalizable to other structured decision-making tasks where DRL is applied. Our findings suggest that XAI-guided feature design is a promising direction for building more efficient, transparent, and adaptable reinforcement learning systems.
Public interest in ranked choice voting mechanisms such as Single Transferable Vote (STV) has increased in recent years, as a potentially more equitable approach, leading to its adoption for political elections in various cities, states, and countries. A recently-studied measure of interest is metric distortion, which captures how much worse an elected candidate is from the socially optimal choice. Building upon the known lower bound of 3 on the metric distortion of any deterministic voting rule shown by Anshelevich et al. [3], and the known upper bound of 15 on the metric distortion of STV on the line by Anagnostides et al. [2], we improve the gap between these bounds by providing an upper bound of 11 on the metric distortion of STV on the line. In addition, we consider the impact of voter turnout on elections, giving a lower bound on the metric distortion of STV on the line that is a function of the voter turnout percentage; in particular, 50
Arabic Sign Language (ArSL) presents formidable communication barriers for 17–23 million deaf individuals across 460 million Arabic speakers. We introduce a neural-symbolic framework addressing dialectal variation through spatiotemporal constraint injection, embedding grammatical rules as optimizable symbolic loss terms. Our multi-agent coordination system employs symbolic validators that inject dialect-specific rules with physics-informed motion synthesis to generate culturally authentic gestures. The framework achieves 40 , ”clear newspaper”) versus ṣuḥuf ( , ”multiple newspapers”)—demonstrating the model’s capacity for morphological decomposition while maintaining cultural coherence.
Estimating the outcome of a negotiation before it is finished allows a party to take effective actions, e.g., exploring outside options, or reporting progress to a human user. However, estimating the outcome is difficult as many (uncertain) factors affect the course of a negotiation. Accordingly, this paper presents a method for predicting the outcome of ongoing bilateral negotiations called PrONeg. We predict the future trajectories of an agent’s own bids and its opponent’s bids using time series forecasting methods. These forecasts are used to determine the agent’s outcome utility distribution, along with the probability of reaching an agreement by the end of the negotiation. Finally, we predict the most likely outcome of the negotiation by combining the outcome utility distribution with preference information available in the negotiation scenario. Our experiments show that Gaussian processes perform best in most settings, including balancing predicting true breakoffs without misclassifying agreements. With its ability to predict the outcome of a negotiation, PrONeg can potentially serve as a negotiation support system in hybrid negotiations.
Dec-POMDPs model cooperative, sequential multi-agent decision problems. They are computationally challenging, and scaling up their performance is difficult. We describe a method for solving Dec-POMDPs in the paradigm of centralized planning with distributed execution. First, we solve a team POMDP in which agent observations are common knowledge. Then, each agent uses imitation learning to try and imitate its part of the centralized policy. Unlike some previous work, the agent not only tries to imitate its behavior within the team, but also its belief state. A final offline synchronization stage improves the likelihood that agents’ policies will be well-coordinated with each other. On standard Dec-POMDP benchmarks, our method performs better than the best Dec-POMDP model-based solution method, and QMIX, a leading multi-agent RL algorithm.
Colored Node Kayles is a combinatorial game played on an undirected graph G=(V,E) with vertex colors from black, gray, white. Two players, Black and White, alternate turns: Black selects a gray or black vertex, while White selects a gray or white vertex. The chosen vertex and its neighbors are then removed from G. The game continues until no valid moves remain, with the last player able to move declared the winner. Node Kayles, the special case of Colored Node Kayles with only gray vertices, has been extensively studied. Due to its simplicity and generality, results for Node Kayles have been applied to a broad range of other combinatorial games, though the partisan variant Colored Node Kayles has received far less attention. A restricted version, called Bigraph Node Kayles, was implicitly introduced by Schaefer in his seminal 1978 work, where one side of a bipartite graph is colored black and the other white. We formally define Colored Node Kayles and study the complexity of deciding the winner. We prove that it is PSPACE-complete even on planar graphs with maximum degree 3. We also show W[1]-hardness with respect to the number of turns and present other hardness results, including for computing game values. On the algorithmic side, Colored Node Kayles is FPT concerning graph structural parameters such as clique deletion number, neighborhood diversity, vertex cover, and twin cover, and is solvable in polynomial time on graphs with bounded vertex integrity or cluster deletion number.
Contemporary Arabic narrative generation systems often fail to capture nuanced cultural authenticity, primarily addressing dialectal variation as a lexical challenge rather than encoding deeper cultural schemas. The distinction between (Egyptian) and (Levantine) for “The story began in Cairo” exemplifies cultural resonance patterns that standardized approaches overlook. We present CULTURA, a neural-symbolic framework orchestrating three specialized agents: (1) CNN-based dialect classification (94.3 p<0.001 .
Defeasible reasoning has always been a central interest of researchers in the fields of Artificial Intelligence (AI) and Multi-Agent Systems (MAS). In fact, this kind of reasoning is central to dealing with conflicting knowledge or beliefs that agents may hold without causing inconsistencies. In the context of languages for Knowledge Representation, many formal approaches have been proposed specifically in Description Logics (DLs) to deal with this phenomenon. With a perspective towards human-centred and agentive AI and building on the DL paradigm we pursue an approach informed by results coming from fields such as linguistics, philosophy and cognitive science. A central problem in the general area of defeasible DLs is to give a principled solution to the question of where preferences originate from in order to provide a notion of defeasibility. To address this issue, a core aspect of our approach is to compute preferences from the knowledge presented in a knowledge base (with standard semantics) itself. We thus present a non-monotonic DL based on a combination of ideas from prototype theory, weighted DLs (aka ‘tooth logic’), and earlier work on justifiable exceptions. A central ingredient in the new framework is the notion of a prototype description, i.e. weighted characterisations of concepts based on the typical features of its members. We show that through such descriptions it is possible to compute a typicality score which allows to define a preference order over models, useful to solve conflicts across exceptional instances. We define two principle ways of computing such preferences, discuss some core semantic properties and finally outline a translation into Answer Set Programming.
Autonomous Agents (AAs) intrinsically have to exploit some forms of causal reasoning, i.e., the ability to understand what actions can bring about intended effects. However, such reasoning is not usually grounded in formally defined and homogeneous causal models, but is instead implicitly represented within the AA model itself. In this paper, we discuss the role that sound causal modelling and learning can play in conceiving and developing AAs and Multi-Agent Systems: reasoning activities would be made available by a uniform, explicit model, that would be amenable to autonomous manipulation by the AAs. Open challenges towards achieving this goal are also discussed.
Urban Air Mobility leverages electric vertical takeoff and landing aircraft to relieve ground congestion and improve passenger and cargo transport within dense cities. Integrating UAM into existing airspace safely and efficiently requires ground control systems that deliver real-time situational awareness, risk assessment, and coordinated command across multiple vehicles. We introduce a distributed multi-agent architecture in which each vehicle functions as an autonomous agent, publishing its position, velocity, attitude, and flight plan over a scalable messaging layer. A recurrent neural network enhanced with attention and sinusoidal positional encoding consumes streaming data in an Earth-Centered, Earth-Fixed frame to produce multi-step trajectory predictions. Compared to the attention-only variant, positional encoding reduces mean squared error from 0.0525 to 0.0517, mean absolute error from 0.1680 to 0.1659, and root mean squared error from 0.2291 to 0.2273. Testing on 23 unseen UAM flights yields an overall RMSE of 0.1897 and an average inference time of 0.86 s, meeting real-time requirements. The system uses these predictions to compute three-dimensional separations and classify collision risk into Safe, Caution, Warning, or Collision categories, achieving 0.9881 classification accuracy. Throughput reaches 171.06 batches per second with end-to-end latency of 7.0 ms (95th percentile 6.5 ms), supporting both streaming and batch modes. To ensure interoperability and resilience, we employ the Model Context Protocol for standardized metadata exchange, synchronized state management, and distributed inference coordination. This protocol enables dynamic agent discovery, context-aware routing, and fault tolerance under network disruptions. By combining high-fidelity deep learning predictions with a protocol-driven multi-agent framework, our platform advances safe, scalable, and compliant UAM ground control in complex urban airspace.
Trust and reputation assessment plays a major role in multi-agent systems to enable seamless and useful interaction among agents. The assessment faces significant challenges due to the volatility of the environment as the agents evolve over time leading to colluding interactions. Such an ecosystem necessitates that the agents perform adaptive learning and have robust privacy mechanisms to ensure that the Trust and reputation thus assessed are context-aware as well. This paper presents a novel and holistic framework that addresses the challenges of privacy and adaptability by uniquely integrating Reinforcement Learning (RL) with Distributed Online Life-Long Learning (DOL3), graph clustering, and Graph Neural Networks (GNN). RL is induced to optimize trust learning strategies within DOL3, enhancing the adaptability of both collusion detection (through Graph Clustering) and privacy-preserving trust score learning (through GNN with differential privacy). This synergistic integration makes the framework more resilient to dynamic threats, and RL actively optimizes how trust is learned and applied within the DOL3 framework along with continuous privacy protection. This, essentially, offers enhanced adaptability, accuracy, and privacy-protection collusion detection. We demonstrate the potential of this framework to improve decision-making in complex domains such as autonomous systems, healthcare, and e-commerce.
Creativity has been a main target of artificial intelligence since its beginning and is still a major challenge today. One well-known option in computational creativity research is conceptual blending where new concepts are created through a selective combination of known ideas. However, existing approaches implementing blending are often either neglecting conceptual aspects, as in image morphing, or they suffer from the high complexity of the creation of the blend. Therefore, we propose a new neuro-symbolic approach to conceptual blending of ontologies based on knowledge graph embeddings which addresses both of these shortcomings. The inherent structure of the embedding space is used both to identify a generic space and to guide the blending process by interpreting blending as path search in the embedding space and by iteratively relaxing the input concepts. This is accomplished by combining a symbolic system for determining a step-wise refinement and analyzing the suitability of these refinements with the help of the embedding. We give an overview on the method and showcase possible heuristics.
Recent advances in generative models enabled the creation of high-quality synthetic text and images, raising concerns about provenance and misuse. We propose MW-MAS (Multimodal Watermarking Multi-Agent System), a unified framework that orchestrates watermarking across text and image modalities via three agents: the Text Watermark Agent, Image Watermark Agent, and Orchestration Agent. The Orchestration Agent adaptively selects optimal agent combinations based on sample characteristics. Evaluated on the WIT dataset, MW-MAS achieves up to 2 × faster runtime than dual-agent baselines while maintaining high fidelity and robust bit-level watermark retrieval, offering a flexible and practical solution for multimodal content watermarking. Code is available at https://github.com/lynnchoi0126/MW-MAS .
This paper presents a cognitively enhanced medical scheduling framework that combines Answer Set Programming (ASP) with L-DINF, an epistemic logic-based agent model. While ASP and Blueprint Personas capture constraints and preferences, they lack adaptability. L-DINF agents overcome this by reasoning about beliefs, intentions, and actions, enabling adaptive and explainable decision-making (Research partially supported by the PNRR Project CUP E13C24000430006 “Enhanced Network of intelligent Agents for Building Livable Environments - ENABLE”, and by PRIN 2022 CUP E53D23007850001 Project TrustPACTX - Design of the Hybrid Society Humans-Autonomous Systems: Architecture, Trustworthiness, Trust, EthiCs, and EXplainability (the case of Patient Care), and by PRIN PNNR CUP E53D23016270001 ADVISOR - ADaptiVe legIble robotS for trustwORthy health coaching).