Indirect reciprocity, supported by simple reputation assessment and social norms, has been demonstrated as an effective mechanism for enabling cooperation in populations of self-interested individuals. However, it has been shown that where there is noise in the performance of actions, or in observers’ perceptions, cooperation may not emerge. Higher-order social norms and generosity have been investigated as potential mechanisms to support cooperation in such environments, but are ineffective without additional, limiting, assumptions. In particular, higher-order norms have typically been investigated in cases where reputation is binary (‘good’ or ‘bad’) and where all agents ascribe the same reputation to an individual, implying full and perfect observation of actions. Generosity with an ‘aligned’ strategy, where all agents have the same likelihood of being generous, has been shown to be ineffective where reputation is binary. In this work we consider reciprocity emergence mechanisms in noisy domains and where agent observations may be incomplete or inaccurate. Our hypothesis is that nuanced reputation scores will enable generosity to be more effective, since an individual act of generosity will have a less extreme impact on reputation. We also investigate whether replacing the ‘aligned’ setting for generosity with a ‘non-aligned’ alternative, which we refer to as forgiveness, will support cooperation in noisy partially observable environments, without the level of ‘unjustified benevolence’ exhibited by generosity. We show both analytically and empirically that generosity when combined with fine grained reputation can help cooperation emerge, and that forgiveness can support cooperation in certain settings.
AI systems are becoming increasingly complex, ubiquitous and autonomous, leading to increasing concerns about their impacts on individuals and society. In response, researchers have begun investigating how to ensure that the methods underlying AI decision-making are transparent and their decisions are explainable to people and conformant to human values and ethical principles. As part of this research thrust, the need for accountability within AI systems has been noted, but this notion has proven elusive to define; we aim to address this issue in the current paper. Unlike much recent work, we do not address accountability within the human organisational processes of developing and deploying AI; rather we consider what it would it mean for the agents within a multi-agent system (MAS), potentially including human agents, to be accountable to other agents or to have others accountable to them. In this work, we make the following contributions: we provide an in-depth survey of existing work on accountability in multiple disciplines, seeking to identify a coherent definition of the concept; we give a realistic example of a multi-agent system application domain that illustrates the benefits of enabling agents to follow accountability processes, and we identify a set of research challenges for the MAS community in building accountable agents, sketching out some initial solutions to these, thereby laying out a road-map for future research. Our focus is on laying the groundwork to enable autonomous elements within open socio-technical systems to take part in accountability processes.
Many semantics for abstract weighted argumentation assume that each argument is associated with a numerical initial weight. Eliciting these initial weights poses several challenges: (1) accurately providing a specific numerical value is often difficult, and (2) individuals frequently confuse initial weights with acceptability degrees in the presence of other arguments. We therefore propose an elicitation pipeline that allows a user to specify their believed final acceptability degree intervals for each argument. We can determine which portion (if any) of these intervals are rational, refining the intervals, or restoring rationality when the intervals are irrational. This allows us to ultimately identify possible initial weights for each argument.
Preference-based argumentation frameworks (PAFs) extend Dung's approach to abstract argumentation (AAFs) by encoding preferences over arguments. Such preferences control the transformation of attacks into defeats, and different approaches to doing so result in different reductions from a PAF to an AAF. In this paper we consider a PAF inverse problem which takes an argumentation graph, a labelling and a semantics as an input, and outputs a “yes" or “no" as to whether there is a preference relation between the arguments which can yield the desired labelling. This inverse problem has applications in areas including preference elicitation and explainability. We consider this problem in the context of the four most widely-used preference based reductions under the complete semantics. We show that in most cases, the problem can be answered in polynomial time.
Large language models have recently reached near-parity with classical planners on well-known planning domains, yet this competence relies on world-knowledge exploitation rather than genuine symbolic reasoning. Goal recognition is a complementary abductive task structurally better suited to LLM strengths: it consists of evaluating consistency with world knowledge rather than generating novel action sequences. This paper provides the first systematic zero-shot evaluation of frontier LLMs as goal recognisers on key classical PDDL benchmarks. Our results show that LLM competence on goal recognition is uneven: some models scale with evidence and approach landmark-based accuracy at full observations, while others remain anchored to world-knowledge priors regardless of how much evidence accumulates. Qualitative analysis of model reasoning traces reveals that this divergence reflects a fundamental difference in evidence integration rather than domain familiarity. These findings position goal recognition as a principled benchmark for the foundational planning knowledge of LLMs.
Weighted gradual semantics provide an acceptability degree to each argument representing its final strength, computed based on factors including the argument's background evidence, and taking into account interactions between the argument and others. We introduce five important problems linking gradual semantics and acceptability degrees. First, we re-examine the inverse problem, seeking to identify each argument's initial weights within the argumentation framework which lead to a specific final acceptability degree. Second, we ask whether the function mapping between argument weights and acceptability degrees is one-to-one. Third, we ask if this mapping is a homeomorphism so that small perturbations in weights lead to small perturbation in acceptability degrees and vice versa. Fourth, we ask whether argument weights can be found when preferences, rather than acceptability degrees for arguments are considered. Last, we consider the geometry of the space of valid acceptability degrees, asking whether ``"gaps" exist in this space. While different gradual semantics have been proposed in the literature, in this paper, and building on the geometry of the acceptability degree space, we identify a large family of weighted gradual semantics which contains many of the existing well-known semantics while maintaining desirable properties such as convergence to a unique fixed point and solving all five aforementioned problems.
Responsibility plays a key role in the development and deployment of trustworthy autonomous systems. In this paper, we focus on the problem of strategic reasoning in probabilistic multi-agent systems with responsibility-aware agents. We introduce the logic PATL+R, a variant of Probabilistic Alternating-time Temporal Logic. The novelty of PATL+R lies in its incorporation of modalities for causal responsibility, providing a framework for responsibility-aware multi-agent strategic reasoning. We present an approach to synthesise joint strategies that satisfy an outcome specified in PATL+R, while optimising the share of expected causal responsibility and reward. This provides a notion of balanced distribution of responsibility and reward gain among agents. To this end, we utilise the Nash equilibrium as the solution concept for our strategic reasoning problem and demonstrate how to compute responsibility-aware Nash equilibrium strategies via a reduction to parametric model checking of concurrent stochastic multi-player games.
Agent interpreters based on the Beliefs, Desires, and Intentions (BDI) model traditionally perform means-ends reasoning using plan libraries composed of reactive planning rules. However, the design of such rules often imposes a heavy knowledge engineering burden on a designer, and trades off flexibility for runtime efficiency. This use of planning rules originates from the limitations of planning technology at the time of the first BDI implementations. While these limitations have gradually been overcome by the integration of various types of planning into existing BDI theories, the corresponding interpreters remain fundamentally plan-library based. In this paper, we develop a novel BDI agent architecture driven by generalised planning as means-ends reasoning, in a radical departure from existing architectures. This architecture has two key properties. First, it more closely resembles the foundations of BDI logic and reasoning. Second, it offers substantial gains in efficiency
This paper evaluates the user interface of an in vitro fertility (IVF) outcome prediction tool, focussing on its understandability for patients or potential patients. We analyse four years of anonymous patient feedback, followed by a user survey and interviews to quantify trust and understandability. Results highlight a lay user's need for prediction model explainability beyond the model feature space. We identify user concerns about data shifts and model exclusions that impact trust. The results call attention to the shortcomings of current practices in explainable AI research and design and the need for explainability beyond model feature space and epistemic assumptions, particularly in high-stakes healthcare contexts where users gather extensive information and develop complex mental models. To address these challenges, we propose a dialogue-based interface and explore user expectations for personalised explanations.
Many semantics for weighted argumentation frameworks assume that each argument is associated with an initial weight. However, eliciting these initial weights poses challenges: (1) accurately providing a specific numerical value is often difficult, and (2) individuals frequently confuse initial weights with acceptability degrees in the presence of other arguments. To address these issues, we propose an elicitation pipeline that allows one to specify acceptability degree intervals for each argument. By employing gradual semantics, we can refine these intervals when they are rational, restore rationality when they are not, and ultimately identify possible initial weights for each argument.
In explainable artificial intelligence (XAI) research, the predominant focus has been on interpreting models for experts and practitioners. Model agnostic and local explanation approaches are deemed interpretable and sufficient in many applications. However, in domains like healthcare, where end users are patients without AI or domain expertise, there is an urgent need for model explanations that are more comprehensible and instil trust in the model's operations. We hypothesise that generating model explanations that are narrative, patient-specific and global(holistic of the model) would enable better understandability and enable decision-making. We test this using a decision tree model to generate both local and global explanations for patients identified as having a high risk of coronary heart disease. These explanations are presented to non-expert users. We find a strong individual preference for a specific type of explanation. The majority of participants prefer global explanations, while a smaller group prefers local explanations. A task based evaluation of mental models of these participants provide valuable feedback to enhance narrative global explanations. This, in turn, guides the design of health informatics systems that are both trustworthy and actionable.
We introduce a family of quantitative measures of responsibility in multi-agent planning, building upon the concepts of causal responsibility proposed by Parker et al. [ParkerGL23]. These concepts are formalised within a variant of probabilistic alternating-time temporal logic. Unlike existing approaches, our framework ascribes responsibility to agents for a given outcome by linking probabilities between behaviours and responsibility through three metrics, including an entropy-based measurement of responsibility. This latter measure is the first to capture the causal responsibility properties of outcomes over time, offering an asymptotic measurement that reflects the difficulty of achieving these outcomes. Our approach provides a fresh understanding of responsibility in multi-agent systems, illuminating both the qualitative and quantitative aspects of agents' roles in achieving or preventing outcomes.
Weighted gradual semantics provide an acceptability degree to each argument representing the strength of the argument, computed based on factors including background evidence for the argument, and taking into account interactions between this argument and others. We introduce four important problems linking gradual semantics and acceptability degrees. First, we reexamine the inverse problem, seeking to identify the argument weights of the argumentation framework which lead to a specific final acceptability degree. Second, we ask whether the function mapping between argument weights and acceptability degrees is injective or a homeomorphism onto its image. Third, we ask whether argument weights can be found when preferences, rather than acceptability degrees for arguments are considered. Fourth, we consider the topology of the space of valid acceptability degrees, asking whether gaps exist in this space. While different gradual semantics have been proposed in the literature, in this paper, we identify a large family of weighted gradual semantics, called abstract weighted based gradual semantics. These generalise many of the existing semantics while maintaining desirable properties such as convergence to a unique fixed point. We also show that a sub-family of the weighted gradual semantics, called abstract weighted (Lp,lambda,mu,A)-based gradual semantics and which include well-known semantics, solve all four of the aforementioned problems.
The MIUA 2023 proceedings deal with medical image understanding and analysis, focusing on image interpretation; radiomics, etc.
We present an extension-based approach for computing and verifying preferences in an abstract argumentation system. Although numerous argumentation semantics have been developed previously for identifying acceptable sets of arguments from an argumentation framework, there is a lack of justification behind their acceptability based on implicit argument preferences. Preference-based argumentation frameworks allow one to determine what arguments are justified given a set of preferences. Our research considers the inverse of the standard reasoning problem, i.e., given an abstract argumentation framework and a set of justified arguments, we compute what the possible preferences over arguments are. Furthermore, there is a need to verify (i.e., assess) that the computed preferences would lead to the acceptable sets of arguments. This paper presents a novel approach and algorithm for exhaustively computing and enumerating all possible sets of preferences (restricted to three identified cases) for a conflict-free set of arguments in an abstract argumentation framework. We prove the soundness, completeness and termination of the algorithm. The research establishes that preferences are determined using an extension-based approach after the evaluation phase (acceptability of arguments) rather than stated beforehand. In this work, we focus our research study on grounded, preferred and stable semantics. We show that the complexity of computing sets of preferences is exponential in the number of arguments, and thus, describe an approximate approach and algorithm to compute the preferences. Furthermore, we present novel algorithms for verifying (i.e., assessing) the computed preferences. We provide details of the implementation of the algorithms (source code has been made available), various experiments performed to evaluate the algorithms and the analysis of the results.
Happiness, or subjective wellbeing, brings lasting positive effects to individuals, communities, and societies. Intentional engagement in kind behaviours can have a significant effect on increasing and sustaining subjective wellbeing in humans. In this paper we investigate the effectiveness of a behaviour change intervention for kindness and subjective wellbeing. Using decision tree learning and training data on personality and susceptibility to Cialdini’s persuasive principles, we developed a machine learning model to predict the most effective persuasive principle for an individual. We conducted a randomised controlled experiment to evaluate two interventions (personalised and non-personalised) to motivate kind behaviours. The results indicate that personalised persuasive messages are more effective at stimulating kind behaviours. However, both interventions were effective at improving behavioural intention and subjective wellbeing. These findings have implications for future work on personalisation and design of adaptive behaviour change interventions.
In this paper we examine the effectiveness of several multi-arm bandit algorithms when used as a trust system to select agents to delegate tasks to. In contrast to existing work, we allow for recursive delegation to occur. That is, a task delegated to one agent can be delegated onwards by that agent, with further delegation possible until some agent finally executes the task. We show that modifications to the standard multi-arm bandit algorithms can provide improvements in performance in such recursive delegation settings.
We propose a new generative model that, given past multi-modal Magnetic Resonance Images (MRI) data with different glioma therapies and exam times, can produce realistic MR images that reflect tumor growth forecasts.We developed this model by extending denoising diffusion probabilistic models (DDPMs) with timing and therapy variables as conditional inputs.We then trained the model on real-world postoperative longitudinal MRI data with treatment information from various exam time series.The model has demonstrated promising performance across a range of tasks, including tumor segmentation, growth prediction, uncertainty estimation, and generation of high-quality synthetic multi-modal MR images.Combined with the synthesized MR images, tumor growth predictions with uncertainty estimates can provide useful information for clinical decision-making. DatasetOne-hundred and twenty-seven MRI exams from 23 patients with histologically confirmed high-grade glioma treated at our institution were included in this study [1].Patients received treatment based on standard protocols for high-grade glioma, including surgery, followed by fractionated radiotherapy approximately four weeks after surgery with concomitant and/ or adjuvant chemotherapy (CRT) with temozolomide (TMZ) [2]. Methods and resultsThe proposed conditional DDPM [3] network incorporates a conditional input encoder into each U-Net layer, which involves summing the treatment and day intervals features (between the target day and each reference day, up to three reference exams).Furthermore, the reference MRI exams (e.g., T1/ T1c/Flair at Day 0, 15, etc, up to three sessions) were concatenated into the Gaussian noise, while the corresponding tumor labels were added to the noise.Finally, the model directly generates MR images for the target days while generating tumor masks using DDPM sampling algorithms.The overall concept of our conditional DDPM U-Net model is depicted in Figure 1.
Research has shown that cooperative action struggles to emerge in the noisy variant of the donation game, a simple model of noisy multi-agent systems where indirect reciprocity is required to maximise utility. Such noise can arise when agents may have an incorrect view of the reputation of their interaction partners, or when the actions themselves may fail. Concepts such as generosity, as well as the use of higher-order norms, have been investigated as mechanisms to facilitate cooperation in such environments, but often are not effective or require additional assumptions or infrastructure in the system to operate. In this paper, we demonstrate both analytically and empirically that a simple form of generosity when combined with fine grained reputation can help cooperation emerge. We also show that the use of individual forgiveness strategies rather than the presence of global generosity can support cooperation in such environments.
Simon Miles合作论文数Aerogility27
Wamberto Weber Vasconcelos合作论文数Department of Computing Science,University of Aberdeen18
Stuart Chalmers合作论文数Royal Institute of British Architects
Department of Computing Science
University of Glasgow6
P. M. D. Gray合作论文数University of Aberdeen;Department of Computing Science6