Cooperative multiagent reinforcement learning (MARL) is a promising approach for complex collaborative tasks. However, practical deployment remains challenging due to ambiguous credit assignment, inefficient exploration, and the cold-start problem, particularly in systems where the number of agents grows dynamically. Inspired by human cognitive mechanisms for cognitive task decomposition and experience-based knowledge transfer, we propose multiagent reward decomposition and knowledge transfer (MARDKT), a unified method that jointly addresses the challenges of credit assignment, exploration inefficiency, and cold-start in dynamic multiagent settings. We introduce a four-channel reward decomposition mechanism that independently separates reward signals along local/global and extrinsic/intrinsic dimensions: local rewards drive individual exploration, global rewards foster cooperation, and intrinsic curiosity at both individual and team levels promotes discovery of novel states. To enable rapid integration of new agents, we further design a teacher-student framework where students inherit knowledge from trained teachers via policy imitation and value function distillation. We prove that MARDKT ensures monotonic policy improvement from a local perspective and demonstrate its effectiveness in multi-vehicle on-ramp merging and cooperative box-pushing tasks. Furthermore, MARDKT achieves rapid convergence even as the number of agents increases dynamically, showcasing strong scalability in dynamic environments.
Evolutionary Multitasking has proven effective in addressing multi-task optimization, with knowledge transfer playing a key role in improving algorithm performance. However, existing studies mainly emphasize the timing and methods of transfer, often constrained by specific task assumptions, while overlooking the potential of components during the process. Additionally, reliance on traditional stochastic evolutionary operators limits search efficiency. To address these limitations, this paper proposes a Diffusion-based Multifactorial Evolutionary Algorithm (D-MFEA), featuring a novel component-level knowledge transfer framework for unconstrained single-objective multi-task problems. This framework integrates a diffusion model as the transfer component, enabling efficient knowledge sharing and collaboration between evolutionary and transfer components. It demonstrates strong generalization, seamlessly adapting to and enhancing various MFEA algorithms. By generating high-quality individuals, the diffusion model facilitates positive transfer, reducing reliance on stochastic evolutionary operators and assumptions about task relationships, thereby significantly improving the efficiency of knowledge transfer. Theoretical analyses ensure the diffusion model’s ability to generate high-quality individuals, while experiments on multiple single-objective multi-task benchmarks and a real-world application demonstrate that D-MFEA achieves faster convergence. Ablation studies confirm the effectiveness and robustness of the framework’s components and analyze the impact of varying noise configurations. Extensive results show that our algorithm outperforms state-of-the-art methods.
Multi-task reinforcement learning aims to enable agents to solve a family of tasks by leveraging shared knowledge, including generalization to previously unseen tasks. However, existing approaches remain limited by how knowledge is represented and shared across tasks, which restricts systematic reuse and undermines zero-shot generalization. In particular, skill-based methods learn reusable behavioral primitives, yet these skills are typically encoded as numerical latent representations without explicit semantic grounding, making it difficult to reason about their applicability to novel tasks. Motivated by the need for representations that support both reuse and interpretability, we introduce Semantic Skill Discovery (SemSD), a representation learning framework that grounds behavioral skills in human-interpretable task semantics to support transferable and compositional decision-making. SemSD aligns learned behavioral primitives with task descriptions in a shared embedding space and organizes skills according to semantic relationships, enabling principled skill selection and composition when encountering unseen tasks. Extensive experiments on the Meta-World benchmark demonstrate that SemSD achieves substantially improved zero-shot generalization performance across diverse tasks, with ablation studies further validating the importance of semantic grounding in multi-task skill learning.
Real-world deployment of Multi-Agent Reinforcement Learning (MARL) in Internet of Things (IoT) systems requires convergence, verifiability, and privacy to be jointly guaranteed, a capability absent in current approaches. While Multi-Agent Trust Region Learning (MATRL) ensures Nash equilibrium convergence, it risks data exposure and malicious attacks. To address this, we propose Trustworthy Distributed Mirror Learning (TDML), the first method to unify convergence, verifiability, and privacy in MARL. TDML theoretically breaks MATRL's centralized architecture into agent-local learning and inter-agent communication. This allows key data to be secured with advanced techniques without compromising the theoretical properties of trust-region learning. Specifically, TDML introduces three core innovations: (1) an information functional that unifies all communication behaviors in distributed MATRL and enables flexible integration of security mechanisms; (2) split advantage computation, which decouples raw inputs from global advantages via intermediate representations to protect local data privacy; and (3) a security scheme that ensures verifiable message exchange by attaching zero-knowledge proofs to inter-agent communications. We prove TDML converges to a Nash equilibrium while providing verifiability and privacy guarantees. More importantly, it constructs a mirror space for trustworthy MARL, where derivative algorithms inherit these theoretical guarantees. Experiments show TDML outperforms state-of-the-art methods, improving attack resilience by up to 76% (sign-flipping attack) and achieving a 90+% privacy reconstruction error, while reducing communication overhead by up to 99% compared to homomorphic encryption. TDML establishes a foundational framework for trustworthy MARL, from which derivative algorithms inherit core guarantees for secure, real-world IoT deployment.
Multi-agent reinforcement learning (MARL) faces a critical dual challenge: balancing exploration and exploitation under non-stationarity, and effectively sharing knowledge without propagating errors. Existing methods often rely on fixed exploration schedules that fail to adapt to dynamic interactions, while unselective knowledge sharing risks spreading suboptimal policies. To address these issues, we propose EntShare, a value-based CTDE framework that utilizes entropy as a unified signal to (i) adaptively regulate exploration intensity and (ii) gate knowledge sharing. EntShare introduces a Dynamic Multi-Scale Entropy (DMSE) controller to measure policy uncertainty across time horizons and an entropy-gated mechanism that disseminates only high-credibility experiences. Theoretical analysis confirms entropy's closed-loop role in the exploration-exploitation trade-off. Across five partially observable environments, EntShare surpasses 10 representative baselines across four categories, demonstrating strong generalization across diverse coordination settings. Furthermore, it maintains superior post-convergence stability, effectively avoiding the catastrophic performance drops observed in comparative methods.
Just-In-Time Defect Prediction (JIT-DP) has become an essential technique for improving software quality by enabling carly detection of risky commits. However, the practical effectiveness of current JIT-DP remains lim-ited due to three key challenges: (1) label inconsistency in real-world multilingual datasets, especially within duplicate code change groups introduced by operations such as "cherry-pick", (2) insufficient discriminative power of existing models in separating defective and non-defective commits; and (3) suboptimal utilization of heterogeneous information derived from code changes and commit messages. To address these challenges, this paper introduces label unification algorithm and the JIT-CLASP, a novel framework that integrates stochastic batch contrastive learning, and semantic-aligned heterogeneous fusion. First, we design Unilabel, a label unification algorithm that resolves inconsistency among duplicate code change groups and produce a higher-quality dataset, LApredictamified Second, we apply Stochastic Batch Contrastive Learning (SBCL.) to multilingual JIT-DP, substantially enhancing the model's ability to capture subtle distinctions between defective and non-defective commits. Third, we propose the Semantic Aligament and Feature Integration (SAFI) module, which effectively aligns and integrates heterogeneous representations from code changes and commit messages Comprehensive experiments on multilingual benchmarks demonstrate that UniLabel consistently improves model performance, particularly in Java projects, The SBCL significantly boosts the discriminative capability of state-of-the-art baselines; and the JIT-CLASP achieves substantial performance gains over all existing baselines, with notable improvements in F1, MCC, AUC-ROC, and AUC-PR. Ablation studies further confirm that cach com-ponent of JIT-CLASP contributes synergistically to its performance. All datasets, code, and replication materials are publicly available.
The continuous expansion of open data platforms and research repositories has led to a fragmented dataset ecosystem, posing significant challenges for cross-source data discovery and interpretation. To address these challenges, we introduce SeDa–a unified framework for dataset discovery, semantic annotation, and multi-entity augmented navigation. SeDa integrates more than 7.6 million datasets from over 200 platforms, spanning governmental, academic, and industrial domains. The framework first performs semantic extraction and standardization to harmonize heterogeneous metadata representations. On this basis, a topic-tagging mechanism constructs an extensible tag graph that supports thematic retrieval and cross-domain association, while a provenance assurance module embedded within the annotation process continuously validates dataset sources and monitors link availability to ensure reliability and traceability. Furthermore, SeDa employs a multi-entity augmented navigation strategy that organizes datasets within a knowledge space of sites, institutions, and enterprises, enabling contextual and provenance-aware exploration beyond traditional search paradigms. Comparative experiments with popular dataset search platforms, such as ChatPD and Google Dataset Search, demonstrate that SeDa achieves superior coverage, timeliness, and traceability. Taken together, SeDa establishes a foundation for trustworthy, semantically enriched, and globally scalable dataset exploration.
Safe reinforcement learning has been widely applied in various domains, including gaming and robotic control. However, traditional safe reinforcement learning methods often suffer from inefficiency and instability due to the inherent limitations of standard optimization and projection-based strategies. To address these challenges, this paper introduces the One-step Anticipatory Policy Selector (OSAPS), a novel algorithm that integrates Policy Selection, One-step Planning, and Adaptive Safety Threshold mechanisms to significantly enhance the efficiency and safety of agent behavior. OSAPS features a conditional dual-path policy network architecture known as the Policy Selection mechanism, comprising a Vertex Policy Network (VPN) and a Vanilla Policy Network (PN). The VPN prioritizes safety through stepwise optimization, thereby mitigating the PN’s potential emphasis on rapid execution at the expense of safety. This design allows for dynamic switching between the two sub-networks, optimizing policy selection based on situational requirements. At the core of OSAPS lies its One-step Planning mechanism, which enables agents to make anticipatory adjustments to actions within safety constraints, mimicking advanced human-like decision-making processes. Furthermore, we introduce an Adaptive Safety Threshold mechanism, which dynamically adjusts safety boundaries to balance exploration and exploitation while ensuring enhanced agent safety. Our empirical evaluations in diverse simulation environments demonstrate significant improvements: OSAPS achieves an 18.52% increase in time efficiency in Pendulum and a 7.67% improvement in Hovercraft, all while maintaining stability and safety standards. The incorporation of the Adaptive Safety Threshold not only amplifies cumulative rewards but also achieves a twofold to fivefold increase in safety measures on average. Consequently, OSAPS offers a more balanced approach, accelerating convergence rates, improving time efficiency, and maintaining overall stability, thus equipping agents with superior capabilities to tackle complex real-world challenges.
Just-in-Time Defect Prediction and Localization (JIT-DP and DL) play a crucial role in software quality assurance by identifying defective code changes and locating faulty lines at the time of code submission. While existing methods leverage either handcrafted expert features or semantic features extracted by deep learning models, few explicitly distinguish or effectively fuse these two types of information. In this paper, we propose JIT-Coka, an improved framework for JIT-DP and DL tasks that combines an encoder-decoder based pre-trained model of code (CodeT5) with KANLinear, which is a spline-based adaptive nonlinear classifier used to model fine-grained nonlinear relationships between semantic and expert features. We conduct comprehensive experiments on the high-quality JIT-Defects4J dataset, evaluating JIT-Coka and representative baselines using multiple metrics. Results show that JIT-Coka significantly outperforms state-of-the-art models in defect prediction, improving F1 and MCC by 6.8
Silent vulnerability fixes (SVFs) are pervasive in open-source ecosystems and pose significant risks to software supply chains due to incomplete or delayed disclosure. Existing SVF identification methods often rely on coarse-grained diff representations, loosely aligned change fragments, or limited modeling of developer intent, restricting their robustness and practical applicability across projects. This paper presents HuBCAP, a hunk-based and context-aware predictor that aligns code-change representations with Git's standardized hunk structure while leveraging pre-trained models (PTMs) for semantic reasoning. HuBCAP explicitly models semantic differences between paired pre- and post-change code fragments, performs hierarchical aggregation at both hunk and file levels, and integrates commit-message semantics within a unified dual-branch architecture. The framework avoids language-specific parsing rules and handcrafted syntactic features, following a language-agnostic design principle based on Git-standard diff structures.We evaluate HuBCAP on a large-scale, manually curated dataset covering Java and Python projects and compare it with state-of-the-art task-specific models and general-purpose LLMs. Results demonstrate consistent performance improvements. Module-level ablation experiments confirm the necessity of hunk-level aggregation, file-level aggregation, and explicit difference modeling, while input-level analyses reveal the complementary roles of code changes and commit messages. Finally, real-world case studies in software supply chain ecosystems show that HuBCAP can uncover undisclosed and cross-project propagated vulnerability fixes. We further analyze false positive patterns and discuss deployment-oriented optimization strategies such as confidence-aware triaging and threshold tuning, highlighting HuBCAP's practical value for improving supply chain risk visibility.
This paper proposes an innovative Multi-Agent Continual Reinforcement Learning (MACRL) framework to address the challenges of continual learning in dynamic multi-agent systems. Traditional reinforcement learning suffers from catastrophic forgetting and inefficient cross-task knowledge transfer in non-stationary environments. To overcome these limitations, we introduce two key components: (1) a Multi-Timescale Replay (MTR) buffer, which hierarchically stores experiences across varying timescales to balance new task learning and prior knowledge retention, and (2) a dynamic task classification mechanism that employs an attention-based contextual encoder to measure task similarity and adaptively route policies, thereby minimizing inter-task interference. Experiments on cooperative multi-agent benchmarks (LBF and PP) demonstrate that our framework achieves up to higher average return compared to baselines in sequential task learning, with superior zero-shot generalization performance. Ablation studies further validate the critical roles of MTR and task classification in mitigating catastrophic forgetting. This work provides a scalable solution for collaborative decision-making in complex, evolving environments.
Expensive constrained multiobjective optimization problems (ECMOPs) are prevalent in real-world scientific research and industrial applications. However, the complexity of feasible regions and the limitation on the number of available function evaluations often prevent most algorithms from achieving satisfactory results. To address these challenges, this article proposes an ensemble-based surrogate framework. Specifically, a global model and multiple local models are constructed as ensemble members to approximate each constraint function, aiming to improve the accuracy of landscape approximation for ECMOPs with complex feasible regions. Additionally, a novel vector-based constrained dominance principle is suggested to maintain the balance between objectives and constraints. By leveraging reference vectors, potential scenarios of the population during the evolutionary process are identified, and the customized selection strategy is devised for each scenario. These two techniques are integrated into a two-stage optimization framework, resulting in a surrogate-assisted evolutionary algorithm for solving ECMOPs. Through extensive experimental investigations, the proposed algorithm demonstrates significant superiority over seven other state-of-the-art peer algorithms on both benchmark test problems and real-world applications.
Federated learning (FL) watermarking has emerged as a key mechanism for protecting the intellectual property (IP) of collaboratively trained models. However, a critical yet underexplored vulnerability arises from the legal transparency requirement: watermarks submitted as evidence in IP disputes must be disclosed to arbitration bodies, potentially exposing them to adversaries and creating an evidence leakage threat. We formalize this threat and propose a federated unlearning attack that uses leaked triggers as forgetting instructions. Inspired by the psychological mechanism of active forgetting, our method constructs a non-knowledge dataset from leaked watermarks and incrementally trains the model to erase the embedded watermark while preserving main-task performance via a Fisher-regularized knowledge retention module. Experiments show that the watermark success rate drops to 0.1 across various FL watermarking frameworks, while main-task accuracy is maintained within a small margin. We further evaluate the attack against adaptive defenses and specialized watermark removal baselines, revealing fundamental limitations in FL watermarking designs. Our work provides a systematic analysis of FL watermarking systems and motivates a rethinking of watermark disclosure protocols.
Differential Privacy Federated Learning (DP-FL) combines Differential Privacy (DP) with Federated Learning (FL), enabling multiple clients to collaboratively train a shared model while protecting data privacy. However, introducing DP into FL will add noise to model parameters, which typically deteriorates model convergence. Although a recent work has revealed the compensation effect by increasing the total batch size, it overlooks the "generalization gap" phenomenon, which is induced by excessively large batch size and has been long discussed in the machine learning field. In order to avoid the other extreme, we strengthen several core components in, and propose an Incentive-driven Differential Privacy Federated Learning (IDP-FL) framework. First, instead of building all hopes on batch sizes, the proposed framework jointly considers the non-IID degrees of local data and clients' privacy budgets, minimizing the difference between the optimal batch size for each selected client and its corresponding critical batch size. Second, we reconfigure the batch size for each selected client by balancing the negative impact of DP noise on convergence and of the "generalization gap" phenomenon. Finally, we design a Stackelberg game-based incentive mechanism that encourages clients to contribute computational resources, and prove the existence of a Stackelberg equilibrium to guarantee stability. Through numerical evaluations on real-world datasets, we show that our IDP-FL framework outperforms existing algorithms in terms of test accuracy and utility. Ablation studies further confirm the effectiveness of each component.
Despite extensive safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. However, existing methods generally lack the capability for continuous learning and self-evolution from interactions, limiting the diversity and adaptability of attack strategies. To address this, we propose ASTRA, an automated framework capable of autonomously discovering, retrieving, and evolving attack strategies. ASTRA operates on a closed-loop “attack-evaluate-distill-reuse” mechanism, which not only generates attack prompts but also automatically distills reusable strategies from every interaction. To systematically manage these strategies, we introduce a dynamic three-tier strategy library (Effective, Promising, and Ineffective) that categorizes strategies based on performance. This hierarchical memory mechanism enables the framework to enhance efficiency by leveraging successful patterns while optimizing the exploration space by avoiding known failures. Extensive experiments in a black-box setting demonstrate that ASTRA significantly outperforms existing baselines.
Cutting-edge bioinformatics research is increasingly intertwined with pre-trained model techniques. However, achieving superior performance of these models in downstream applications typically requires large amounts of accurately labeled experimental data for fine-tuning, which poses substantial practical challenges due to the difficulty in preparing such datasets at scale. To address this limitation, we propose a novel few-shot fine-tuning framework, the Evolution-Aware Adaptation (Evo-AA). It aligns fine-tuning with pre-training objectives while integrating prompt learning and biological coevolutionary insights. Additionally, we introduce reinforced prompting and lambda ranking loss to further improve performance. Extensive experiments demonstrate that Evo-AA with limited training set, enhances the spearman correlation in fine-tuning tasks, while achieving superior precision and recall rates in homology search tasks. Our findings suggest that Evo-AA holds great potential to drive advancements in protein engineering and computational biology.
Deep learning models have driven significant progress in predicting protein function and interactions at the protein level. While these advancements have been invaluable for many biological applications such as enzyme engineering and function annotation, a more detailed perspective is essential for understanding protein functional mechanisms and evaluating the biological knowledge captured by models. This study introduces VenusX, the first benchmark designed to assess protein representation learning with a focus on fine-grained intra-protein functional understanding. VenusX comprises three major task categories across six types of annotations, including residue-level binary classification, fragment-level multi-class classification, and pairwise functional similarity scoring for identifying critical active sites, binding sites, conserved sites, motifs, domains, and epitopes. The benchmark features over 878,000 samples curated from major open-source databases such as InterPro, BioLiP, and SAbDab. By providing mixed-family and cross-family splits at three sequence identity thresholds, our benchmark enables a comprehensive assessment of model performance on both in-distribution and out-of-distribution scenarios. For baseline evaluation, we assess a diverse set of popular and open-source models, including pre-trained protein language models, sequence-structure hybrids, structure-based methods, and alignment-based techniques. Their performance is reported across all benchmark datasets and evaluation settings using multiple metrics, offering a thorough comparison and a strong foundation for future research. Our code (https://anonymous.4open.science/r/VenusX-4674), data (https://huggingface.co/collections/anonymous-researcher-123/venusx-68cc5163ade527b0974bab29), and a leaderboard (https://anonymous-researcher-816.github.io/) are provided as open-source resources.
Mobile-based cloud computing (MBCC) has become a paradigm shift that greatly en-hances the computing power of the mobile devices with limited resources by enabling them to access in-the-cloud resources of highly scalable infrastructures at considerable distance. Nevertheless, the constant and unregulated data transfer between mobile customers and remote cloud providers is deep bandwidth expenses, and at the same time, interfere with the overall system performance, address undesirable energy utili-zation, and sensitive information to security weaknesses. These are especially the problem in bandwidth-limited geographical locations or cost-aware enterprise settings in which the network usage is directly translated into expensive operational costs. This paper proposes a hierarchical, edge-aware architecture of the system in an approach to provisioning a holistic optimization of bandwidth usage as long as end-users remain robust in performance, in no less than three-way integration of mobile devices, local-ized, edge, and centralized cloud data centers. In this proposed ecosystem tasks to be offloaded are strictly encrypted and coded before their offloading and preserving sen-sitive user information as well as mathematically minimizing the encoded payload to the minimum. Also, we introduce a new hybrid approach that combines the data com-pression algorithms and game-theory-based optimization of tasks offloading, dynamic synchronization processes, and protocols of secure transmission through the encrypt-ed and encoded task packing. Finally, this study can be used in regards to the basic development of more sustainable, secure, and cost-efficient mobile-cloud systems that are highly essential in the next generation of Internet of Things (IoT) applications in cost-sensitive environments.
To balance the gap between data privacy and the need for data fusion, federated learning (FL) has been proposed and has become a hot-point method to address data silos and privacy issues. However, AI models exchanged in FL face risks such as illegal copying, redistribution and/or free-riding. To address these risks, FL watermarking frameworks have been proposed to assert and protect the intellectual property (IP) of models, which are resistant to popular watermark removal attacks. Knowledge distillation has recently been of significant contribution to FL convergence performance optimization but brings vulnerability to FL watermark robustness with distillation attack, which enables attackers to maintain high performance on the main task while erasing the watermarks. In response, we introduce a new FL watermarking framework called FedRW, which focuses specifically on anti-distillation. FedRW employs model regularization techniques to bind the main task parameters with the watermark task parameters, thereby enhancing resistance to distillation attacks. Extensive experiments confirm the threat of distillation attacks in FL and demonstrate that FedRW is more resistant to distillation compared to existing FL watermarking frameworks.