The rapid globalization of supply chains (SC) has opened vast prospects but also increased disruption risks and substantial uncertainties in supply chain systems-of-systems (SCSoSs). Supply chain reconfiguration (SCR) has emerged as a pivotal strategy for mitigating these risks. This paper proposes a multi-agent reinforcement learning-based resilience reconfiguration approach for SCSoSs to address the agile, stable, and spatio-temporal requirements of SCR under disruption risks. It begins by detailing the SCR issue involving suppliers, manufacturers, distributors, and consumers amid disruption risks and introduces three resilience strategies: filling, repairing, and recruiting. A three-phase model for calculating resilience and reconfiguration costs is then developed, grounded in the supply chain directed network (SCDN). Following this, the reconfiguration process is modeled as a partially observable Markov decision process (POMDP), with the state space representing SC elements and the action space including available strategies. The reward function balances resilience and costs considerations. Utilizing the multi-agent proximal policy optimization (MAPPO) technique, the method enables dynamic reconfiguration of SCSoSs, demonstrating its effectiveness through experimental simulations. The analysis also explores how different attributes affect reconfiguration outcomes. Results indicate that the MAPPO approach substantially enhances reconfiguration performance under disruption risks compared to other baselines, providing valuable insights for modern SC management.
We can use multi-agent reinforcement learning algorithms such as Multi-Agent Deep Deterministic Policy Gradient (MADDPG) to effectively solve the organization problem of multi-agent systems. However, the training process of these algorithms is challenging due to sparse rewards, which means it typically takes a long time to finish the training. To address the challenge, we propose the Large Language Model (LLM) Assisted MADDPG (LLM-MADDPG) algorithm that utilizes LLM to guide the training of multi-agent systems. The core original contribution is two-fold. First is the dual-track mechanism, which combines the autonomous policy learning capability of the distributed Actor network with the cross-domain, high-level priori knowledge guidance provided by the Large Language Model (LLM). They both work together to improve training efficiency. The second is knowledge quality assurance through human-in-the-loop screening. The expert strategy knowledge generated by LLM is refined through manual screening and transformed into a code-based reward function to ensure logical rationality and accelerate the generation of correct guidance from LLM.
Multi-agent systems must operate in increasingly dynamic environments, where sudden changes such as the appearance of moving or unforeseen obstacles pose major challenges to the system’s adaptability. In this paper, we propose a bi-level Function-Behavior-Structure (FBS) self-closed-loop framework to enable both individual agents and clusters of agents to reason adaptively from task requirements and to change their behavior in response to environmental changes. Inspired by the FBS model in product design, this framework introduces self-reasoning at two hierarchical levels: the cluster-level and the individual agent-level. To enhance real-time adaptability, we develop an environment prediction method based on incremental changes in the location of obstacle pixels observed across time steps. This allows for the construction of dynamic prediction maps. Key to this is the introduction of a system event influence value, to quantify environmental disruptions and determine whether the system should perform global re-planning, local re-planning, or continue with the current plan, This helps the researcher establish a balance between responsiveness and computational efficiency. The framework is demonstrated with a collaborative box-pushing scenario involving multiple agents navigating evolving environments with successive levels of complexity. Comparative experiments show that the proposed method outperforms a traditional bi-level planning approach in both success rate and computational resource usage, particularly in high-complexity scenarios. The framework offers a generalizable approach to adaptive design for multi-agent systems and can be applied to domains such as smart manufacturing, logistics, and autonomous operations.
In dynamic environments, such as box-pushing tasks, multi-agent systems (MAS) face significant challenges in coordinating agents within high-density settings while managing uncertainties arising from fluctuations in agent configurations and environmental dynamics. In this study, we explore the integration of surrogate response surface modeling (SRSM) with optimization algorithms, comparing Stochastic Gradient Descent (SGD) with a fixed learning rate, and adaptive learning rate (ALR) optimizers—including Adaptive Moment Estimation (ADAM) and Adaptive Approximate Direction Method Algorithm (AADMA)—to enhance MAS performance metrics, such as Agent Collision Rate (ACR), Agent Movement Frequency (AMF), and Task Completion Time (TCT). Through systematic experimentation across five scenarios, SRSM is employed to uncover key trends in MAS performance and identify configurations that improve scalability and adaptability. From the analysis of simulation data, it has been observed that SGD struggles significantly in dynamic environments, while ADAM demonstrates moderate improvements. However, AADMA consistently outperforms both by reducing loss, lowering collision rates, increasing movement efficiency, and achieving shorter task completion times. Performance comparison charts and loss function graphs emphasize AADMA’s superiority in addressing the complexities of real-time coordination and adaptability. Through this study, we highlight the critical role of combining SRSM with ARL to design an MAS that is capable of thriving in complex, dynamic, and high-density environments. By addressing key scalability and adaptability challenges, the proposed framework significantly advances MAS design, paving the way for improved multi-agent coordination in real-world applications.
Multi-agent self-organizing systems (MASOS) exhibit key characteristics including scalability, adaptability, flexibility, and robustness, which have contributed to their extensive application across various fields. However, the self-organizing nature of MASOS also introduces elements of unpredictability in their emergent behaviors. This paper focuses on the emergence of dependency hierarchies during task execution, aiming to understand how such hierarchies arise from agents' collective pursuit of the joint objective, how they evolve dynamically, and what factors govern their development. To investigate this phenomenon, multi-agent reinforcement learning (MARL) is employed to train MASOS for a collaborative box-pushing task. By calculating the gradients of each agent's actions in relation to the states of other agents, the inter-agent dependencies are quantified, and the emergence of hierarchies is analyzed through the aggregation of these dependencies. Our results demonstrate that hierarchies emerge dynamically as agents work towards a joint objective, with these hierarchies evolving in response to changing task requirements. Notably, these dependency hierarchies emerge organically in response to the shared objective, rather than being a consequence of pre-configured rules or parameters that can be fine-tuned to achieve specific results. Furthermore, the emergence of hierarchies is influenced by the task environment and network initialization conditions. Additionally, hierarchies in MASOS emerge from the dynamic interplay between agents' "Talent" and "Effort" within the "Environment." "Talent" determines an agent's initial influence on collective decision-making, while continuous "Effort" within the "Environment" enables agents to shift their roles and positions within the system.
The question we address in this chapter follows: How can students leverage Learning Statements collected from former students in a design, build, and test course to enhance their own learning by reflecting on doing? A Learning Statement is a structured [Experience|Learning|Value] text-based construct for students in AME4163 Principles of Engineering Design to record what they learned by reflecting on authentic immersive experiences throughout the semester. The immersive experiences include lectures, assignments, reviews, building, testing, and a post-analysis of an electro-mechanical device to address a given customer need. Over the past four years, at the University of Oklahoma-Norman, we have collected almost 32,000 Learning Statements from about 600 students. Based on these Learning Statements, in our earlier work, we proposed a text-mining framework to improve our understanding of what students have learned by reflecting on doing and thence how we might improve the delivery of the course in the future. Our focus, in the earlier work, was on leveraging the historical data to improve the course from the instructor’s perspective. In this chapter, we look at the data from the student’s perspective and focus on demonstrating how students can utilize the former students’ Learning Statements as a knowledge base to tutor themselves in their own learning through reflection on doing. In this chapter, we describe a self-tutoring system that allows students to compare their own Learning Statements to that of the former students in three ways, namely, Experience, Learning, and Value. The core of the self-tutoring system is a semantics-based similarity metric that enables the automatic searching of historical Learning Statements with similar experience, learning, or value at the semantic level. The value of the self-tutoring system is that it provides an opportunity for students to augment what they have written by reflecting on prompts from the self-tutoring system to augment what they eventually share as experience, learning, and value in their own Learning Statements.
In this study, we explore the fundamentals of design and deployment of Multi-Agent Swarm Engineering Systems (MASESs), a comprehensive framework aimed at improving decision-making and operational efficiency in industrial environments. Swarm engineering systems draw inspiration from collective behaviors observed in natural swarms (such as ants or bees), offering a decentralized approach to problem-solving in complex systems. We provide a detailed theoretical foundation for MASES, focusing on key elements such as spatial control formation, adaptive coordination, decision-making dynamics, and critical factors such as scalability, adaptability, and robustness in practical applications. The analysis is enriched with insightful tables that highlight current innovations and challenges, underscoring the importance of balancing decentralization with scalability. In transitioning to practical applications, we discuss design methodologies and performance evaluation strategies, emphasizing the transformative potential of swarm engineering systems. Advanced frameworks and simulation environments have been shown to play a crucial role in ensuring that MASES operates effectively, fostering intelligence and adaptability to meet complex industrial demands. By showcasing various applications, we highlight the potential of MASES to enhance operational landscapes across industries and identify pathways for future research to refine and perfect this integration.
Advanced multi-agent systems are capable of executing tasks in complex and unknown domains that are unsuitable for humans. However, the design of multi-agent systems faces challenges in balancing global and local decision-making and enabling the system to adaptively generate desired behaviors in a dynamic environment. To address these challenges, in this article, we propose a multi-agent system design framework based on bilevel closed-loop planning. We use the multi-agent box-pushing problem as an example to verify the framework. Within this framework, the upper-level planning (which is used for box position prediction) and the lower-level planning (which is used for agent position allocation) are designed to connect and coordinate between the global and local decisions. The influence of states based on planning creates a closed-loop control mechanism with temporary targets as input, allowing the system to adapt to various environments. In this article, we use webots as the simulation platform to conduct multi-agent box-pushing experiments and compare the results with the rule-based method to demonstrate the effectiveness and advantages of our approach.
Multi-agent Self-organizing Systems (MASOSs) feature scalability, flexibility, and robustness, and their application in complex assembly tasks has garnered increasing attention. However, a pressing challenge persists in effectively curbing potential negative outcomes in MASOSs while guiding them towards positive ones. To address this challenge, in this paper the concept of Shared Mental Models (SMMs) is applied to MASOSs, and an SMMs-based collaboration method in assembly tasks for MASOSs is proposed. Integrating individual mental models from multiple agents, an SMMs structure for MASOSs is initially constructed. Building upon this structure, a collaboration method for MASOSs to execute assembly tasks is afterward proposed, and the impact of SMMs on task performance and emergent behaviors of MASOSs is explored through an “L-shape” assembly case. The results demonstrate that the SMMs-based method significantly improves task reliability, task efficiency, time efficiency, and energy efficiency. Additionally, we find that increasing agent team size initially leads to positive emergent behaviors due to scale advantages, but surpassing the optimal size results in negative emergent behaviors due to coordination disadvantages. Furthermore, the degree of knowledge sharing among agents also has a significant impact on task performance and emergent behaviors. SMMs are expected to serve as a mechanism to regulate and optimize team emergent behaviors, ultimately achieving optimal system performance.
In this article, we address the following question: How can the designers of multi-agent systems use the relationships among relevant parameters to quantitatively characterize how a multi-agent system produces processes involving collective behavior to better control and design multi-agent systems? To address this challenge, we propose a model that couples environmental complexity with rule adoption rates to characterize quantitatively the interaction between the environmental changes and the adaptation of behavioral rules. This model is applied in a dynamic environment with a box-pushing task, simulated using the webots platform. In the model, we define static complexity (including obstacle distribution and target positioning) and dynamic complexity (based on the box's position and orientation) while incorporating rule adoption rates to capture the agents' behavioral adjustments. Experimental results indicate that multi-agent system behavior evolves with four distinct patterns: (1) initialing pattern, (2) adjusting pattern, (3) stabilizing pattern, and (4) ending pattern. These findings provide valuable insights into the dynamics of collective behavior in multi-agent systems operating in complex, dynamic environments and propose a real-time control framework that can be applied to practical scenarios such as logistics sorting, disaster relief, and other domains requiring robust multi-agent system performance.
This paper presents a decision tree-based resource recommendation system for aerospace manufacturing. During the resource selection stage of the aerospace manufacturing process design, the recommendation system first acquires the parts and their current features. The decision tree model is then used to determine the type of equipment required for these features, forming an initial set of equipment. Subsequently, the process rule base screens this set, matching the equipment that meets the requirements of the parts and the current process, forming a set of equipment to be recommended. Finally, the suitable equipment is recommended to the process personnel. By constructing and training the decision tree model, the system achieves rapid and accurate manufacturing resource recommendations, thereby enhancing the efficiency of process design for aerospace parts.
When natural or man-made disasters occur at sea, a maritime unmanned rescue system-of-systems (MURSoSs), as an important guarantee for the safety of people's lives and property, has a rapid response to emergency rescue services. In the design process of MURSoSs, it is often faced with problems such as unexplainable mechanisms and inaccurate modeling due to the characteristics of multi-level coupling and irregular emergence. A design method of MURSoSs based on an attention Transformer is proposed. First, this paper decomposes the MURSoSs into a multi-level organizational structure composed of individual performance, group structure, and overall effectiveness. Second, the maritime unmanned rescue simulation environment is constructed, to obtain the multi-level evolution data of group structure and overall effectiveness under different individual performance, and the attention Transformer is used to mine the functional correlation relationship between levels. Finally, the experimental simulation prediction is carried out by using this function relationship. Taking the design of MURSoSs as an example, the results show that the constructed model accurately quantifies the correlation relationship between levels and effectively reveal the internal logic of the emergence process of maritime rescue.
Self-organizing systems are capable of executing tasks in complex and unknown domains that are unsuitable for humans. However, designing of self-organizing systems is challenging due to balancing global and local decision-making and in enabling the system to adaptively generate desired behaviors in different environments. To address these challenges, in this paper we propose a self-organizing system design framework based on stigmergy and bi-level planning. We use the multi-agent box-pushing problem as an example to verify the framework. Within this framework, the upper-level planning (which is used for box position prediction) and the lower-level planning (which is used for agent position allocation) are designed to connect and coordinate between the global and local decisions. The influence of states based on stigmergy creates a closed-loop control mechanism with "pheromones" as input, allowing the system to adapt to various environments. In this paper we use Webots as the simulation platform to conduct multi-agent box-pushing experiments and compare the results with an ant colony handling method, to demonstrate the effectiveness and advantage of our approach.
A process route recommendation engine for aviation structural components based on ANN-LCS algorithm is introduced. By processing the machining process knowledge of aviation structural components, defining their basic information and process route data, and then constructing and training an ANN (Artificial Neural Network) to classify parts. Based on part classification, the LCS (Longest Common Substring) algorithm based on dynamic programming is used to generate process routes. Through the testing of process route recommendation examples, the results show that this method can effectively obtain and select feasible process routes in the process design stage.
In this paper, we address the following question: How can multi-robot self-organizing systems be designed so that they show the desired behavior and are able to perform tasks specified by the designers? Multi-robot self-organizing systems, e.g., swarm robots, have great potential for adapting when performing complex tasks in a changing environment. However, such systems are difficult to design due to the stochasticity of system performance and the non-linearity between the local actions/interaction and the desired global behavior. In order to address this, in this paper, we propose a framework for designing self-organizing systems using Multi-Agent Reinforcement Learning (MARL) and the compromise Decision-Support Problem (cDSP) construct. The proposed framework consists of two stages, namely, preliminary design followed by design improvement. In the preliminary design stage, MARL is used to help designers train the robots so that they show stable group behavior for performing the task. In the design improvement stage, the cDSP construct is used to explore the design space and identify satisfactory solutions considering several performance indicators. Surrogate models are used to map the relationship between local parameters and global performance indicators utilizing the data generated in the preliminary design. These surrogate models represent the goals of the cDSP. Our focus in this paper is to describe the framework. A multi-robot box-pushing problem is used as an example to test the framework's efficacy. This framework is general and can be extended to design other multi-robot self-organizing systems.
Addressing the business logic and relevant requirements of process personnel during the selection stage of processing methods for aerospace structural parts, this paper introduces a recommendation system for aircraft parts processing methods based on a hybrid of rules and instances. Firstly, the typical features and processing methods of aircraft parts are defined, vectorized, and their feature similarity is calculated using the Euclidean distance method. Based on industry standards, a rule library for selecting processing methods is established to filter recommendation results. For the characteristics of the parts, the recommendation system retrieves the most similar features and their processing methods from the feature instance library for recommendation. Validation tests demonstrate that this recommendation system can effectively recommend suitable processing methods based on the features of aircraft parts.
Surrogate models have been widely used in engineering design for approximating a simulation system with high computational cost. Complex system design typically is a multi-stage and multi-discipline design problem, which requires a large number of surrogate models. The choice of surrogate modeling method (SMM) is critical as it directly impacts the performance of both the surrogate models and the designed systems. With the growing variety of SMMs, designers face challenges in selecting the appropriate methods for their specific applications. To address this, we propose a representation and recommendation framework for surrogate modeling methods based on knowledge graph. Firstly, we develop an ontology to formally represent core concepts involved in the recommendation for surrogate modeling methods, including surrogate modeling method, surrogate model, and data sets,etc. Secondly, we extract 460 samples from 46 benchmark functions using Latin hypercube sampling to construct a knowledge graph with 8,343 nodes and 16,100 relationships, which involves 7,820 surrogate models generated from 17 surrogate modeling methods. Finally, we propose a knowledge graph-based recommendation method for surrogate modeling method named KGRSMM to facilitate the selection of an appropriate surrogate modeling method. We test the efficacy of KGRSMM using examples of theoretical problems and engineering problems of hot rod rolling respectively. It is shown in the results that KGRSMM is capable of recommending surrogates with appropriate accuracy, robustness, and time to satisfy designers’ preferences.
Spatio-temporal crowdsourcing (STC) is a typical case of complex system-of-systems (SoSs) design, wherein the primary objective is to allocate real-time tasks to suitable groups of workers. Over time, the STC allocation has gradually evolved into a dynamic matching involving three distinct entities: tasks, workers, and workplaces. Aiming at addressing the problems of poor convergence, slow response and sparse actions caused by the spatial complexity and time dynamics of the STC, this paper proposes an improved proximal policy optimization algorithm based on an invalid action masking (IAM-IPPO) for the SoSs design of the STC. Initially, the ternary dynamic matching (TDM) of tasks, workers and workplaces in the STC is described. Furthermore, the STC allocation is formulated as a Markov decision process, with the corresponding definition of state space, action space, and reward mechanism. On this basis, an invalid action masking (IAM) method is mainly introduced to update the policy-based network of proximal policy optimization (PPO), realizing sampling only from valid actions to masking invalid action selection. Subsequently, the algorithmic framework of IAM-IPPO is elaborated upon, and the model is trained to generate an effective allocation scheme. Comparative experiments are conducted on authentic datasets, aiming to assess performance indicators of the presented approach. The findings demonstrate a substantial enhancement in performance for the IAM-IPPO algorithm compared to other baselines, which is helpful in exploring excellent design schemes of the crowdsourcing SoSs, especially in dynamic large-scale cases.
In this paper, we address the following question: How can self-organizing system designers steer the system towards expected global behaviors that satisfy multiple conflicting performance indicators? Self-organizing systems (SOS) have advantages in performing tasks in exploratory and hazardous domains that are not suitable for humans. The design of the SOS is however difficult because negative emergence with unwanted behaviors is likely to happen and multiple conflicting performance indicators need to be considered. To address this challenge, in this paper, we propose an SOS design method using surrogate models and the compromise Decision Support Problem (cDSP) construct. Surrogate models are used to capture the relationship between low-level rules or parameters and high-level emerging system performance. And the cDSP construct is used to explore "good enough" solutions (characterized by rule adoption rates) while managing the trade-offs among conflicting performance indicators. The efficacy of the proposed method is illustrated using a multi-agent box-pushing problem in the Webots simulation environment. It is shown in the results that our method leads to a 6.9% improvement in time efficiency, an 8.4% improvement in energy efficiency, and 26.2% in system reliability.