INTERNATIONAL JOURNAL OF PRODUCTION RESEARCH(2026)
Nanjing Univ Aeronaut & Astronaut
被引用0|浏览0
摘要
As the demand for personalised customisation increases, manufacturing enterprises are continually enhancing their capabilities to meet diverse market requirements. This paper investigates a two-stage assembly flowshop scheduling problem, focusing on minimising total tardiness while accounting for resource flexibility and dynamic product arrivals using multi-agent deep reinforcement learning (MADRL). To explore the operational benefits of resource flexibility in scheduling environments, a skill matrix is introduced to assess the flexibility level of the processing machines. The problem is modelled as a Markov decision process (MDP), and a Multi-Agent Proximal Policy Optimisation algorithm with an Epsilon-greedy approach (E-MAPPO) is proposed for product sequencing and processing machine allocation. Experimental results show that, across different numbers of arrival products and flexibility levels, the scheduling agent trained with the E-MAPPO effectively learns to apply suitable dispatching rules. Notably, the algorithm outperforms 18 composite dispatching rules and two other deep reinforcement learning algorithms, namely VDN and standard MAPPO. Our findings demonstrate that enhanced resource flexibility positively impacts scheduling performance, which indicates that enterprises must balance the benefits of flexibility against its implementation costs to satisfy customer demand.