Accurate and timely mechanical fault diagnosis is essential for the safe operation of industrial equipment, yet in practical settings-characterized by small sample sizes and complex operating conditions-the accuracy and robustness of conventional methods remain limited. To address these challenges, we propose the spectrum analysis and perception multimodal large language model (SAP-MLLM), a fine-tuned multimodal large language model that accepts spectrogram inputs and produces multi-task outputs for industrial fault diagnosis. The architecture integrates a spectrum-aware visual encoder, a spectral vector encoder, and a cross-modal fusion module, enabling the model to capture image-based spectral patterns and parse per-frequency numerical amplitudes from vibration signals. Using multi-channel vibration data, we construct image and numerical spectra via weighted fusion and envelope spectral analysis, and evaluate the approach on industrial datasets featuring numerous fault categories, limited samples, and class imbalance. Experimental results demonstrate that, compared with existing deep learning models such as vision transformer and traditional methods, SAP-MLLM exhibits superior diagnostic accuracy and enhanced robustness under small-sample and complex operating conditions. In addition, the method supports on-premises deployment to protect data security, underscoring its practical relevance and potential for industrial adoption.
Decision-making is an effective way to resolve production anomalies, restore stable production, and ensure on-time order delivery. However, in actual manufacturing workshops, rescheduling not only increases adjustment costs but also introduces greater uncertainty. Although existing research on real-time and predictive decision-making can improve decision accuracy, balancing decision accuracy with production stability remains challenging. This paper proposes a two-stage decision-making method to enhance the applicability of decision-making techniques in manufacturing workshops. In the first stage, a decision recommendation model is constructed. To reduce data redundancy caused by random walk strategies, a subgraph model of production decision is developed. Then a decision recommendation model based on Finite-order DeepWalk is proposed to address the question of what level of decision-making is required. For production scenarios that recommend rescheduling, the second stage builds a scheduling generation model. Considering the real-time requirements of production decision-making, a deep reinforcement learning model is employed to learn historical decision strategies and enable precise adjustment. Finally, the proposed method is validated using a real manufacturing workshop. In the decision recommendation task, the proposed method achieves the best performance, with Accuracy, Precision, Recall, and F1-score reaching 0.9962, 0.9951, 0.9968, and 0.9959, respectively, outperforming all baseline models. In the scheduling generation task, the proposed method delivers feasible rescheduling plans and effectively limits completion deviations under production anomalies such as equipment failure and processing delay, thereby mitigating the influence of production anomalies on order delivery.
In the complex and dynamic discrete manufacturing environment, accurate prediction of the workshop operation situation (WOS) is crucial to ensure on-time delivery of orders. However, the spatio-temporal (ST) coupling characteristics of the manufacturing process, dynamic fluctuations of workshop performance, and varying contributions of samples to the prediction model make WOS prediction more challenging. To address these issues, this paper proposes a ST parallel ensemble learning approach for WOS prediction. Specifically, based on workshop production data, a temporal data model and a dynamic graph model are constructed to comprehensively characterize the ST characteristics of the production process. Subsequently, this paper proposes a ST parallel ensemble learning method, named Adaboost-GLT, which integrates three ST weak learners (GCN, LSTM, and TGCN) to effectively capture the ST characteristics. Furthermore, a dynamic optimal selection mechanism is designed to adaptively select the best-performing weak learner at each stage, enabling the prediction method to evolve synchronously with the dynamic changes of the manufacturing process. Additionally, a sample weight updating strategy that takes into account sample timeliness and prediction error is introduced to improve the rationality of Adaboost-GLT's attention allocation to samples during training. Finally, the performance of Adaboost-GLT is experimentally validated on real workshop production datasets. The experimental results show that Adaboost-GLT can fully exploit the ST characteristics, effectively cope with the dynamic fluctuations of workshop performance, and thereby achieve high-precision prediction of WOS.
As a critical stage in helicopter manufacturing, the assembly process relies on effective scheduling to ensure both efficiency and quality. Traditional, experience-based scheduling methods are often inadequate for complex shop-floor environments, especially given the heterogeneity of worker skills and complex technological constraints. Therefore, this article proposes a deep reinforcement learning approach based on proximal policy optimization with an attention mechanism (PPO-AM) to solve the helicopter assembly workshop scheduling problem (HASP) with consideration for worker skill proficiency. First, an assembly time prediction model was established that integrates workers' skill levels, task criticality, and the dynamic evolution of proficiency. This model provides a precise foundation for task-time estimation in assembly workshop scheduling. Based on this foundation, a shop scheduling model incorporating multiple process constraints was constructed. Furthermore, the state space, action space, and reward function were dynamically adjusted according to production progress, thereby formulating the problem as a Markov decision process (MDP). Within this framework, a PPO-AM method incorporating a self-attention mechanism was proposed, and an assembly shop scheduling agent was developed based on this method to enable flexible and efficient worker allocation in practical scenarios. The PPO-AM method leverages the self-attention mechanism to assess the importance of state features, enabling the agent to adaptively focus on critical information, thereby enhancing its state awareness and policy generalization capability. Experiments were conducted on assembly shop cases of various scales as well as in real-world engineering scenarios, and the results demonstrated that PPO-AM exhibits superior performance and practical value in complex assembly scheduling tasks.
To address the characteristics of complex aviation assembly process knowledge, including diverse knowledge types, complex hierarchical structures, strong semantic correlations, and continuous evolution, a top-down construction method for an assembly process knowledge graph (KG) is proposed. First, a multi-level assembly process ontology is developed to provide unified and standardized modeling of product information, assembly process knowledge, tools and equipment, historical cases, and their cross-level relationships, enabling structured knowledge representation and semantic consistency enforcement. Second, to cope with the semantic complexity of assembly process texts, a large language model (LLM)-based multi-agent collaborative triple information extraction method is proposed. By integrating ontology constraints and a self-consistency chain-of-thought (CoT) mechanism, the stability, accuracy, and interpretability of extraction results are significantly improved. Finally, a semantic embedding-based dynamic KG update strategy is adopted to maintain entity uniqueness and semantic consistency during continuous knowledge evolution. Experimental results demonstrate the effectiveness and applicability of the proposed method in complex aviation assembly process scenarios.
Digital twin (DT) model can accurately predict the future state of the shop floor and promptly identify potential problems, abnormal situations, or optimization opportunities. However, traditional production simulation method without considering the temporal characteristics of entities’ attributes. In the life cycle of physical entities, its attributes’ change will increase the DT simulation parameters’ error. Therefore, deep learning algorithms are used to model and evolve the simulation parameters of the digital twin shop floor (DTSF) to improve simulation accuracy. Firstly, the interaction mechanism between deep learning and discrete event simulation is designed. Then, a sequential regression variational autoencoder (SRVAE) is proposed to model the DT temporal parameters. Furthermore, the online instructor algorithm is proposed to update SRVAE through online data. This approach improves the simulation accuracy of DTSF while allowing its parameters to be self-maintained. And the effectiveness of the proposed method is verified by a case study.
In the manufacturing sector, abnormal production can disrupt production schedules, leading to significant economic and reputational losses for manufacturers. To address this issue, in this study, we present an explainable mechanism for production process anomalies (EM2PA) designed to clarify the complex coupling relationships among various manufacturing factors, analyze the impact of these factors on the production process, identify abnormal production, provide explanations for its causes, and enable trace-back analysis. EM2PA consists of three modules: the data augmenter, the influence factor recognizer, and the causal interpreter. Specifically, the data augmenter generates small sample data of abnormal production, the influence factor recognizer decouples the complex coupling relationships and identifies the factors influencing abnormal production, and the causal interpreter provides causal explanations. Furthermore, through a case study based on the actual production process of a discrete manufacturing workshop, we demonstrate the effectiveness of EM2PA in identifying the root causes of problems, while highlighting the importance of explainability and causal analysis of production process anomalies. This study presents EM2PA, an explainable mechanism for analysing production process anomalies. It enables trace-back analysis by identifying abnormal production, clarifying the relationships among manufacturing factors, and providing explanations for their causes.
Production anomaly has always been one of the main influencing factors that prevent discrete manufacturing workshops from maintaining stability and agility. Proactive anomaly detection can evaluate the production state and serves as a crucial foundation for preventive maintenance decision. Knowledge graph enables the use of multi-source manufacturing data as a data foundation for proactive anomaly detection. Although rich manufacturing data can comprehensively depict complex manufacturing process, constructing an accurate proactive anomaly detection model remains challenging because of insufficient analysis of the local and temporal features of the manufacturing process. This paper presents a link prediction model based on a deep graph neural network to solve the problem. Specifically, the manufacturing knowledge graph is constructed through OPC UA information model, Bert model and OWL semantic mapping model to organize multi-source heterogeneous data. The deep autoencoder model with local graph learning and the Seq2Seq model with attention mechanism are trained to analyze the neighboring relationship and the temporal correlation of the manufacturing elements, respectively. Finally, the link prediction model is designed by integrating both local and temporal features, with a restructured loss function to improve training effectiveness. Experiments suggest that the designed link prediction model has better prediction performance and is at least 25.6 % higher than the baseline models on the mean reciprocal rank.
In assembly sequence planning (ASP) of complex products, hierarchical structure, geometric feasibility, assembly tool changes, and assembly direction changes should be fully considered. Since the traditional information models of ASP are mostly static and abstract matrices, the outdated information reduces sequence rationality and increases time costs. To address these limitations, a novel decision-making framework based on dynamic knowledge graph (DKG) is proposed to build an intuitive semantic information model, planning the assembly sequences of complex products. An automated DKG generation method with updating mechanisms is designed to ensure that the DKG remains effective throughout the construction and maintenance process. A double-layer degree ordering algorithm (DDOA) is proposed to analyze sequence constraints within the DKG for obtaining a higher-quality assembly sequence. Comparative experiments demonstrated that the DDOA exhibits optimal performance. The solved assembly sequence possesses the best value of the objective function, with the fewest changes in assembly directions and assembly tools, as well as the shortest runtime.
Diagnosing spindle failures in CNC machine tools poses challenges due to data imbalance and multi-condition complexity. With normal samples vastly outnumbering failure samples, traditional models struggle to effectively identify rare-class failures. Insufficient feature extraction under multiple working conditions further compromises diagnostic accuracy. To address these issues, a parallel feature extraction network (PFEN) is proposed. This approach utilizes deep convolutional neural networks with wide first-layer kernel (WDCNN) to extract coarse-grained features from vibration signals. It innovatively designs a one-dimensional convolutional neural network (1DCNN) with Vision transformer (1DViT), integrating a parallel structure of 1DCNN and Transformer encoders to balance local details and global dependency features. A soft attention-based feature fusion method effectively combines WDCNN and 1DViT features, enhancing key feature representation. To address the issue of traditional loss functions overemphasizing majority classes and neglecting minority classes under imbalanced data conditions, this paper introduces a joint mechanism of weighted cross-entropy loss and contrastive loss. By dynamically adjusting weighting factors to enhance the importance of minority class samples, the model better learns rare and challenging features, thereby improving sensitivity to subtle differences in complex working conditions. Finally, experimental validation on the CK210sp CNC machine tool spindle imbalance fault dataset demonstrates that the proposed method (WDCNN-1DViT-PFEN) achieves superior diagnostic accuracy compared to existing methodologies while maintaining stable diagnostic performance across multiple working conditions. The results validate the effectiveness of this approach in addressing sample imbalance issues in CNC machine tool spindle fault diagnosis.
In discrete manufacturing workshop where disturbances occur frequently, the dynamic scheduling problem that considers transportation resource constraint is complex and challenging. Additionally, rescheduling without evaluating the impact of disturbances may adversely affect production stability of workshop. To address these issues, this paper proposes a dynamic scheduling framework based on digital twin and deep reinforcement learning. Specifically, the digital twin environment is constructed to provide a high-fidelity training environment for scheduling agents and serve as a simulation means for evaluating the impact of disturbances. Furthermore, a rescheduling trigger discriminator mechanism is designed to dynamically determine the necessity of rescheduling. In particular, the multi-agent proximal policy optimization with multiple critics (MAPPO-MC) is proposed to efficiently solve the discrete manufacturing workshop dynamic scheduling problem with transportation resource constraint. The innovation of MAPPO-MC lies in using the global critic to facilitate collaboration among scheduling agents from a global perspective, while employing individual critics to guide corresponding agents to learn their specialized scheduling knowledge from a local perspective, thereby achieving optimal scheduling decisions. Finally, extensive experiments have demonstrated the effectiveness of the proposed framework and the superiority of the MAPPO-MC. Under this framework, MAPPO-MC can respond promptly and effectively to disturbances in the workshop while ensuring stable production.
Nozzle health prediction is a pressing challenge in Surface Mount Technology (SMT) production, and traditional unimodal methods are limited by low accuracy and poor interpretability, which are difficult to meet today's production requirements. This study proposes an innovative method for nozzle health online prediction in SMT production, combining multimodal deep learning with a Multimodal Gated Attention (MGA) mechanism and an interpretability module. By integrating vacuum data and nozzle images, the method overcomes the limitations of traditional unimodal approaches, offering a comprehensive reflection of nozzle health. The Multimodal Fusion (MF) mechanism, based on the Kronecker product, along with the MGA mechanism, significantly improves prediction accuracy and stability. The interpretability module, incorporating Integrated Gradients (IG) and Gradient-weighted Class Activation Mapping (Grad-CAM), provides clear explanations for the model's predictions, supporting effective maintenance strategies. Experimental results demonstrate the superior performance of the proposed multimodal method, achieving an accuracy of 99.5%, precision of 100%, recall of 99.8%, and F1-score of 99.9%. This study presents a promising solution for efficient and reliable nozzle health prediction, combining high predictive accuracy with actionable insights for maintenance.
This paper utilizes an artificial neural network (ANN) model to predict the melt pool morphology (including melt pool width, height and depth) of Titanium (Ti) alloy by adjusting welding process parameters (laser power, welding current, and welding speed). The correlation between welding process parameters and melt pool morphology are analyzed using the Pearson correlation coefficient, which the result indicates that the input features are directly related to the output features. Four deep learning (DL) models are evaluated, which ANN model has higher coefficient of determination (0.8725 and 0.9691), lower root mean square error (RMSE) (0.4322mm and 0.1405mm) and lower mean absolute error (MAE) (0.025mm and 0.112mm). When the number of hidden layers is five, the R2 and RMSE values of ANN model are 0.8024, 0.9838, 0.5356mm and 0.1019mm, respectively. To assess the accuracy of the model, experiments are conducted to validate the predicted values, and the errors between the experimental and predicted values are 0.01mm, 0.10mm, and 0.02mm, respectively.
In the field of aviation manufacturing, the increasing variety of parts poses adaptability challenges for process template recommendation, especially for new types of parts. Collaborative filtering (CF)-based recommendation methods have limitations in capturing complex features of parts and templates, as well as in mining implicit interaction data, which affects their recommendation accuracy and adaptability in handling the complex relationships between parts and templates. Therefore, a two-channel CF recommendation algorithm is proposed, which combines a residual convolutional attention network (RCAN) and a gated graph convolutional network with initial residual and identity mapping (GGCNII). Firstly, statistical analysis is conducted on the aviation parts and process template features. An expert rating mechanism is employed to establish the correlation between “parts” and “process templates”. Based on this, initial embeddings for part features and process template features are obtained. Secondly, RCAN and GGCNII are utilized to extract node features and structural features in two separate channels. In the aggregation and propagation layers of the graph convolutional network (GCN), different-order collaborative signals of part features and template features are obtained through initial residual connections and identity mappings. During the node feature update phase, a gating mechanism is incorporated to further enhance the model’s flexibility and expressive power. Finally, the extracted features are fused for rating prediction, and the process template with the highest predicted rating is selected as the recommendation. Experimental comparisons were conducted between the proposed method and baseline methods using MovieLens-1M, MovieLens-20M, and aviation conduit datasets. The results demonstrate that the proposed process template recommendation method, a two-channel CF network based on the RCAN and GGCNII (RCAN-GGCNII-2C), outperforms baseline methods across various metrics.
In the manufacturing industry, the compilation of part machining process instructions relies on manual editing, leading to low efficiency, a high risk of errors, and difficulties in rework. An intelligent generation method for process instructions based on the combination of large-scale pre-trained language model (PLM) and multi-reasoning is proposed, aiming to improve the accuracy and automation of process instruction compilation. First, a knowledge graph (KG)-based multi-source heterogeneous knowledge representation model is constructed to support the structured processing of process information. Then, RoBERTa-ResBiLSTM-CRF model with attention fusion (RRCAF) is designed. Feature extraction of bidirectional long short-term memory (BiLSTM) is optimized through residual connections (ResBiLSTM), and an element-wise weighting method is used to fuse the output features of robustly optimized BERT approach (RoBERTa) and ResBiLSTM, significantly enhancing the performance of process element category recognition. At the decision-making level, a multi-reasoning model integrating feature query, case retrieval, and neural network reasoning is proposed. Case retrieval adopts a weighted fusion method based on the Jaccard similarity of process sequences and part features, while neural network reasoning achieves intelligent prediction of process parameters through an AdaBoost combined with a regularization-enhanced back propagation neural network (BPNN) algorithm (AdaBPNN-R). Experimental results show that compared with manual and existing instruction compilation methods, proposed method demonstrates significant advantages in the accuracy of process instruction compilation, offering an innovative solution for the automation and intelligence of process planning in manufacturing.
Multi-modal multi-objective optimisation problems (MMOPs) involve multiple equivalent Pareto sets that share the same Pareto front, and a special subclass is known as MMOPs with local Pareto fronts (MMOPLs). Conventional multi-modal multi-objective optimisation evolutionary algorithms (MMOEAs) struggle to effectively identify local Pareto fronts, and existing methods tailored for MMOPLs exhibit notable limitations. Additionally, commonly used performance metrics lack systematic evaluation and suffer from inherent shortcomings. A simple yet effective MMOEA that leverages the non-dominance range in the decision space is proposed to address these challenges. By prioritising individuals with above-average non-dominance ranges, the algorithm enhances the identification and retention of both global and local optima. Convergence is further improved using the local outlier factor method. The limitations of existing performance metrics are analysed and six new performance metrics tailored for MMOPLs are introduced. To facilitate evaluation, benchmark problems are modified to create scenarios in which global and local optima coexist on the same Pareto set. Extensive experimental results confirm that the proposed algorithm achieves competitive performance, effectively identifying global and local optima while ensuring well-distributed solutions.
Human-robot collaboration in industrial assembly demands real-time sensing for seamless actions. However, existing methods for perceiving industrial assembly behaviours, particularly in aerospace applications, still suffer from insufficient prediction accuracy, slow processing speeds, and limited transferability. This study proposes an assembly process recognition and prediction method tailored for aerospace human-robot collaborative assembly scenarios. By integrating behavioural feature parameter extraction techniques with frame-to-frame matching strategies, it effectively mitigates parameter complexity and reduces data bias introduced by human and environmental factors. An assembly behaviour prediction network named GCSM is constructed, incorporating attention mechanisms and skip connections to accelerate network training and prediction speeds while enhancing prediction accuracy. The experiments prove that the neighbour frame matching method can effectively identify the assembly progress with more than 97% accuracy. The prediction accuracy and real-time performance of the GCSM network are significantly improved, and it has a good migration ability, with about 50% improvement in RSE, 4% improvement in Correlation, 10% improvement in prediction speed, and 65% improvement in model training speed compared with the traditional prediction network. This method is applicable for assembly process recognition and prediction in complex human-robot collaborative assembly scenarios within the aerospace industry. See https://github.com/WeiZihan5/Human-Robot-Collaboration-Assembly-Dataset for dataset details.
Production anomalies, being one of the main causes of disrupted production schedules and product quality issues, have driven the manufacturing industry to focus on real-time monitoring and effective management, as these measures undoubtedly ensure production continuity and enhance efficiency. The complexity of discrete manufacturing workshops-characterized by diverse products, complex process routes, and frequent disturbances-leads to a corresponding complexity in the occurrence and evolution of production anomalies. Unlike point-to-point models for root cause analysis of production anomalies, this paper proposes a multi-level root cause analysis model for production anomalies to reveal the key influencing factors in their evolution process. First, to address the challenge of single-dimensional manufacturing data failing to effectively represent complex production states, production state representation models of manufacturing elements are built based on a first-order graph model of manufacturing elements, enabling consistent expression of production states. Second, a production anomaly evolution pattern analysis model based on a nonlinear Granger model is proposed to answer the questions of how production anomalies arise and evolve. Then, considering the imbalance in production anomalies, a meta-learning Transformer model is designed to learn evolution patterns of production anomalies and enable root cause analysis. Finally, using a real discrete manufacturing workshop as an example, the proposed method can accurately analyze the evolution patterns of production anomalies. In the evolution pattern learning task, it achieves better learning performance and is at least 69.5 % higher than the baseline models on the root mean square error. Additionally, the method achieves an accuracy of 81.67 % in identifying the top three root cause states of production anomalies. The research results demonstrate that nonlinear networks can effectively analyze the complex evolution processes of production anomalies and enhance the Granger model's accuracy in identifying the evolution patterns. The meta-learning framework improves the generalization ability of the evolution pattern learning model, enabling more precise root cause identification. Consequently, the proposed method offers a new perspective for evolution analysis and root cause analysis of production anomalies.
Effective and efficient online quality inspection of poor soldering in Surface Mount Technology (SMT) solder joints, whose defect characteristics are not visually discernible, remains a formidable obstacle in the electronics manufacturing industry, despite the availability of various inspection methods. To overcome this challenge, a novel multimodal fusion method for soldering quality inspection is proposed. First, the method innovatively introduces information from different modalities as input to the inspection model, with the aim to provide more comprehensive information for the decision making of the inspection model. Then, a combination of multimodal gated attention and tensor fusion is used to fuse the features extracted from each modality to form a comprehensive multimodal representation. Finally, this multimodal representation is used to conduct soldering quality inspection. The experimental results demonstrate that the proposed method improves the detection rate of poor soldering significantly, from 93.6 to 99.4%, with a precision level of nearly 100%. The detection rate for visually unapparent soldering defects increases from 49.3 to 95.4%. This meets both manufacturing and customer requirements and to some extent addressing the industry challenge of online SMT soldering quality inspection. This remarkable performance surpasses that of the six mainstream ResNet, ResNext, RegNet, ShuffleNet, EfficientNet and NoisyNet inspection models currently available.
Cloud manufacturing (CMfg), as a manufacturing mode to realize resource collaboration and service sharing, can help enterprises reduce costs, increase efficiency and enhance competitiveness. Existing CMfg relies on historical service evaluation results as the basis for task matching and thus lacks effective methods for the precise assessment of the adaptability to current tasks and the sustainability for future endeavors; throughout the service process, deficiencies occur including a lack of real-time assessment, prediction, and optimization capabilities for service execution outcomes. Motivated by the shortcomings of existing CMfg, this article focuses on how to enhance the dynamism and timeliness of the CMfg service process. A novel digital twin (DT)-driven CMfg service model is proposed. Specifically, a framework and an operation mechanism of DT-driven CMfg service are proposed; three key technologies related to DT in CMfg service are proposed, including model migration based matching modeling technology, multi-agent reinforcement learning based CMfg service composition optimization, and DT data-driven CMfg service performance prediction, which provide a theory and method for the DT and CMfg integration construction. The experiment demonstrates that the proposed methods have good performance for the production of cast aluminum engine fan bracket and provide a holistic understanding of DT in CMfg service.