Vibration signals are often severely disturbed by noise during the acquisition process. To effectively eliminate noise interference, a direct noise learning-based multiscale dynamic weighted (M-DW) multidimensional joint (M-DJ) residual convolutional autoencoder model is proposed. Firstly, the combination of a convolutional autoencoder and residual learning is utilized to learn and reconstruct noise. Meanwhile, a M-DW module is added to effectively extract multiscale features and adaptively adjust feature weights according to their denoising contribution. Furthermore, a one-dimensional to two-dimensional (1D-2D) and 2D-1D M-DJ framework is constructed, which effectively integrates adjacent M-DW features using a 2D convolutional neural network and reconstructs noise using a 1D convolutional neural network. The experimental results show that the network effectively filters out noise components in constant-speed, variable-speed and high-noise conditions.
The flexible architecture of the open radio access network (O-RAN) provides effective support for the deployment of federated learning (FL). However, existing FL schemes in wireless networks often suffer from packet errors, which lead to reduced model test accuracy and increased training delay. In this paper, we propose a robust FL scheme in O-RAN, named adaptive retransmission-based hierarchical FL (HFedAR), which improves model test accuracy and reduces training delay. Specifically, HFedAR first analyzes the similarity of local gradients using a clustering algorithm to associate users with edge servers. Subsequently, we employ an adaptive retransmission mechanism for both edge and global aggregation, thereby facilitating rapid convergence of FL on non-independent and identically distributed data. Considering the limitations of spectrum resources and user energy, we formulate a multi-objective optimization problem to minimize FL training loss and delay. Due to the implicit nature of the objective function, the coupling of decision variables, and network dynamics, it is difficult to solve the problem through traditional convex optimization and machine learning algorithms. Therefore, we derive a convergence upper bound for HFedAR and present a two-stage resource allocation algorithm. The algorithm can jointly make optimal user power and computing resource allocation decisions in the first stage, and user scheduling, retransmission selection, and radio resource block allocation decisions in the second stage. Extensive simulation results on the MNIST and CIFAR-10 datasets demonstrate that the HFedAR scheme significantly improves model test accuracy by 3.5% and 6.5% and reduces training delay by 40.1% and 34.2% compared with existing FL benchmarks.
In this paper, we study an intent-driven service level objective (SLO)-aware routing problem for cloud-edge multi-domain networks, in which service intent signals are provided by an upstream large-model-assisted intent inference component to guide automated routing and orchestration. The key challenge is that inferred intents can be uncertain and directly acting on hard (top-1) intent predictions may mislead the closed-loop controller, causing unnecessary SLO violations and performance instability, especially when the network operates near congestion. To address this issue, we propose a lightweight and deployable Soft-Intent Guard (SIG) framework that sits between intent inference and control to make intent usage more reliable, without changing the underlying routing policy. The proposed framework consists of two simple mechanisms: 1) soft intent usage, which treats inferred intents as graded signals to avoid abrupt control changes, and 2) confidence-based guarding, which triggers a safe fallback behavior when the intent signal is unreliable. SIG is designed to mitigate intent uncertainty propagation in closed-loop control, where small inference errors can be amplified into large performance degradation. Simulation results in a multi-domain routing setting with controlled intent error rates show that SIG can consistently enhance robustness: at a critical load point, SIG improves SLO satisfaction by up to 6.13 percentage points for the most affected service class, while keeping tail latency and packet drop behavior comparable to benchmark schemes.
Multi-source bearing fault diagnosis for industrial edge deployment requires a careful balance among diagnostic accuracy, robustness to noise, and inference latency, especially under few-label conditions. To address this challenge, we propose a jointly optimised diagnostic framework that combines an attention-enhanced wide-kernel convolutional neural network (ACWDCNN) with a GRU autoencoder (GAU). The ACWDCNN front end extracts and reweights fault-relevant features, whereas the GAU imposes reconstruction-based temporal regularisation during candidate training. Hyperparameters are selected using QPSO-AWA, which searches a mixed discrete-continuous design space spanning both structural and training-related variables. During candidate training, the GAU influences representation learning through the loss term, whereas QPSO-AWA ranks candidate configurations using a validation objective; at inference, the deployed model remains a fixed sequential pipeline. Experiments on the NSK 6205 DDU dataset and a motor compound-fault dataset show that the framework achieves 98.5% accuracy on the NSK dataset and 97.4% on the motor dataset, remains robust at-4 dB SNR, and satisfies edge-deployment constraints with an inference latency of 1.8 ms on an RTX 4060 and 9.8 +/- 0.5 ms on a Jetson Nano. These results indicate that the proposed framework offers a practical balance among diagnostic performance, noise robustness, and computational efficiency for few-label, multi-source bearing fault diagnosis.
As a core transmission component of rotating machinery, the gearbox's early weak fault features are often masked by strong background noise under actual working conditions. Moreover, acquiring sufficient fault samples is difficult, posing significant challenges to high-precision intelligent diagnosis. Existing pure data-driven methods lack deep utilization of mechanical physical mechanisms, and traditional static convolutions struggle to adaptively track transient impact features in non-stationary time-varying signals, leading to limited model generalization performance under extreme working conditions. To this end, this paper proposes a few-shot fault diagnosis method integrating a physical mechanism knowledge graph, named Fault Concept-centered Knowledge Graph (FCKG), with dynamic adaptive spatio-temporal signal processing. First, an FCKG centered on the gearbox fault evolution mechanism is constructed to transform multivariate time-series signals into a time-series graph embedding physical semantic correlations. Second, a Knowledge Graph-guided Dynamic Adaptive Sequential Convolution for Time-series (KG-DASC-T) is proposed. Serving as a semantic-energy dual-driven time-varying filter, this module can accurately track and enhance the sensitive frequency band features of weak faults. Building upon this, a Cross-Scale Residual Coupling Unit (CSR-CU) is utilized to capture the long-range dynamical dependencies of fault evolution, and a Lightweight Temporal Attention (LTA) mechanism is employed to focus on critical transient impact responses. Validation results on datasets from the Politecnico di Torino and the laboratory QPZZ-II gearbox test rig demonstrate that the proposed method exhibits excellent robustness under extreme few-shot (only 5 samples/condition) and strong background noise scenarios, achieving diagnostic accuracies of 99.8% and 98.7%, respectively. Furthermore, the model has a parameter size of only 0.35M, fulfilling the requirements for lightweight deployment and real-time diagnosis on industrial edge devices while achieving a deep synergy between physical mechanisms and data-driven approaches.
The growing scale of Large-Model (LM) Artificial Intelligence (AI) service deployments brings heterogeneous requests with distinct Service-Level Objectives (SLOs), thus posing a challenge to achieve long-term service quality at scale. In this paper, we investigate multi-class routing for LM AI services deployed across multiple autonomous domains interconnected over a wide-area network (WAN). In such WAN-scale settings, privacy and autonomy policies limit sharing of domain-internal states, and heavy cross-domain interfaces hinder their deployment. Under this limited visibility, we formulate joint service-domain selection and in-domain dispatch in a constrained stochastic optimization framework to improve long-term service quality. To address this problem, we propose a Two-Layer coordination framework with Two-Stage execution (TLTS) that separates slow-timescale cross-domain coordination from fast-timescale per-request execution. The Coordination layer compresses policy-compliant aggregated feedback into a per-class pressure signal that summarizes sustained SLO stress, so that exchanging only this pressure enables Two-Layer coordination without sharing domain-internal states. We further develop a Service-Identification-Class based TLTS (SIC-TLTS) algorithm to instantiate TLTS, where the Execution layer performs Two-Stage per-request decisions conditioned on the SIC label and the latest pressure. Simulations on a real-world backbone topology show that SIC-TLTS consistently improves service quality while satisfying per-class SLO constraints, compared with representative baselines.
To solve the fault diagnosis difficulties in autonomous underwater vehicle (AUV) thrusters, a semi-supervised AUV fault diagnosis method based on dynamic decay learning strategy and hypergraph attention network (HGAN) is proposed. Firstly, an attention mechanism is introduced into hypergraph convolutional networks (HGCN) to construct HGAN. Then, the HGAN and graph convolutional network (GCN) are designed in parallel architecture to capture both the dynamic and static features of the input graph signal simultaneously. Finally, a dynamic decay learning strategy is introduced, which improves the training efficiency of the proposed model. The diagnostic precision of the proposed approach is verified through experiment analysis. Besides, its superiority over the other related methods is also verified through comparison study.
Deterministic transmission and computation are essential for open radio access networks (O-RANs) to support latency-critical and computation-intensive applications. How ever, existing O-RAN and time-sensitive networking integrations mainly focus on deterministic guarantees in wired fronthaul transmission, while lacking a unified mechanism to coordinate wireless transmission, wired transmission, and computation. This limitation makes it difficult to provide bounded end-to-end latency under time-varying wireless channels and increasing computational demands, which reduces flow scheduling success rates and causes resource wastage. In this paper, we propose a hierarchical deterministic O-RAN framework, named DetO RAN, which ensures deterministic transmission and computation for flows through a unified queuing and resource allocation mechanism. DetO-RAN introduces a three-queue model (i.e., wireless, wired, and computing queues) to support transmission and computation of flows, and establishes a closed-loop resource allocation process based on model training, inference, and updating. Based on this framework, we formulate a multi-objective optimization problem aiming to maximize the flow scheduling success rate and resource utilization. Due to the time-varying wireless channels, strongly coupled resources, and network dynamics, traditional optimization algorithms are inefficient. Thus, we decouple the problem into a radio resource block (RB) allocation subproblem as well as a time slot and computing re source allocation subproblem. Furthermore, we propose a meta learning-based dual-timescale resource allocation (ML-DTRA) algorithm to solve them. Specifically, ML-DTRA performs RB allocation via a graph neural network at a small timescale and makes time slot and computing resource allocation decisions via deep reinforcement learning at a large timescale. Extensive simulation results demonstrate that the ML-DTRA algorithm significantly improves the scheduling success rate by 12.7% and the resource utilization by 15.9% compared with state-of-the art benchmarks, showing its effectiveness in supporting end-to end deterministic transmission and computation while improving resource efficiency in dynamic O-RAN environments.
Microservice architecture dominates cloud data centers due to its agility and scalability, yet optimizing microservice deployment and user request routing for strict Quality of Service (QoS) remains challenging. Existing approaches often overlook the frequent data interactions between microservices and their distinct resource consumption patterns, which cause aggravated performance bottlenecks as the microservice diversity scales. In this paper, we propose a model that evaluates service latency by modeling user requests as multiple chains and leveraging parallel and serial processing characteristics. Building on this, we formulate the joint optimization problem of microservice deployment and multi-chain request routing to reduce the average service latency and improve load balancing. To tackle this challenge, we first introduce the Communication-Aware Microservice Grouping (CAMG) algorithm, which minimizes cross-server communication by grouping strongly data-dependent microservices into the same deployment units. We then propose the Transformer Enhanced Twin Delayed Deep Deterministic Policy Gradient Microservice Deployment (TETD3-MD) algorithm, where a Transformer-based encoder extracts resource type features to optimize load balancing. Integrated with the TD3 framework, this algorithm generates deployment decisions implementing CAMG's group or individual deployments. Furthermore, we design an A-star queue-aware multi-chain scheduler that dynamically routes user request chains to suitable microservice instances by integrating queue states, thereby reducing service latency. Extensive simulations demonstrate that the proposed algorithms reduce service latency by 22.5% and improve load balancing efficiency by 35.1% compared to the latest baselines, while adapting to varying scenarios.
In Industrial Internet of Things (IIoT) networks, Time Slotted Channel Hopping (TSCH) networks play a critical role in providing deterministic wireless connectivity. However, traditional static schedulers struggle to adapt in real time to dynamic operational modes such as emergency traffic bursts, while standard Deep Reinforcement Learning (DRL) approaches suffer from poor sample efficiency and limited generalization across heterogeneous facilities. To address these challenges, we develop DETER (DEterministic Topology-aware mEta-learning Resource scheduler), an intelligent scheduling framework designed to guarantee deterministic communication in IIoT networks. The DETER framework decouples policy optimization from schedule execution: a topology-aware meta-learner generates link-level scheduling weights, which are then converted into candidate TSCH schedules by a priority-based scheduler and verified by a network-calculus-based validator. This design enables fast adaptation while preserving strict quality-of-service compliance and feasibility guarantees. By acquiring transferable meta-parameters, the proposed framework can be efficiently fine-tuned to previously unseen operational scenarios within only a few gradient updates. Simulation experiments demonstrate that the DETER framework generalizes effectively to unseen network topologies and achieves rapid convergence in dynamic environments, significantly outperforming state-of-the-art DRL approaches in terms of deterministic performance guarantees, specifically reliability and latency.
The rapid development of cloud computing, coupled with the proliferation of microservices, has led to an explosive growth in data processing demands. To overcome the limitations of the default Kubernetes scheduler, which only considers CPU and memory, this study proposes DQKS-Double Deep Q-Network-based Kubernetes Scheduler. DQKS models scheduling as a Markov Decision Process, using CPU, memory, disk I/O, network usage, and application demands as inputs. It learns optimal strategies by interacting with the cluster and evaluates decisions based on latency, resource utilization, and load balance. For model training, a resource collector is designed to gather local resource usage data from worker nodes. Evaluation results show DQKS improves resource utilization by 17.10%, enhances load balance by $1.69 \times$, and reduces decision latency to $0.85 \times$ compared to the default scheduler.
This paper addresses the issues of low efficiency in topology discovery and link state information collection, as well as the high overhead of control messages in software defined networks. To address these issues, an efficient collection algorithm based on graph partitioning is proposed. The algorithm uses graph partitioning techniques to divide large-scale networks into multiple subgraphs. Information collection is then performed in parallel within each subgraph. This approach significantly improves collection efficiency. The proposed algorithm is validated through theoretical analysis and simulation experiments. The results show that the algorithm effectively reduces control message overhead.
The emerging intelligent services, spurred by the rise of the intelligent Internet, are placing multidimensional requirements on the network to collaboratively guarantee computing and networking resources. In this article, we propose a unified end-to-end intelligent resource scheduling method for converging networks [e.g., Internet of Things (IoT)], which can always globally abstract the available resources from different networks with a unified model description, and jointly planning the resources from end-to-end by deep reinforcement learning (DRL) algorithms supporting both discrete and continuous variable decisions. The method proposes a three-layer architecture, including service layer, network layer, and adaption layer, which aims at optimizing the flow transmission performance. Through the general Markov decision process (MDP) transformation from the model, the DRL-assisted algorithm can further solve the optimization problem. We categorize heterogeneous network resource scheduling into horizontal and vertical scenarios, applying the proposed architecture to both. Compared with the existing diverse learning (DiLearn) and naive (DiNaive) approaches, the proposed approach is not only time-saving but also can schedule 28.4% and 8x more flows in horizontal scheduling scenarios, and improve 54.2% and 3.5x flows in vertical scheduling scenarios, respectively.
Computing-aware networks (CAN) can provide ubiquitous AI inference services for intelligent applications. However, due to the huge differences in the demand for inference services of intelligent applications and the continuous innovation of computing devices, traditional protocol-based scheduling methods make it difficult to schedule complex heterogeneous computing resources. In this paper, we propose a CAN resource scheduling method, named SMAD, which can adapt to external environmental changes without human modification of the protocol mechanism. Aiming at the scheduling problem of complex heterogeneous computing resources and concurrent random diverse inference tasks, a constrained multi-objective optimization problem of scheduling service quantity, accuracy, and delay is formulated. Through the general Markov Decision Process (MDP) transformation from the model, the Deep Reinforcement Learning (DRL)-based AI-native scheduling algorithm framework can further solve the optimization problem. Meanwhile, the AI-native framework aggregates diverse device states into a data tensor, integrates the DRL algorithm in a data-driven manner to generate a dynamic action tensor, and reversely drives full-stack resource scheduling for global closed-loop optimization in CAN, aligning with inference service demands. Extensive simulation results show that the proposed SMAD has good convergence performance. Compared with the traditional DRL algorithm, it significantly increases the number of concurrent schedulable tasks and reduces the inference service delay.
Convolutional neural network (CNN) was widely applied to the data-driven-based fault diagnosis. However, it often needs to artificially transform the signal into a 2-D image with the help of time-frequency transformation; furthermore, the alternative 1-D CNN can only extract single-scale features but cannot adaptively reveal the relationship between scales even if 1-D convolution kernels of multiple scales are applied at the same time. Accordingly, this article proposes a mixed CNN model, namely, a multiscale dynamic weighted 1-D to 2-D (1D-2D) CNN model (1D-2D MDWCNN), which resembles multiwavelet-based CNN method but uses different scale convolution kernels in 1-D convolution neural network to facilitate the multiscale feature extraction and accomplish adaptive features fusion in the constructed 1D-2D seamless joint network framework. Specially, to reflect the global contribution of different convolution kernels, a dynamic weighted (DW) network layer is constructed to adaptively adjust the global weight of each convolution kernel, so as to improve the fault diagnosis ability. The validity of the model is verified by motor bearing fault diagnosis experiment and gearbox fault diagnosis experiment. The diagnosis results proved the developed 1D-2D MDWCNN model superior to the latest CNN-based models.
Under the dual influence of time-varying working conditions and noise interference, accurately faults diagnosing in mechanical equipment poses significant challenges. Therefore, this paper proposes a limited sample fault diagnosis method using dilation kernel gated recurrent dropout attention unit (GRDAU) for time-varying speed based on interference suppression. Firstly, dilation kernel parameters were integrated into traditional convolution to suppress high-frequency noise. Secondly, to enhance the model’s robustness against noise interference and variations in speed, a GRDAU was developed based on gated recurrent unit (GRU). Additionally, a global cyclic dynamic decay learning strategy was implemented within the GRDAU to better adapt to complex speed variation conditions. Finally, two case studies were conducted to validate the robustness and interference suppression capabilities of the GRDAU. When compared to a range of existing advanced diagnostic methods, it demonstrated superior performance and stronger generalization ability.
A single type of sensor signal cannot fully represent the operational status of mechanical equipment, leading to incomplete state characterization and inaccurate diagnostics. This paper proposes an innovative fault diagnosis method based on Convolutional AutoEncoders combined with multivariate information fusion to accurately identify the overall health status of bearings by analyzing various sensor data. Our approach leverages the Convolutional AutoEncoders to effectively integrate heterogeneous sensor data from multiple sources, including vibration and sound signals, with data augmentation and normalization techniques for preprocessing, thereby improving the model’s generalization capability and accuracy. Furthermore, the integration of K- means clustering and a Sparse Attention mechanism enables precise recognition of critical fault features. The model’s effectiveness is validated through comprehensive performance evaluation using confusion matrices and visualization techniques. Experimental results demonstrate that our method achieves high accuracy and robustness in fault diagnosis tasks, offering a significant advancement in intelligent maintenance and fault prediction of rolling bearings by addressing the limitations of traditional methods.
Autonomous underwater vehicles (AUVs) are widely used in ocean exploration, scientific research, and other fields that perform tasks in complex underwater environments. Since propellers fault samples are very scarce and difficult to collect in AUV practice, traditional fault diagnostics face the challenge of insufficient data. For this reason, a channel attention residual transfer learning (ECRTN) model based on dual-loss nonlinear independent component estimation (DLNICE) is proposed to expand data for few-shot fault diagnosis of AUVs. Specifically, DLNICE is first used to augment fault samples by combining both time and frequency-domain information of AUV propellers vibration signals with a dual-loss function. The augmented fault samples and normal data are fed into the channel-attention residual network, which is fine-tuned by the large language model. The fine-tuned model is then used to diagnose real vibration samples in the target domain. Experimental results show that the proposed ECRTN achieves an average diagnosis accuracy of 96.81 %, outperforming other state-of-the-art fault diagnosis models for dealing with few-shot fault diagnosis tasks. The present method provides an effective technical solution for the few-shot fault diagnosis of AUV propellers and has a better prospect for practical applications.
Industrial Internet of Things (IIoT) applications, such as industrial process control, demand ultrahigh reliability and bounded delay. The reliable and available wireless (RAW) initiative within the Internet engineering task force DetNet working group addresses these needs by applying IEEE 802.15.4 time-slotted channel hopping (TSCH) technology and leveraging techniques like packet replication, elimination, and ordering functions (PREOFs) to ensure deterministic performance for IIoT. However, while PREOF improves reliability, its redundant transmission mechanism inevitably increases energy consumption, conflicting with the energy constraints of TSCH nodes. The existing multipath routing approaches struggle to address this challenge, failing to jointly consider both energy efficiency and deterministic performance. Additionally, these approaches often overlook the delay variation caused by multipath transmissions of different lengths-a key factor that can undermine deterministic performance by increasing buffering requirements and affecting the predictability of data flows. In this article, we investigate a multipath optimization problem aiming at improving energy efficiency and minimizing delay variation while meeting the requirements of bounded reliability and delay for deterministic flows. Considering the above multipath routing optimization problem, which aims to satisfy multiple objectives under multiple constraints, is typically NP-hard, solving these challenges with traditional methods is highly complex. Thus, we further propose a energy-efficient multipath routing (EEMR) algorithm that utilizes deep reinforcement learning (DRL) to optimize the multipath selection, effectively enhancing energy efficiency for deterministism. EEMR can be extended to solve optimization problems in holistic-deterministic multidomain scenarios, such as smart factories integrating 5G and DetNet. We compare the performance of our proposed method with several baseline methods. Empirical evaluations show that EEMR significantly reduces energy comsumption and delay variation compared to baseline methods under various environment settings.
Unmanned Aerial Vehicle (UAV) networks are essential for data transmission in emergency scenarios, serving as relays to transmit data from ground users to base stations. Traditional UAV routing focuses primarily on unicast routing, which requires that a specific destination be identified for the data before transmission begins. This approach encounters significant challenges in dynamic networks due to frequent topology changes. When multiple base stations are available within the network, routing data to several base stations can enhance transmission efficiency. However, existing routing algorithms are not well-suited for such scenarios. This paper redefines the routing of UAV networks with multiple base stations as anycast routing tailored for dynamic networks. We introduce a distributed anycast routing method named QAR to boost data transmission performance. In the QAR, Q-learning parameters are dynamically adjusted, and a multi-base station transmission value function is crafted to calculate rewards and update the Q-table. Simulation results indicate that QAR surpasses existing Q-learning based routing methods in multiple base station scenarios, delivering superior performance in terms of delay, packet delivery ratio, and throughput.