Data-driven deep neural networks have significantly enhanced high-resolution Direction of Arrival (DOA) estimation. However, standard convolutional architectures rely on localized receptive fields, making them deficient at capturing global information. While advanced network models can aggregate global features, they introduce substantial parameter burdens and computational costs. This local-global trade-off limits DOA estimation performance in challenging acoustic environments, particularly under low SNR, narrow angular separations, and limited snapshots. To overcome these bottlenecks, a Residual Fourier Network (ResFFTNet) is proposed. The architecture introduces a task-specific local-global residual mechanism, in which a Fourier-domain branch is embedded into the residual module and operates in parallel with the conventional convolutional pathway. The convolutional branch captures local spatial-spectrum characteristics, whereas the Fourier branch establishes global interactions across the angular spectrum, enabling complementary local-global feature fusion for high-resolution DOA reconstruction. By fusing these representations, the model enhances feature extraction and DOA estimation accuracy. Furthermore, the proposed residual Fourier structure enables faster convergence during network training. Simulations demonstrate robust estimation under low SNR, narrow angular intervals, and limited snapshots. Finally, experimental validations using SWellEx-96 S59 and South China Sea trial data corroborate the practical effectiveness of the proposed architecture in complex underwater scenarios.
The rapid advancement of domain adaptation algorithms has significantly accelerated the deployment of intelligent diagnostic technologies. However, existing domain adaptation methods predominantly focus on closed-set fault diagnosis and rarely address data privacy concerns, limiting their applicability in industrial settings. To this end, a universal source-free domain adaptation method is proposed. Initially, a source model is pre-trained using labeled source data. This pre-trained model then processes the target data to decompose the target features into common class components and unknown class components, while simultaneously generating target prototypes and source anchors. Subsequently, the distribution of the unknown class components is estimated using a Gaussian mixture model with two components. Finally, a confidence estimation strategy is developed to derive instance-level decision boundaries by evaluating the distance between target prototypes and source anchors, thereby completing the classification task. Experimental results on gearbox and rolling bearing datasets demonstrate that our approach excels in handling fault diagnosis under varying conditions while ensuring data privacy.
It is essential to perform fault acoustic source localization (FASL) for the health maintenance of mechanical systems under strong noise and multiple interference environments. While sparse optimization methods offer high localization accuracy and excellent generalization, their performance heavily depends on manually tuned hyper-parameters. On the other hand, deep neural network can learn network parameters from training data to perform fault localization directly, but its generalization capability and physical interpretability are weak. To address these issues, a hyper-Parameter Learning Network (PLN) is proposed to perform high-accuracy fault source localization rapidly and adaptively. The PLN integrates sparse optimization with learnable hyper-parameters (step size and shrinkage threshold). These parameters are trained in a end-to-end way via MSE minimization, which combines model-driven robustness with data-driven adaptability. Simulations and experiments show that the PLN significantly improves both localization accuracy and convergence speed over sparse methods, and provides an efficient and scalable solution for FASL in complex environments.
Deconvolution beamforming (DecBF) is a robust and high-resolution way to estimate direction of arrival (DOA) of interesting acoustic sources in sonars and radars. However, its estimation accuracy is unsatisfactory under low SNR and few snapshot conditions, due to that deconvolution procedure is a highly ill-posed problem. Leveraging the sparse distribution pattern of target sources in the angular domain, a non-convex sparse deconvolution beamforming (NSDB) method is proposed by incorporating a sparse regularization function (scale-invariant ℓp/ℓq norm) into the optimization model of the DecBF. The highlight of NSDB is that the ℓp/ℓq norm approximates the ideal sparse function ℓ0 norm reliably to describe the focused sparse pattern, and also mitigates the ill-posedness of DecBF effectively. Therefore, the recovered beam has very low sidelobe levels and sharp peaks in the target-source directions. Moreover, a smoothing strategy is employed to minimize the ℓp/ℓq function and FFT-based fast convolution is also utilized to accelerate algorithmic convergence. Simulation results show that the proposed NSDB method has excellent beam resolution, sidelobe suppression, and noise immunity under all SNRs, array aperture, and few snapshot conditions. Anechoic chamber experiments are conducted, and comparative results verify that the proposed NSDB method achieves state-of-the-art DOA estimation accuracy and is feasible and effective in practical acoustic scenarios.
This letter proposed a sparse deconvolution localization method (FFT-L1ML2) driven by non-convex L1−αL2 regularization that more closely approximates the ideal L0 norm. It is an alternative that explores the sparse structure of sound sources to enhance localization accuracy, while the original sparse deconvolution beamforming lacks a sufficiently accurate sparse description. An optimization solver composed of forward gradient descent and backward proximal operator is then developed for the FFT-L1ML2 model to reconstruct the beamforming map. Both simulation and experimental results show the effectiveness and superiority of the proposed method in localization accuracy, energy concentration, pseudo source reduction, and computational cost.
The classification of MTS is a common challenge with wide-ranging applications across different fields. Although deep learning methods have demonstrated potential in this area, graph neural networks (GNNs) have become a effective method for capturing the complex interconnections within MTS data. However, most current GNN-based methods depend on predefined graphs, which restricts their ability to adjust to changing relationships between time series samples. To overcome this limitation, we present MTSGNN (Graph Neural Network for MTS), an innovative data-oriented model for classifying MTS. MTSGNN integrates a graph learning module to capture intricate relationships between time series samples and a graph interaction module to facilitate the exchange of information between static and dynamic graph representations. The multi-layer perceptron (MLP) is used for the down streaming classification task. Through experiments on ten real-world MTS datasets, it shows that MTSGNN outperforms current state-of-the-art methods significantly. This emphasizes the effectiveness and potential of MTSGNN for real-world applications.
Impulsive blind deconvolution (IBD) is a popular method to recover impulsive sources for bearing fault diagnosis. Its underpinnings are in the design of objective functions based on prior knowledge of impulsive sources and a transfer function to describe transmission path influences. However, popular objective functions cannot retain waveform impulsiveness and periodicity cyclostationarity simultaneously, and the single convolution operation of IBD methods is insufficient to describe transmission paths composed of multiple linear and nonlinear units. Inspired by the MaxPooling period modulation intensity (MPMI) and convolutional sparse learning (CSL), an adaptive multi-D-norm-driven sparse unfolding deconvolution network (AMD-SUDN) is proposed in this paper. The core strategy is that one target vector with simultaneous impulsiveness and cyclostationarity is constructed automatically through the MPMI; then, this vector is substituted into the multi D-norm to design objective functions. Moreover, an iterative soft threshold algorithm (ISTA) for the CSL model is derived, and its iterative steps are unfolded into one deconvolution network. The algorithm’s performance and the hyperparameter configuration are investigated by a set of numerical simulations. Finally, the proposed AMD-SUDN is applied to detect the impulsive features of bearing faults. All comparative results verify that the proposed AMD-SUDN achieves a better deconvolution accuracy than state-of-the-art IBD methods.
Direction-of-Arrival (DOA) estimation is widely applied in acoustic source localization. Recent deep learning methods have achieved satisfactory estimation accuracy, attributed to their excellent nonlinear description capabilities in exploiting latent structures embedded in datasets. However, the prior knowledge of array structure is often ignored in the design stage of these networks, resulting in a compromise on model transparency. To address this issue, we propose an interpretable, sparse optimization driven framework for DOA estimation called Array Manifold Integrated LISTA (AMI-LISTA). Firstly, the array manifold matrix is fixed as a constant dictionary in every iterative step and integrated into sparse optimization backbone network LISTA, and thus the intrinsic knowledge of array structure is maintained to ensure network interpretability. Secondly, our approach introduces a layer-fused strategy into the loss function. The layer-fused mechanism facilitates incremental improvements and enhances training stability, leading to faster convergence. Numerical experiments confirm the excellent interpretability of AMI-LISTA and its superior performance, particularly under small angle interval conditions.
Robot technology equipped with vision system is expected to facilitate deep, in-hole, in-situ nondestructive recognition of water content in loess slopes. This advancement could considerably aid in the study of the spatial and temporal evolution of loess water content. However, a robust image recognition method capable of accurately recognizing the water content in loess across different regions is essential for the widespread adoption of this technology. Thus, we collected loess samples from the western (Lanzhou), central (Lantian), and eastern (Yan'an) parts of the Chinese Loess Plateau as a case study to explore the feasibility of a cross-regional intelligent recognition method for loess water content. Initially, we simulated the environmental conditions encountered during in-hole recognition to design an image collection platform, and subsequently prepared a loess water content dataset comprising 32,940 images from these three regions. Based on domain adaptation, we proposed a deep learning recognition method, namely, self-attention-based domain adaptation with Deep EXpectation (DADEX Swin). The proposed method demonstrates reasonable accuracy in estimating the water content of loess across regions, with a mean absolute error of 0.807%-1.137%, mean absolute percentage error of 0.074-0.154, and root mean square error of 0.939%-1.546%. The findings of this study provide valuable insights for addressing cross-regional issues, and DA-DEX Swin shows promise for application in robot detection technology to monitor variations and distributions of water content within slopes.
Soil water content (SWC) plays a vital role in agricultural management, geotechnical engineering, hydrological modeling, and climate research. Image-based SWC recognition methods show great potential compared to traditional methods. However, their accuracy and efficiency limitations hinder wide application due to their status as a nascent approach. To address this, we design the LG-SWC-R3 model based on an attention mechanism to leverage its powerful learning capabilities. To enhance efficiency, we propose a simple yet effective encoder–decoder architecture (PVP-Transformer-ED) designed on the principle of eliminating redundant spatial information from images. This architecture involves masking a high proportion of soil images and predicting the original image from the unmasked area to aid the PVP-Transformer-ED in understanding the spatial information correlation of the soil image. Subsequently, we fine-tune the SWC recognition model on the pre-trained encoder of the PVP-Transformer-ED. Extensive experimental results demonstrate the excellent performance of our designed model (R2 = 0.950, RMSE = 1.351%, MAPE = 0.081, MAE = 1.369%), surpassing traditional models. Although this method involves processing only a small fraction of original image pixels (approximately 25%), which may impact model performance, it significantly reduces training time while maintaining model error within an acceptable range. Our study provides valuable references and insights for the popularization and application of image-based SWC recognition methods.
It is a challenging problem to extract aero-engine fault signals which contaminated by non-Gaussian noises under complex operation conditions. A Learnable Wavelet Packet De-noising Network (LWPD-Net) is proposed in this paper to address it. The highlights of LWPD-Net are to performs multi-scale decomposition of the original signal to extract features across various frequency bands, and more importantly the Double-Sharp threshold function with learnable parameter is employed to suppresses the non-Gaussian noise. Moreover, supervised learning strategy is adopted to learn LWPD-Net parameters for adaptively enhancing the fault features through conducting training samples from fault signals corrupted by non-Gaussian noises. Simulation experiments show that LWPD-Net can recover fault feature frequencies in the envelope spectrum for varying noise levels, and achieves satisfying fault feature extraction performances. Additionally, experiments conducted on aero-engine gear hubs confirm that the proposed method can recover the spectral features of vibration signals interfered with non-Gaussian noise under different operation conditions.
Direction-of-arrival (DOA) estimation methods are divided into model-driven and data-driven algorithms according to the driving method. However, traditional model-driven algorithms suffer from poor accuracy at low signal-to-noise ratios (SNR), while data-driven models fail to consider the significance between data, leading to a high number of algorithm parameters, inaccurate DOA estimation. To address these issues, we propose the Denoised Attention Neural Network (DANN), which utilizes the channel attention module, spatial attention module, and multi-head attention module to assign weights to features and learn their importance adaptively. In addition, to reduce the interference of noise on the estimation results, a threshold denoising module is introduced to achieve noise abatement. Simulation results demonstrate that our method achieves higher estimation accuracy at low SNR with smaller network parameters compared to existing methods.
The deconvolution method (DM) is an effective tool for enhancing the impulsive features of rolling bearings. Deep network-based deconvolution methods transform complex numerical computations into network optimization, improving the performance of DMs. However, current deep deconvolution methods do not consider the possibility of utilizing nonlinear operations to further enhance the performance. Inspired by convolutional sparse learning and algorithm unfolding, a bearing fault feature extraction method, Multi D-norm driven algorithm unfolding network (MDN-AUN) is proposed. First, a convolutional sparse coding model is established with the fault impulse as the target sparse code. Then the Iterative Soft Thresholding Algorithm (ISTA) for solving the model is unfolded from the iterative direction and the algorithm unfolding network is generated. Finally, the multi D-norm, which can evaluate the periodicity and impulsiveness of the signal, is chosen as the objective function of the sparse code to train the neural network. MDN-AUN is compared with several deconvolution methods through simulation and real vibration signals. The results show that MDN-AUN has the superior ability to enhance the impulsive features, and the fault diagnosis performance is significantly better than the compared methods.
To address the problem that the traditional fault diagnosis method for chemical processes under big data relies too much on expert experience and fault features are difficult to distinguish, a deep learning-based fault diagnosis method is proposed, which combines convolutional neural network (CNN), long and short-term memory (LSTM) and attention mechanism (AM). In this method, the spatial sequence features of the input signal are extracted by the CNN adaptively, while the LSTM extracts the time-series features of the signal. Finally, the model performance is enhanced by introducing the attention mechanism and using the SoftMax layer as a classifier for fault diagnosis, so that the model can notice the important features of the faults with the interference of noise. Simulation validation of the method in this paper is performed using the TE chemical process data set, and it is demonstrated that the method can be used for chemical process fault diagnosis studies. Finally, compared with other fault diagnosis methods, the method is more accurate and has certain superiority.
Long-term in-situ monitoring of loess moisture content is important for revealing the mechanism and preventing the loess disaster, but the current detection methods have limitations in long-term, non-destructive, deep and large-scale. While the intelligent detection can provide new direction for this problem. The pixel-level differences of loess in different regions are different, which leads to the current intelligent detection method can not predict loess moisture content in unknown regions from known regions. Therefore, we propose a cross-regional loess water content prediction method. Firstly, a loess dataset with three Loess Plateau regions is built and the domain adaptation is introduced to minimize the pixel-level differences, however, the dispersion distribution of classes makes the traditional DA methods unable to solve the regression problem of the loess water content prediction. Therefore, we carry out deep classification with Deep EXpectation on the class space. The feature distribution of loess images is global, which requires the model to extract the features powerfully, while the convolution neural network has poor ability to extract the semantic relationship between global features. Therefore, we build a feature extractor based on self-attention to extract the global features. Finally, the effectiveness of our method is verified by experiments.
Aeroengine has a complex mechanical structure, and its working condition varies widely. Its fault signals are thus modulated through complex nonlinear transfer paths and influenced by non-Gaussian noises. The popular data-driven multiscale diagnostic model, however, does not sufficiently consider the embedded noises and consequently keeps the noise content in the advanced discriminative features, which decreases the diagnostic accuracy. Therefore, a multiscale attention network with adaptive noise reduction (MANANR) is proposed in this article for aeroengine bearing fault diagnosis. The MANANR first divides the original vibration signal into different scales via the average method of adjacent points. Then, two stacked multiscale noise reduction modules, B1 and B2, are designed for noise reduction. The core strategy behind B1 and B2 is the threshold noise reduction (TNR), which removes the noises from multiscale convolution features adaptively. Based on the physical principles that global average pooling (GAP) and max pooling (MAP) can extract periodic and impulsive characteristics of the fault signals respectively, the threshold value of TRN is thus constructed through fusing GAP and MAP outputs. Furthermore, an attention mechanism is employed to enhance the discriminative capability of multiscale features by globally capturing the relations among different scales and channels. Finally, one two-layer classifier is introduced to confirm bearing fault patterns. Experimental results demonstrate that the proposed method has the feature-level noise reduction property, and more importantly, achieves satisfying intelligent diagnosis precision for aeroengine bearing with the minimum peeling area of 0.5 mm2 under the continuous acceleration condition from 12 000 to 12 550 r/min, which outperforms state-of-the-art intelligent diagnostic methods.
Accurate Remaining Useful Life (RUL) prediction of turbofan engines not only reduces labor and maintenance costs but also mitigates the risk of aircraft accidents. To solve the noise interference problem, we present a Legendre memory model enabled Attention Network (LAN). The LAN utilizes Legendre memory model and dual multi-head self-attention. Legendre memory model contains two sub-layers: Legendre Projection Unit (LPU) and Frequency Selection Layer (FSL). The sensor data is projected into the state space using the LPU to retain the historical data information. Next, frequency components are selected using the FSL to reduce the impact of noise. Meanwhile, the dual multi-head self-attention is applied to capture useful feature information from both the time step and sensor dimensions for RUL prediction. Lastly, the extracted features are inputted into the residual convolutional network for RUL prediction. Experiments are performed on the widely adopted C-MAPSS dataset, and the results substantiate the efficacy of the proposed approach.
The in-situ detecting robot equipped with vision system can be used for deep in-situ detecting of loess water content, which provides an effective means to reveal loess disaster mechanism and loess disaster early warning. Current image recognition methods mainly focus on the relationship between a type of soil image and its water content, but rarely study the cross-layer method and its feasibility. However, the loess-layer and paleosoil-layer exist alternately in the loess region, which makes the research of cross-layer moisture content recognition(CMCR) method necessary. Therefore, we take loess images to recognize the water content of paleosol images to explore the feasibility of CMCR and propose a new method. Firstly, a dataset of 19 categories of loess and paleosol images is built for model training and validation. Then, the high redundancy of image information is illustrated by an encoder-decoder, and a model is proposed with the idea of ignoring redundant information to CMCR. Finally, fusion model is designed to combine representative areas in the model. Experimental results show that: mean absolute error is 1.149, mean absolute percentage error is 0.082 and root mean square error is 1.761, which shows the feasibility of CMCR and provides algorithmic support for the visual system.
It is a challenge problem to accurately recognize damage distribution pattern for multi-stage industrial gearboxes in filed, due to entangled relationships between strong interferences/noises and complicate transfer path modulations. In this work, a tailored two-stage strategy (LR-CSL) based on low-rank representation and convolutional sparse learning is proposed. Based on the periodic similarity of focused features, a weighted low-rank stage is firstly utilized to suppress strong interferences and noises, which provides a cornerstone to enhance blind deconvolution methods. Then, a convolutional sparse stage is adopted to mitigate the transfer path modulation by enforcing one nonnegative bounded regularizer, which guarantees the reliable recovery of impulsive source envelopes. Lastly, the damage distribution patterns could be reliably confirmed by directly referring to the recovered source envelopes (rather than modulated waveforms) and gearbox dynamics. Comprehensive health evaluations to one 750 kW wind turbine drivetrain are performed blindly and gear surfaces with multiple weak spalling patterns are recognized accurately. Moreover, the spalling fault evolution process is deduced and maintenance guidances are allocated. Further analysis also confirms the first low-rank stage plays a necessary and important role in boosting LR-CSL's deconvolution capability. Lastly, quantitative evaluations demonstrate that our LR-CSL method achieves a higher diagnostic accuracy than state-of-the-art fault diagnosis techniques. (c) 2020 Elsevier Ltd. All rights reserved.