Abstract Accurate and real-time coal-gangue identification is quintessential for advancing the automation of top-coal caving. To tackle the persistent challenges of intense underground background noise and the non-stationary nature of impact signals, this paper proposes a Cross-Modal Spatiotemporal Gated Fusion (CMSGF) method. Specifically, a MoE-based Spectrum-Aware Dynamic Vibration Encoder is developed to accommodate the transient impact characteristics of vibration signals, achieving sample-level adaptive kernel aggregation via VMD-decomposed multi-channel inputs. Concurrently, to resolve the physical property coupling in acoustic features, a Time-Frequency Decoupled Three-Stream Audio Pyramid Network is designed to reconstruct coal-gangue physical attributes across three orthogonal dimensions, significantly enhancing fine-grained discrimination. To bridge the semantic gap between heterogeneous modalities, we further introduce a Bidirectional Symmetric Multi-head Cross-modal Attention (BSCCA) module and a dual-dimensional spatiotemporal gating mechanism, which adaptively filter high-SNR features through a dynamic competition strategy. Experimental results on a simulated top-coal caving platform demonstrate that the proposed method achieves a recognition accuracy of 96.94%. This research provides a low-cost, high-reliability for online coal-gangue monitoring.
Mechanical fault diagnosis under industrial noise interference poses significant challenges, including low feature fusion efficiency and inadequate dynamic noise adaptation. To address these issues, this paper proposes a noiserobust fault diagnosis method that leverages multi-scale feature enhancement and dynamic cross-modal interaction, thereby enhancing noise robustness through hierarchical feature enhancement and an adaptive interaction mechanism. First, using the time and frequency domain features of the original time series data after noise reduction, a Multi-scale time-frequency noise separation and feature enhancement methods (MSTF) framework is constructed to achieve the synergistic optimization of noise suppression and fault feature enhancement. Secondly, a dual-branch architecture is designed: the one-dimensional vibration signal branch uses a multi-scale deep residual network (MS-ResNet) to extract time-domain multi-granularity features, and the time-frequency image branch parses the time-frequency features generated by MSTF through a multi-scale attention interleaving network (MAIN) to realize multi-scale feature extraction. The dynamic cross-modal interactive fusion (DCIF) module is proposed to dynamically adjust the modal weights through a noise-aware gated attention mechanism, addressing the adaptive defects of traditional fixed fusion in response to changes in noise intensity. The multimodal joint optimization strategy is further proposed to construct a multilevel constraint mechanism. Experimental results on the CWRU dataset (SNR = -6 dB) show that the Accuracy of DMRF-Net reaches 97.51 %, which is significantly better than that of the traditional method. The ablation experiments further demonstrate the superiority of the proposed framework. The proposed framework in this study will provide an effective solution for fault diagnosis under complex operating conditions.
Coal-gangue recognition technology plays an important role in the intelligent realization of integrated working faces and coal quality improvement. However, the existing methods are easily affected by high dust, noise, and other disturbances, resulting in unstable recognition results that make it difficult to meet the needs of industrial applications. To realize accurate recognition of coal-gangue in noisy environments, this paper proposes an end-to-end multi-scale feature fusion convolutional neural network (MCNN-BILSTM) based gangue recognition method, which can automatically learn and fuse complementary information from multiple signal components of vibration signals. It combines traditional filtering methods and the idea of multi-scale learning, which can expand the breadth and depth of the feature learning process. the breadth and depth of the feature learning process. Moreover, to strengthen the expression of key features, a feature weighting method based on the attention mechanism is combined to give adaptive weights to different features. Finally, the experimental platform of a tail beam of coal-gangue impact hydraulic support is built, and several comparative experiments are carried out. The comprehensive comparison experiments show that the method shows strong adaptability, robustness, and noise resistance under various complex noise environments, and is suitable for complex practical industrial sites.
To address the issues of severe interference of equipment operating noise and information loss caused by single extraction methods during coal gangue audio feature extraction, a coal gangue audio classification method based on improved EfficientNet is proposed. The method adopted a feature extraction approach combining Mel spectrogram and Gammatone frequency cepstral coefficients to effectively capture low-frequency information and detailed features in gangue audio. EfficientNet-B0 was selected as the backbone network, and the following improvements were made: the original multi-scale channel attention module was replaced with a convolutional block attention module, resulting in the Convolutional Attention Feature Fusion (CAFF) module. This module allowed the network to autonomously assign different weight information to features in different spatial positions, generating new effective features. Additionally, a Frequency-domain Channel Attention (FCA) module was embedded in parallel within the original MBConv module, strengthening the representation ability of feature maps and thereby improving overall network performance. The experimental results demonstrated that after introducing the CAFF module, the model's accuracy improved by 0.61%, the F1 score increased by 0.52%, and convergence was faster, indicating that the CAFF module effectively enhanced the model's ability to capture spectral features. After integrating the FCA module, accuracy improved by 0.45%, and the F1 score increased by 0.62%, showing that combining these modules further enhanced the model's generalization ability and its ability to process complex features. The improved EfficientNet model achieved an accuracy of 91.90%, with a standard deviation of 0.108, significantly outperforming other comparable audio classification models.
Fast and accurate coal-gangue identification techniques are essential for intelligent integrated mining and improved coal quality. However, existing methods are susceptible to high dust, noise, and other disturbances, resulting in unstable recognition results that cannot meet the demands of industrial applications. To address these challenges, this paper proposes a coal-gangue recognition method based on the fusion of a multi-scale parallel MCNN-BITCN network with an attention mechanism and the Improved Sparrow Search Algorithm (ISSA). The method combines a bidirectional spatio-temporal convolutional network (BITCN) with a multi-branch convolutional neural network (MCNN) to deeply mine and expand the extracted features. The time-frequency features are then fused through a cross-attention mechanism and fed into a fully connected layer. An improved sparrow search algorithm (ISSA) is used to generate the input weights and biases of the hidden layer nodes. This paper constructs an experimental platform to simulate the impact of coal-gangue on the tail beam of hydraulic support and conducts several comparative experiments. Results show that the MCNN-BITCN model maintains 83.76
Infrared and visible image fusion seeks to integrate complementary information from both modalities, generating a single, comprehensive image for subsequent visual tasks. However, existing fusion methods often inadequately capture local details and long-range contextual information, resulting in information loss and edge blurring. In this paper, we propose a Dual-branch Multi-cascade Hierarchical Network (DMHNet) for infrared and visible image fusion. DMHNet utilizes a dual-branch architecture to extract both local and global features. Furthermore, we introduce a feature fusion module that employs feature embedding and a bi-directional calibrated attention mechanism to facilitate global information exchange and achieve efficient fusion of complementary features. Additionally, DMHNet employs progressive feature linking for cross-stage feature fusion. This strategy preserves low-level details and enhances high-level semantic information, while simultaneously minimizing gradient loss from the source images. Extensive experiments on benchmark datasets demonstrate that DMHNet surpasses stateof-the-art methods across both qualitative and quantitative metrics, confirming its efficacy for multimodal image fusion.
Accurate identification of coal and gangue is a crucial guarantee for efficient and safe mining of top coal caving face. This article proposes a coal-gangue recognition method based on an improved beluga whale optimization algorithm (IBWO), convolutional neural network, and long short-term memory network (CNN-LSTM) multi-modal fusion model. First, the mutation and memory library mechanisms are introduced into the beluga whale optimization to explore the solution space fully, prevent falling into local optimum, and accelerate the convergence process. Subsequently, the image mapping of the audio signal and vibration signal is performed to extract Mel-Frequency Cepstral Coefficients (MFCC) features, generating rich sample data for CNN-LSTM. Then the multi-head attention mechanism is introduced into CNN-LSTM to speed up the training speed and improve the classification accuracy. Finally, the IBWO-CNN-LSTM coal-gangue recognition model is constructed by the optimal hyperparameter combination obtained by IBWO to realize the automatic recognition of coal-gangue. The benchmark function proves that IBWO is superior to other optimization algorithms. By building an experimental platform for the impact of coal and gangue falling on the tail beam of hydraulic support, multiple experimental data collection is carried out. The experimental results show that the proposed coal-gangue recognition model has better performance than other recognition models, and the accuracy rate reaches 95.238%. The multi-modal fusion strategy helps to improve the accuracy and robustness of coal-gangue recognition.
The coal-gangue recognition technology plays an important role in the intelligent realization of fully mechanized caving face and the improvement of coal quality. Although great progress has been made for the coal-gangue recognition in recent years, most of them have not taken into account the impact of the complex environment of top coal caving on recognition performance. Herein, a hybrid multi-branch convolutional neural network (HMBCNN) is proposed for coal-gangue recognition, which based on improved Mel Frequency Cepstral Coefficient (MFCC) as well as Mel spectrogram, and attention mechanism. Firstly, the MFCC and its smooth feature matrix are input into each branch of one-dimensional multi-branch convolutional neural network, and the spliced features are extracted adaptively through multi-head attention mechanism. Secondly, the Mel spectrogram and its first-order derivative are input into each branch of the two-dimensional multi-branch convolutional neural network respectively, and the effective time-frequency information is paid attention to through the soft attention mechanism. Finally, at the decision-making level, the two networks are fused to establish a model for feature fusion and classification, obtaining optimal fusion strategies for different features and networks. A database of sound pressure signals under different signal-to-noise ratios and equipment operations is constructed based on a large amount of data collected in the laboratory and on-site. Comparative experiments and discussions are conducted on this database with advanced algorithms and different neural network structures. The results show that the proposed method achieves higher recognition accuracy and better robustness in noisy environments.
Stability of obstacle–crossing and structural optimization are important issues in the research of tracked mobile robots. In this paper, in order to fully understand the obstacle–surmounting ability of the robot, the relationship between the position of the center of gravity and the posture of the front and rear swing arms is analyzed. Based on the motion mechanism of the robot crossing obstacles, the geometric model and the dynamic model are established for the key states in the obstacle crossing process. Based on these models, a multi-objective optimization problem for the maximum obstacle–crossing height and minimum driving torque is established during the obstacle crossing process of the robot, which must meet geometric, slip, and stability constraints. To effectively handle the optimization problem of tracked mobile robots, an improved non–dominated sorting genetic algorithm with elite strategy version II based on adaptive genetic strategy (NSGA-II-AGS) is proposed in this paper. Some meaningful relationships between the objective function and the design variables are obtained through sensitivity analysis. Finally, the robot's obstacle-crossing ability was verified through virtual simulation and prototype experiments. These excellent performances enable the proposed NSGA-II-AGS to be qualified for dealing with the multi-objective optimization problem.
An event camera is a neuromimetic sensor inspired by the human retinal imaging principle, which has the advantages of high dynamic range, high temporal resolution, and low power consumption. Due to the interference of hardware and software and other factors, the event stream output from the event camera usually contains a large amount of noise, and traditional denoising algorithms cannot be applied to the event stream. To better deal with different kinds of noise and enhance the robustness of the denoising algorithm, based on the spatio-temporal distribution characteristics of effective events and noise, an event stream noise reduction and visualization algorithm is proposed. The event stream enters fine filtering after filtering the BA noise based on spatio-temporal density. The fine filtering performs time sequence analysis on the event pixels and the neighboring pixels to filter out hot noise. The proposed visualization algorithm adaptively overlaps the events of the previous frame according to the event density difference to obtain clear and coherent event frames. We conducted denoising and visualization experiments on real scenes and public datasets, respectively, and the experiments show that our algorithm is effective in filtering noise and obtaining clear and coherent event frames under different event stream densities and noise backgrounds.
The mechanical fault diagnosis of HVCB is important to ensure the stability of electric power systems. Aiming at the problem of poor diagnostic performance of deep learning methods under limited samples, this paper proposes an HVCB operating mechanism fault diagnosis model (multi-channel CNN-SABO-SVM, MCCSS) based on multimodal data fusion features and Subtraction-Average-Based Optimizer (SABO). This model extracts and fuses features from the input two-dimensional data using a multi-channel CNN network and then uses the multimodal data fusion features to diagnose HVCB faults. Additionally, the SVM is used instead of the Softmax classifier to classify the fused features of vibration and sound, compensating for the poor diagnostic performance and generalization ability of the CNN network in small sample data scenarios. To further enhance the fault diagnosis performance of the SVM, the SABO is introduced for hyperparameter optimization of the SVM classifier. An HVCB fault test platform was established to train and test the model with limited data. The experimental results show that, compared with the multi-channel CNN-SVM and the CNN model based on unimodal signals, the proposed multi-channel CNN-SABO-SVM model improves the accuracy by 2.66% and 10.66%, respectively, and effectively addresses the challenge of circuit breaker fault diagnosis with limited samples.
Traditional coal-gangue recognition methods usually do not consider the impact of equipment noise, which severely limits its adaptability and recognition accuracy. This paper mainly studies the more accurate recognition of coal-gangue in the noise site environment with the operation of shearer, conveyor, transfer machine and other device in the process of top coal caving. Mel Frequency Cepstrum Coefficients (MFCC) smoothing method was introduced to express the intrinsic feature of sound pressure more clearly in the coal-gangue recognition site. Then, a multi-branch convolution neural network (MBCNN) model with three branches was developed, and the smoothed MFCC feature was incorporated into this model to realize the recognition of falling coal and gangue in noisy environment. The sound pressure signal datasets under the operation of different device were constructed through a great deal of laboratory and site data acquisition. Comparative experiments were carried out on noiseless dataset, single noise dataset and simulated site dataset, and the results show that our method can provide higher correct recognition accuracy and better robustness. The proposed coal-gangue recognition approach based on MBCNN and MFCC smoothing can not only recognize the state of falling coal or gangue, but also recognize the operational state of site device.
To solve the problem of inaccurate object segmentation caused by unbalanced samples for in-vehicle point cloud, an improved semantic segmentation network RangeNet++ based on asymmetric loss function (AsL-RangeNet++) is proposed, which uses asymmetric loss (AsL) function and Adam optimizer to calculate and adjust object weights, achieve optimal point cloud segmentation. AsL-RangeNet++ can solve the problem of unbalance between positive and negative samples and label error in multi-label classification by calculating the weights of positive and negative samples respectively and more accurately segments the point cloud of small targets. A large number of experiments on the widely used SemanticKITTI dataset show that the proposed method has higher segmentation accuracy and better adaptability than the current mainstream methods.
The article proposes a qualitative identification scheme of fluorescent immunoassay strips based on residual networks to address problems such as poor strip positioning accuracy and inadequate strip size specifications in current photoelectric fluorescent immunoassay quantitative detection systems, which result in low accuracy and detection efficiency. The proposed method employs the Hough line detection algorithm, which is based on Canny edge detection, to extract the tilt angle of the test strip. It then combines this with contour extraction of the strip image to calculate its contour center. By utilizing the test strip tilt angle and contour center, the article accurately locates the test strip position. The residual network is utilized for extracting strip features, while the extreme learning machine is employed for discriminating the validity and positive/negative results of the fluorescent strips. Validation experiments on a novel coronavirus fluorescent immunoassay strip demonstrate that the residual network model based on extreme learning machine proposed in this article achieves 100% accuracy, precision, recall, and F1-score values for strip classification, effectively improving the recognition accuracy and detection efficiency of the fluorescent immunoassay quantitative detection system.
The surface electromyography(sEMG) is one of the basic processing techniques to the gesture recognition because of its inherent advantages of easy collection and non-invasion. However,limited by feature extraction and classifier selection, the adaptability and accuracy of the conventional machine learning still need to promote with the increase of the input dimension and the number of output classifications. Moreover, due to the different characteristics of sEMG data and image data, the conventional convolutional neural network(CNN) have yet to fit sEMG signals. In this paper, a novel hybrid model combining CNN with the graph convolutional network(GCN) was constructed to improve the performance of the gesture recognition. Based on the characteristics of sEMG signal, GCN was introduced into the model through a joint voting network to extract the muscle synergy feature of the sEMG signal. Such strategy optimizes the structure and convolution kernel parameters of the residual network(ResNet) with the classification accuracy on the NinaPro DBl up to 90.07%. The experimental results and comparisons confirm the superiority of the proposed hybrid model for gesture recognition from the sEMG signals.
Gastric cancer is the third most common cause of cancer-related death in the world. Human epidermal growth factor receptor 2 (HER2) positive is an important subtype of gastric cancer, which can provide significant diagnostic information for gastric cancer pathologists. However, pathologists usually use a semi-quantitative assessment method to assign HER2 scores for gastric cancer by repeatedly comparing hematoxylin and eosin (H&E) whole slide images (WSIs) with their HER2 immunohistochemical WSIs one by one under the microscope. It is a repetitive, tedious, and highly subjective process. Additionally, WSIs have billions of pixels in an image, which poses computational challenges to Computer-Aided Diagnosis (CAD) systems. This study proposed a deep learning algorithm for HER2 quantification evaluation of gastric cancer. Different from other studies that use convolutional neural networks for extracting feature maps or pre-processing on WSIs, we proposed a novel automatic HER2 scoring framework in this study. In order to accelerate the computational process, we proposed to use the re-parameterization scheme to separate the training model from the deployment model, which significantly speedup the inference process. To the best of our knowledge, this is the first study to provide a deep learning quantification algorithm for HER2 scoring of gastric cancer to assist the pathologist's diagnosis. Experiment results have demonstrated the effectiveness of our proposed method with an accuracy of 0.94 for the HER2 scoring prediction.
关节力矩预测在康复医学、临床医学和运动训练等领域有着重要作用,对力矩连续、实时地预测可以使人机交互设备更好地反馈、复刻人体运动意图.为了给患者提供一个安全、主动、舒适的康复训练环境,提升人机交互设备的柔顺性,提出了一种改进型递归小脑模型神经网络模型关节力矩预测方法.该方法采用肌肉协同分析对采集的相关肌肉的表面肌电信号(sEMG)进行降维,将降维后的sEMG特征向量与关节角速度、关节角度作为输入信号,并在小脑模型神经网络中加入递归单元和模糊逻辑规则,以小波函数作为隶属度函数,对非疲劳、过渡疲劳及疲劳这3种状态下的踝关节背屈跖屈运动的动态力矩进行连续预测.力矩预测值与实际值之间的平均皮尔逊相关系数和平均标准均方根误差分别为0.933 5和0.159 8,实验结果验证了该方法对下肢关节力矩连续预测的准确性和有效性.
Abstract. A reliable optimization of dynamic vibration absorber (DVA) parameters is extremely important to analyze its dynamic damping characteristics and improve its vibration suppression performance. In this paper, we will discuss a parameter optimization method of the Voigt and three-element DVA models according to the H∞ optimization criterion. The particle swarm optimization method is an effective heuristic optimization algorithm; however, it is easy to lose diversity and fall into local extremum. To solve this problem, the adaptive multiswarm particle swarm optimization (AM-PSO) is used to search the solution of the DVA models. Particles in AM-PSO are adaptively divided into multiple swarms, and the variable substitution learning strategy is utilized to reduce their computational complexity and improve the algorithm's global search capability. In addition, the AM-PSO method is employed to optimize the parameters of DVA models and compared with the genetic algorithm and PSO. The simulation results show that the AM-PSO algorithm has superior performance. Also, the adaptive multiswarm numerical design method discussed herein will push the field towards practical applications, including traditional DVA and related complex three-element DVA.
The lifting pipe is a key component of deep sea mining whose dynamic response directly affects the safety of the lifting operation. The objective of this paper was to investigate the effects of heave motion and sailing velocity of mining vessel and the buffer mass on the dynamic response of lifting pipe. First, an equivalent model of the lifting pipe was established, and the natural frequency and dynamic response of the lifting pipe equivalent model were determined with consideration of the wave action by the method of separated variables. Secondly, the reliability of the equivalent model was verified by simulating a 5000 m stepped pipe with OrcaFlex software. Then the dynamic displacement, axial tension, axial stress of the lifting pipe under different sea conditions and sailing velocities were studied, and the main factors affecting the dynamic response of the pipe described. By comparing the simulation results of actual and equivalent models, the equivalent model can be used to analyze the longitudinal vibration characteristics of the lifting pipe. The sailing velocity of the mining vessel has little effect on the dynamic response of the lifting pipe, but the surface wave has a significant effect.
为了研究复杂阶梯状扬矿管在采矿船升沉运动和海流作用下的纵向振动特性,利用连续弹性杆振动理论,对5 000 m长扬矿管纵向振动性能进行分析.首先,根据达朗贝尔原理建立扬矿管纵向振动数学模型,采用分离变量法推导管道固有频率方程;然后,进行振型的质量归一化处理;最后,利用ABAQUS软件建立扬矿管有限元模型,对管道的纵向动态响应进行研究.研究结果表明:扬矿管的一阶纵向共振频率处于矿区海浪能量集中的频带内,随着中间矿仓质量的增加扬矿管固有频率减小,中间矿仓质量对高阶固有频率的影响更加明显;随着海浪频率的增加,纵向振幅、轴向力和轴向应力先增大后减小,并在一阶固有频率时达到峰值,其峰值分别发生在扬矿管5 000、0、1 000 m处;随着采矿船升沉幅值的增加,扬矿管的动态响应逐渐增大,当升沉幅值大于1.5 m时,扬矿管动态响应的增长速度变缓;扬矿管发生一阶纵向共振时,振动位移和轴向力先增大后作等幅稳态振荡;随着海水深度的增加,沿管长方向的振动幅值逐渐增大,振动平衡位置发生下移,振动响应时间发生延迟,同时轴向力和轴向应力逐渐减小,且轴向应力在每两级阶梯管间急剧变大.