Pulmonary embolism (PE) detection in Computed Tomography Pulmonary Angiography (CTPA) remains challenging due to the presence of small embolic lesions, low contrast, and complex three-dimensional vascular structures. To address these challenges, we propose IVT-PENet, an Information-Gated and Variance-Enhanced U-Net with Direction-Aware Positional Encoding and a Channel-Spatial Transformer. First, a Direction-Aware Positional Encoding (DAPE) module is embedded throughout the encoder–decoder pathway to preserve directional anatomical priors and maintain spatial consistency. Second, an Information-Gated and Variance-Enhanced Atrous Spatial Pyramid Pooling (IG-VEASPP) module is developed to enhance multi-scale contextual representation and improve the discrimination of subtle embolic candidates. Third, a Channel-Spatial Attention Transformer (CSAT) is introduced at the bottleneck to capture long-range three-dimensional contextual dependencies through sequential channel–spatial attention. In addition, a dynamic probability-annealing strategy is employed to adaptively regulate random noise and intensity perturbations during training, thereby improving model robustness. The proposed framework adopts a detection-oriented design, where vessel-aware anatomical guidance and embolic candidate generation jointly facilitate the subsequent detection process. Experimental results on the CAD-PE dataset demonstrate that IVT-PENet achieves a sensitivity of 0.895, a FROC score of 0.812, and an AUC of 0.851. Furthermore, without any retraining or fine-tuning, IVT-PENet attains a sensitivity of 0.942, a FROC score of 0.764, and an AUC of 0.881 on the external FUMPE dataset, demonstrating strong cross-dataset generalization capability. These results indicate that IVT-PENet effectively captures multi-scale embolic characteristics and provides reliable pulmonary embolism detection in complex vascular environments. The source code is available at https://github.com/GuYuIMUST/IVT-PENet.
The burst-like and high-amplitude characteristics of impulsive noise, which markedly differ from those of Gaussian noise, render methods based on the Gaussian assumption unable to accurately characterize signals under impulsive noise. Moreover, when dealing with multicomponent signal, existing impulsive noise suppression methods inevitably introduce cross-term interference. To address these issues, this paper proposes an impulsive noise suppression method based on the torque clustering (TC) algorithm, and thus establishes an accurate representation of multicomponent linear frequency modulation (LFM) signal under impulsive noise. First, the theoretical analysis is conducted to elucidate the inherent limitations of existing noise suppression methods that inevitably cross-term introduction. A novel impulsive noise suppression technique based on TC is developed, fundamentally eliminating cross-term interference. Subsequently, two signal representation methods for multicomponent LFM signal under impulsive noise are proposed, namely, TC-fractional Fourier transform (TC-FRFT) and TC-synchrosqueezing transform (TC-SST). These methods enable accurate characterization of multicomponent LFM signal under impulsive noise, and facilitate precise extraction of signal features. Finally, a mathematical model for parameter estimation of multicomponent LFM signal is established using TC-FRFT, enabling high-precision estimation of center frequency and chirp rate in the presence of impulsive noise. Simulation results show that the proposed method can effectively suppress impulsive noise, avoid cross-term interference in existing methods, and achieve accurate characterization of multicomponent LFM signal under impulsive noise. Furthermore, the proposed TC-FRFT outperforms existing parameter estimation methods in terms of stability, accuracy, and robustness against noise.
As multimedia technology continues to advance, the quality of multichannel audio data has become particularly important for enhancing the user experience. However, audio signals are susceptible to noise interference and data loss during transmission and processing, resulting in degradation of sound quality. This study focuses on the recovery of multichannel audio with deletions using tensor correlated total variation (t-CTV) regularization, which aims to recover the complete audio tensor from partially observed data. By integrating the t-CTV regularization term into the tensor singular value decomposition (t-SVD) framework, a regularization model incorporating low-rankness and local smoothness is constructed and optimized using the alternating direction multiplier method (ADMM). Experimental results on two publicly available datasets, VoiceHome-2 and Alimeeting, show that the method performs well in a variety of data loss scenarios, especially when the audio signal missing rate is up to 75%, it still can significantly recover the audio quality with a PESQ value of 3.666, which is much higher than other algorithms. This study demonstrates the applicability of the t-CTV regularization method for audio restoration tasks, while exploring potential extensions of tensor completion techniques in acoustic applications.
Pneumonia classification from chest x-ray images remains a challenging task because different pneumonia categories often exhibit highly similar visual manifestations, while subtle lesion regions require both fine-grained local detail perception and effective global contextual modeling. To address these challenges, we propose PneumoMamba, a dual-path network for computer-aided pneumonia diagnosis that combines a convolutional neural network branch for local texture extraction with a Mamba-based branch for efficient long-range dependency modeling. In the proposed framework, the convolutional branch focuses on capturing detailed local structures, whereas the Mamba-based branch is designed to model global contextual information with linear computational complexity. To better preserve spatial continuity for sequence modeling, we further introduce an eight-directional scan strategy (8DScan), which converts two-dimensional feature maps into multidirectional scan sequences for more comprehensive spatial dependency modeling. In addition, a Multi-Scale Asymmetric Convolution (MSAConv) and a focused feature module (FFM) are incorporated to enhance the representation of subtle pneumonia-related patterns. We evaluate the proposed method on the Pneumonia-CXR-Database, which contains four classes of chest x-ray images, under a 6:2:2 train/validation/test split. Experimental results show that PneumoMamba achieves a test accuracy of 94.94%, together with strong F1-score, sensitivity, and precision, outperforming representative comparison methods. These results demonstrate the effectiveness of the proposed framework for computer-aided pneumonia classification.
Combining deep learning and bird sound recognition strongly supports monitoring bird species and maintaining ecological balance. However, in outdoor environments, the extraction of bird sound features is often hindered by environmental noise, making it challenging for models to learn the fine-grained features of bird sounds fully. And single-scale feature extraction is harrowing to cover the time-frequency domain feature information of bird sounds in multiple dimensions. To address these issues, this paper proposes a multi-grained detail-enhanced and patch-aware network. The model utilizes densely connected time delay neural network as the backbone network and introduces the multi-grained detail-enhanced convolution, which combines vanilla convolutions with differential convolutions in the horizontal, vertical, angular, and central levels, and incorporates multi-grained pooling strategies to learn fine-grained acoustic features at different levels. To further overcome the limitations of single-scale feature extraction, the branch patch-aware attention module is proposed. This module collaboratively captures local details and global contextual information through a multi-branch structure and patch partitioning of different sizes. On the three datasets, the method achieved accuracies of 96.29%, 86.51%, and 97.40%, respectively. This achievement demonstrates the precise capture and parsing ability of the method for audio feature information.
Due to the memory and non-local characteristics of fractional calculus, fractional-order tracking differentiator (FOTD) performs excellently in suppressing impulse noise. However, the parameters of FOTD need to be manually adjusted according to the scene requirements, and cannot automatically maintain optimal performance in scenarios where the signal and noise intensities change dynamically. To address this issue, this paper proposes a multi-parameter optimization-driven FOTD (OFOTD) based on envelope entropy, enhancing the adaptability of FOTD in complex scenarios. Furthermore, a fractional multisynchrosqueezing transform (FRMSST) is developed, and OFOTD-FRMSST is established to accurately represent the signal under impulsive noise. Finally, OFOTD-FRMSST is applied to parameter estimation of linear frequency modulation (LFM) signal, demonstrating its superiority in accuracy, noise robustness, and practicality. Experimental results demonstrate that, from both time domain and time-frequency plane, OFOTD achieves enhanced noise suppression performance through adaptive parameter optimization. Furthermore, in comparison with existing methods, OFOTD-FRMSST yields a more accurate signal representation under impulsive noise, thereby improving accuracy and noise robustness of parameter estimation.
The automatic classification of pulmonary nodules is important for early cancer diagnosis. In this work, we aim to find the optimal state space model (SSM) for pulmonary nodule classification by leveraging Differentiable Architecture Search (DARTS). To achieve DARTS, we design the Differentiable State Space Supergraph (DSSS) to search for the optimal SSM. DSSS removes convolution, dimensionality expansion, and spatial attention around the SSM of Mamba while incorporating sampling-based multi-scale feature fusion. Additionally, anatomical scan provides multiple visual sequences of three-dimensional (3D) pulmonary nodules. During the supergraph search process, a knockout (KO) strategy is introduced to progressively remove low-probability candidate operations, thereby reducing the discrepancy between the continuous supergraph and the final discrete architecture while improving search efficiency. The discrete architecture obtained by DSSS is subsequently augmented with output feedback. By stacking the feedback-augmented cells, the final model, termed the Feedback-Augmented State Space Model (FASSM), is constructed. We conducted extensive experiments on the Lung Nodule Analysis 2016 (LUNA16) dataset, and the results show that the FASSM delivers outstanding performance. The architecture search was completed in less than six hours. The resulting FASSM achieved a mean classification accuracy of 93.22% across folds 5-9 with 8.59 M parameters. FASSM achieves a favorable balance between classification performance and model scale, with 8.59 M parameters and a computational cost of 20.41 GFLOPs. Related code and results have been released at: https://github.com/GuYuIMUST/DSSS-FASSM.
Self-training semisupervised semantic segmentation has emerged as a significant approach, yet the quality of pseudo labels remains a critical factor influencing its performance. This study aims to enhance the quality of pseudo labels in self-training semisupervised semantic segmentation by introducing a novel image search module. This module augments the utilization of labeled data, thereby improving the generation of pseudo labels. Additionally, we design an adaptive soft pooling mechanism within the pooling feature map of the image search module to better capture and model intricate global feature dependencies. Furthermore, a data augmentation method named Cutin is proposed to boost the model's generalization ability and training performance. Experiments conducted on the PASCAL VOC2012 dataset demonstrate the effectiveness of our approach. Specifically, at annotation ratios of 1/16, 1/8, and 1/4, our model achieves improvements of 0.53%, 0.87%, and 0.85% in mIoU accuracy, respectively, compared to the baseline model. These results validate the progressiveness of our model, which generates higher-quality pseudo-labels to enhance segmentation performance. The related codes and results have been released at https://github.com/GuYuIMUST/STI.
Radar emitter signal recognition, as a critical component of modern electromagnetic spectrum sensing systems, plays a pivotal role with significant strategic importance and practical value in the domains of national defense, military applications, and civilian technology. Impulse noise, characterized as a typical non-Gaussian noise, these existing radar emitter signal recognition methods with a Gaussian assumption exhibit a significant degradation or even failure in the presence of impulse noise. To address this issue, this article proposes a high-precision automatic recognition method for radar emitter signals under impulse noise. First, an impulse noise suppression technique based on an outlier detection algorithm is developed, and a comprehensive analysis of the noise suppression capabilities for five outlier detection algorithms is conducted. Second, multisynchrosqueezing transform (MSST) is employed to transform the denoised signal into a time-frequency image (TFI), and fuzzy C-means (FCMs) clustering is utilized to further eliminate redundant noise. Finally, a dataset is constructed using the TFIs of the radar emitter signal, and a convolutional neural network (CNN) is trained to establish the intelligent recognition model of radar emitter signals. The simulation results indicate that, compared to existing nonlinear transform methods, the impulse noise suppression technique based on outlier detection demonstrates superior noise suppression performance, and effectively enhances the time-frequency characteristics of radar emitter signals under impulse noise. Moreover, under an impulse noise with a GSNR >= -3 dB, the proposed signal recognition method achieved an average PSR of 100% for six signals. This outcome confirms that the method is capable of accurately recognizing radar emitter signals in the presence of impulse noise, thereby successfully addressing the challenge of unreliable signal recognition under such an environment.
In the presence of impulse noise modeled by the alpha-stable distribution, conventional noise suppression methods inevitably introduce cross-terms when processing multi-component signal, leading to significant deviations in subsequent signal representation and parameter estimation. To effectively address this issue, this paper develops an impulsive noise suppression technique based on K-medoids cluster (KMC), and proposes two representation methods for multi-component linear frequency modulation (LFM) signal under impulse noise. Firstly, the reason for cross-terms introduction is analyzed from the mathematical perspective, and subsequently a KMC-based impulsive noise suppression technology is developed. Secondly, KMC-fractional Fourier transform (KMC-FRFT) and KMC-synchrosqueezing transform (KMC-SST) are proposed, enabling precise characterization of multi-component LFM signal in the fractional domain and time-frequency domain, respectively. Finally, KMC-FRFT is applied to the parameter estimation of multi-component LFM signal under impulsive noise. Simulation experiments demonstrate that, from fractional domain and time-frequency domain, KMC not only suppresses high-amplitude burst impulsive noise, but also completely resolves the cross-terms problem inherent in existing methods. On this basis, under impulsive noise, KMC-FRFT and KMC-SST effectively capture the fractional spectral characteristic and time-frequency distribution characteristic of multi-component LFM signal from complementary perspectives. For both simulated and measured impulsive noise, RMSE demonstrates that KMC-FRFT can accurately estimate the parameters of weak component signal when GSNR >= 6dB, addressing the issue of incorrect parameter estimation caused by the cross-terms interference.
Pulmonary embolism (PE) is a life-threatening clinical problem where early diagnosis and prompt treatment are essential to reducing morbidity and mortality. While the combination of CT images and electronic health records (EHR) can help improve computer-aided diagnosis, there are many challenges that need to be addressed. The primary objective of this study is to leverage both 3D CT images and EHR data to improve PE diagnosis. First, for 3D CT images, we propose a network combining Swin Transformers with 3D CNNs, enhanced by a Multi-Scale Feature Fusion (MSFF) module to address fusion challenges between different encoders. Secondly, we introduce a Polarized Self-Attention (PSA) module to enhance the attention mechanism within the 3D CNN. And then, for EHR data, we design the Tabular Transformer for effective feature extraction. Finally, we design and evaluate three multimodal attention fusion modules to integrate CT and EHR features, selecting the most effective one for final fusion. Experimental results on the RadFusion dataset demonstrate that our model significantly outperforms existing state-of-the-art methods, achieving an AUROC of 0.971, an F1 score of 0.926, and an accuracy of 0.920. These results underscore the effectiveness and innovation of our multimodal approach in advancing PE diagnosis.
In the event extraction task, the existing models use trigger words as a bridge to extract structured information, but the extraction effect is not ideal when faced with police texts without trigger words or fixed trigger words. To solve this problem, an end-to-end trigger-free word overlapping event extraction model was proposed—TFOEE. In this model, the task of extracting overlapping events without triggering words is transformed into a task of identifying relationships based on grid filling strategy, event types and word fragments. Experiments show that the accuracy, recall rate and F1 value of TFOEE model are better than those of baseline model on police text dataset. And the F1 value of the TFOEE model reached 94.1%.
Accurate differential diagnosis of pneumonia remains a challenging task, as different types of pneumonia require distinct treatment strategies. Early and precise diagnosis is crucial for minimizing the risk of misdiagnosis and for effectively guiding clinical decision-making and monitoring treatment response. This study proposes the WSDC-ViT network to enhance computer-aided pneumonia detection and alleviate the diagnostic workload for radiologists. Unlike existing models such as Swin Transformer or CoAtNet, which primarily improve attention mechanisms through hierarchical designs or convolutional embedding, WSDC-ViT introduces a novel architecture that simultaneously enhances global and local feature extraction through a scalable self-attention mechanism and convolutional refinement. Specifically, the network integrates a scalable self-attention mechanism that decouples the query, key, and value dimensions to reduce computational overhead and improve contextual learning, while an interactive window-based attention module further strengthens long-range dependency modeling. Additionally, a convolution-based module equipped with a dynamic ReLU activation function is embedded within the transformer encoder to capture fine-grained local details and adaptively enhance feature expression. Experimental results demonstrate that the proposed method achieves an average classification accuracy of 95.13% and an F1-score of 95.63% on a chest X-ray dataset, along with 99.36% accuracy and a 99.34% F1-score on a CT dataset. These results highlight the model’s superior performance compared to existing automated pneumonia classification approaches, underscoring its potential clinical applicability.
This study evaluates the performance of mainstream large language models (LLMs) in Chinese security generation tasks, examines the potential security risks associated with these models, and proposes strategies for mitigating these risks. To this end, we developed the multidimensional security question answering (MSQA) dataset and the multidimensional security scoring criteria (MSSC). This study compares the performance of three models across six distinct security tasks. Pearson correlation analysis was conducted using GPT-4 and questionnaires, while automatic scoring was implemented using GPT-3.5-Turbo and Llama-3. Experimental results reveal that ERNIE Bot excels in ideology and ethics evaluation, ChatGPT demonstrates strong performance in assessing rumors, false information and privacy security, and Claude performs well in evaluating factual fallacies and social biases. Additionally, the fine-tuned model showed effectiveness in security scoring tasks, and the proposed Security Tips Expert (ST-GPT) successfully mitigates security risks. Despite the promising results, all models exhibit inherent security risks. Based on these findings, we recommend that both domestic and international models adhere to the legal frameworks of their respective jurisdictions, minimize AI hallucinations, continuously expand training corpora, and undergo regular updates and iterations to enhance their reliability and safety.
Acoustic scene classification aims to recognize the scenes corresponding to sound signals in the environment, but audio differences from different cities and devices can affect the model’s accuracy. In this paper, a time–frequency–wavelet fusion network is proposed to improve model performance by focusing on three dimensions: the time dimension of the spectrogram, the frequency dimension, and the high- and low-frequency information extracted by a wavelet transform through a time–frequency–wavelet module. Multidimensional information was fused through the gated temporal–spatial attention unit, and the visual state space module was introduced to enhance the contextual modeling capability of audio sequences. In addition, Kolmogorov–Arnold network layers were used in place of multilayer perceptrons in the classifier part. The experimental results show that the proposed method achieves a 56.16% average accuracy on the TAU Urban Acoustic Scenes 2022 mobile development dataset, which is an improvement of 6.53% compared to the official baseline system. This performance improvement demonstrates the effectiveness of the model in complex scenarios. In addition, the accuracy of the proposed method on the UrbanSound8K dataset reached 97.60%, which is significantly better than the existing methods, further verifying the generalization ability of the proposed model in the acoustic scene classification task.
Recently, semi-supervised learning has demonstrated significant potential in the field of medical image segmentation. However, the majority of the methods fail to establish connections among diverse sample data. Moreover, segmentation networks that utilize fixed parameters can impede model training and even amplify the risk of overfitting. To address these challenges, this paper proposes an adversarial consistency-based semi-supervised segmentation method, leveraging a dual multiscale mean teacher model. First, by designing a discriminator network with adaptive feature selection and training it alternately with the segmentation network, the method enhances the segmentation network's ability to transfer knowledge from the limited labeled data to the unlabeled data. The discriminator evaluates the quality of the segmentation network's results for both labeled and unlabeled data, while simultaneously guiding the network to learn consistency in segmentation performance throughout the training process. Second, we design a Triple-attention dynamic convolutional (TADC) module, which allows the convolution kernel parameters to be adjusted flexibly according to different input data. This improves the feature representation capability of the network model and helps reduce the risk of overfitting. Finally, we propose a novel feature selection and fusion module (FSFM) within the segmentation network, which dynamically selects and integrates important features to enhance the saliency of key information, improving the overall performance of the model. The proposed adversarial consistency-based semi-supervised segmentation method is applied to the MosMedData dataset. The results demonstrate that the segmentation network outperforms the baseline model, achieving improvements of 3.83%, 3.97%, 3.14% in terms of Dice, Jaccard, and NSD scores, respectively, for the segmentation of pneumonia lesions. The proposed segmentation method outperforms state-of-the-art segmentation networks and demonstrates superior potential for segmenting pneumonia lesions, as evidenced by extensive experiments conducted on the MosMedData and COVID-19-P20 datasets.
To enhance the resolution of synchrosqueezing transform (SST) in non-stationary signal representation, an optimization synchrosqueezed fractional wavelet transform (SSFRWT) is proposed, which possesses rigorous mathematical principle and high resolution. First, the definition, properties, and principles of SSFRWT are presented. On this basis, a time-fractional-frequency (TFF) analysis method is established utilizing SSFRWT. The experimental results demonstrate that SSFRWT is capable of establishing a high-resolution TFF representation for chirp-type signals, surpassing existing methods in terms of noise robustness and energy concentration. Lastly, leveraging the signal TFF representation, SSFRWT is successfully applied to the chirp signal parameter estimation and multi-component signal separation, yielding superior estimation results and reconstructed signal compared to SST. Notably, SSFRWT is also innovatively employed in the field of optical measurement, achieving high-precision measurement of the curvature radius of convex lens.
Currently, convolutional neural networks have demonstrated outstanding efficiency in heart sound detection and automatic diagnosis of cardiovascular diseases. However, due to the non-stationary nature and complex data patterns caused by environmental noise and stethoscope differences, traditional neural networks are limited in extracting discriminative features. This article proposes a convolutional neural network based on tensor decomposition to address this issue. This model uses a convolutional neural network with four parallel structures to extract audio features of heart sound signals and introduces a tensor network to use tensor decomposition to perform low-rank approximation on the convolutional kernel, compress model parameters, reduce redundancy, and improve performance. When processing feature data, the model divides large areas of features into locally unordered small areas to achieve feature compression and reorganization, ensuring that crucial information is preserved while compressing parameters. The model can accurately capture spatial structural information and critical features by refining the matrix product state layer. Experiments were conducted on the 2016 PhysioNet/CinC Challenge and the Yaseen heart sound public dataset, the experimental results show that the proposed method has an accuracy of 96.4% and 99.2% on two datasets, specificity of 99.1% and 99.8%, demonstrating its excellent generalization ability and diagnostic accuracy.
Introduction: This study aims to propose and evaluate a two-stage semi-supervised segmentation framework with dual multiscale uncertainty estimation and graph reasoning, addressing the challenges of obtaining high-precision pixel-level labels and effectively utilizing unlabeled data for accurate pneumonia lesion segmentation. Methods: First, we design a guided supervised training strategy for modeling aleatoric uncertainty (AU) at dual scales, reducing the impact on segmentation performance caused by aleatoric uncertainties introduced by blurred lesions and their boundaries in the image. Second, we design a training strategy for multi-scale noisy pseudo-label correction to reduce the cognitive bias problem caused by unreliable predictions in the model. Finally, we design a new combination of fused feature interaction graph reasoning (FIGR) and attention modules, which enables the network model to better capture image features in small infected regions. Results: Our study was validated using the MosMedData public dataset. The proposed algorithm improves the performance by 1.25%, 1.03%, 2.98%, and 0.59% on Dice, Jaccard, normalized surface dice (NSD), and average distance of boundaries (ADB), respectively, compared to the baseline model. Discussion: Our semi-supervised pneumonia segmentation framework, through two-stage multi-scale uncertainty estimation and modeling, significantly improves segmentation performance by leveraging unlabeled data and addressing uncertainties, offering clinical benefits in pneumonia diagnosis while facing challenges in generalization and computational efficiency that future work will target with GAN-based data synthesis and architecture optimization. Conclusion: It can be convincingly concluded that the proposed algorithm is of profound importance and value in the domain of clinical practice.
Speech emotion recognition, as an important research area, aims to automatically identify and classify the emotional state of the speaker from the speech signals. It is often difficult for a single feature to fully reflect the emotional information in speech, and multiple different features are complementary to each other, therefore, combining multiple features is often the key to improve the accuracy of speech emotion recognition. Based on this problem, this paper proposes a speech emotion recognition model SXANet based on multifeature fusion, which fuses three different features as inputs, and learns discourse-level features and frame-level features of the speech information through parallel spatial channel attentional convolution module and extended long and short-term memory network module, respectively. In addition, important channel and spatial features are also focused in the spatial channel attention convolution module due to the presence of spatial and channel attention mechanisms, thus providing richer features. In addition, the output of the parallel module is balanced by the weighted attention fusion mechanism, which improves the overall efficiency and effectiveness of the model. Extensive experiments are conducted on RAVDESS and EMODB datasets, and the experimental results show that the proposed method has an accuracy of $91.97 \%$ and $98.15 \%$ on the two datasets, respectively, which proves the effectiveness and superiority of SXANet.