Speech emotion recognition (SER) in noisy environments remains challenging because noise distortions may obscure emotion-related speech cues and weaken the generalization ability of models trained on clean speech. To address this problem, we propose an Uncertainty-Balanced Exponential Synergistic Distillation (UESD) framework for robust SER under noisy conditions. UESD consists of three main components. First, an Exponential Gain Attention (EGA)-Former backbone is designed, which uses confidence-guided reweighting and exponential enhancement to emphasize emotion-salient regions while suppressing noise-corrupted components. Second, a Synergistic Strategy Distillation mechanism is introduced to improve clean-to-noisy knowledge transfer. Through a meta-learning-based bilevel optimization process, the teacher model adaptively adjusts its teaching strategy according to the student model’s representation bottlenecks under noisy inputs. Finally, we integrate an uncertainty-based loss fusion mechanism that leverages the principle of homoscedastic uncertainty to adaptively balance the weights of heterogeneous tasks, thereby stabilizing multi-objective optimization. Experiments on IEMOCAP, CASIA, and EMODB under multiple noise types and signal-to-noise ratio (SNR) levels show that UESD achieves competitive performance and improved robustness, especially under severe noise conditions. At 0 dB SNR, compared with the best competing result for each metric under the same setting, UESD achieves absolute gains of 1.76%, 3.20%, and 2.90% in WA, UA, and WF1 on IEMOCAP, respectively. The corresponding gains in WA, UA, and WF1 are 1.76%, 1.76%, and 1.50% on CASIA, and 0.45%, 1.61%, 1.52% on EMODB, respectively. These results suggest that UESD is particularly effective in preserving emotion-related cues under low-SNR conditions. The code and model are available at: https://github.com/kcc98353-wq/UESD.
MDA5+ dermatomyositis (MDA5+DM) frequently complicates interstitial lung disease (ILD), especially rapidly progressive ILD (RP-ILD), which worsens prognosis. Traditional assessments cannot directly reflect fibroblast activation status, a key driver of pulmonary fibrosis. 68Ga-FAPI-04 PET/CT targets fibroblast activation protein (FAP), but its value in MDA5+DM remains understudied. Fourteen MDA5+DM patients and 5 controls were enrolled. 68Ga-FAPI-04 PET/CT was performed to obtain SUVmax of lungs, muscles, and other organs. Clinical data, pulmonary function tests, and deep learning-quantified HRCT indices were collected. Follow-ups for pulmonary function and HRCT were conducted. Correlations between SUVmax, clinical features, and follow-up outcomes were analyzed. Whole-lung SUVmax was significantly higher in MDA5+DM patients (3.32 vs. 0.67, p < 0.0001), with left lower lobe showing the most prominent difference (p = 0.0003). SUVmax strongly correlated with HRCT inflammation/fibrosis indices (e.g., left lower lobe vs. PII at 3 months: r = 0.9863, p < 0.0001) and small-to-medium airway function (lung vs. MMEF 75/25: r=-0.8168, p = 0.0133). Baseline SUVmax predicted 3–12-month lesion progression and lung function changes, with deceased patients having higher baseline whole-lung SUVmax (p = 0.008). This exploratory study suggests that 68Ga-FAPI-04 PET/CT may reflect fibroblast activation and the active inflammatory-fibrotic process, assesses early pulmonary impairments and monitoring lesion progression in MDA5+DM. It provides a valuable hypothesis-generating tool for clinical evaluation, warranting further validation.
To address the difficulty of jointly modeling local and global information in deep clustering based single channel speech separation and the poor adaptability of clustering under unknown speaker conditions, a single channel speech separation method based on frequency domain joint attention is proposed. The method enhances structural representation through the smoothed modulation amplitude spectrum, models multi scale modulation features using a multi scale frequency domain convolution module, and improves feature discriminability through joint attention and triple gated fusion. In addition, a confidence screening based adaptive clustering strategy is introduced to automatically match the clustering structure and reconstruct the target speech. Experimental results show that the proposed method outperforms current state of the art methods in PESQ, SDRi, and SI-SDRi on mixed datasets with unknown speakers.
Speech emotion recognition (SER) hinges on extracting discriminative features from complex acoustic signals. While existing methods capture local time-frequency patterns well, they often neglect non-local semantic dependencies and struggle to model dynamic emotional variations with sufficient adaptivity. To address these limitations, we propose a speech emotion recognitionmodel based on a Structural--Semantic Joint Dual-Branch Graph Convolutional Network (SJDGCN) and a Multi-Head Hybrid Dynamic Attention Network (MH-HDAN). First, the SJDGCN leverages a structural--semantic adjacency strategy to model both sequential structure and semantic similarity across temporal and frequency dimensions, capturing local and non-local emotional dependencies. Second, the MH-HDAN projects features into multiple subspaces, where each head combines Gaussian and Laplace attention mechanisms to learn complementary distribution characteristics. This enables adaptive focus on salient emotional regions with varying smoothness and intensity, enhancing discrimination of subtle emotional differences. Experiments on IEMOCAP, CASIA, and EMODB corpora achieve weighted accuracy (WA)/unweighted accuracy (UA) of 77.48%/76.32%, 61.13%/61.13%, and 86.96%/86.18%, surpassing several state-of-the-art methods and validating the effectiveness of the proposed approach. The code and model are available at:https://github.com/Re799/SSGCN_with_MH-HDAN_for_SER.
Automated bird species recognition (BSR) is crucial for biodiversity monitoring, but its accuracy is often hampered by geographic variation in bird vocalizations, which is frequently described as dialectal at the population level. This work uses the term dialect region in an operational sense to refer to geographically separated recording regions defined in benchmark corpora that are dominated by different vocal variants across populations, rather than to fine-grained, species-specific dialect types. We aim to address the challenge of generalizing across such dialect-dominated regions by introducing a novel approach leveraging well-established auditory-inspired features known for their computational robustness. We propose Human Auditory Representation Learning (HARL), a framework that integrates Gammatone- and Mel-spectrogram features to capture frequency selectivity and invariance to acoustic variations through their spectral efficiency and empirical success in audio processing. These complementary auditory representations are processed by a dual-stream ResNet50 architecture, with a multi-head attention mechanism to emphasize discriminative spectral–temporal patterns. In cross-dialect evaluation on the D3BV benchmark and cross-site tests on two field datasets (S1 and S2), the approach outperformed strong baselines, raising F1-score by up to 24.31% on D3BV and 28.23% on S1 and S2. Performance remained stable across noise conditions from −10 to +10 dB signal-to-noise ratio, indicating robustness for real-world deployment. These findings showed that bridging HARL with deep learning delivered a scalable and accurate solution for biodiversity monitoring, enabling reliable species recognition across diverse geographies and acoustic conditions.
To improve the reliability of diesel engine and predict its remaining service life under different operating conditions, this study proposes a piston life prediction method based on digital twin and multi-physical field coupling. First, a Simulink mechanism model of the engine is constructed, and parameters such as rotational speed and load under different operating conditions, specified rotational speeds, different loads, and different crank angles are inputted into the mechanism model to obtain the in-cylinder pressures and in-cylinder temperatures under multiple operating conditions, specified rotational speeds, different loads, and different crank angles. The mechanism model is calibrated using the sparrow search algorithm to improve the accuracy and reliability of the model; the mechanism model and the optimization search algorithm form a digital twin model, and the constructed digital twin model can generate a large amount of high-fidelity twin data. Secondly, the temperature field, stress field and thermal-mechanical coupling stress field of the piston were calculated and analyzed using finite element analysis to obtain the force of the piston under different physical fields. Finally, life prediction of the piston was performed by adding a fatigue analysis module on the basis of thermal stress analysis. The accuracy of the life prediction was verified using experimental data. The results show that the fatigue life of the piston under actual operating conditions is predicted based on the established life prediction model. The calculated results show that the minimum fatigue life of the piston under the operating conditions of 2000 r/min and 90% load is 48,330 hours, which satisfies the manufacturer’s design life of 47,000 hours.
Cross-domain fault diagnosis of marine diesel engines presents significant challenges due to variations in data distribution and the limited availability of labeled fault samples under different operating conditions. To address this, an unsupervised domain-adaptive diagnostic framework is proposed, integrating stepwise diffusion and iterative bidirectional optimization to enhance fault identification. First, the quadratic axial attention transformer introduces a fourth weight in the axial computation to effectively capture the long-range spatio-temporal correlations in the time-frequency representations and strengthen the cross-axis contextual dependence. Next, the domain stepwise diffusion bridge utilizes Markov transform to gradually refine the significant distributional differences across domains into continuous sub-distributions, ensuring a smoother adaptation process. Finally, an iterative bidirectional optimization strategy is proposed to dynamically coordinate the interaction between stepwise diffusion and fault classification, where two complementary learning directions are alternately executed to preserve the semantic integrity of features. Experimental validation on a self-constructed dataset covering multiple operating conditions demonstrates the effectiveness of the proposed approach, achieving 93.80% average accuracy, 93.75% precision, and 93.45% recall. This approach not only breaks through the limitations of existing domain alignment methods and provides a brand new solution for cross-domain fault diagnosis, but also provides a wide range of implications for future research and applications in this field. The code and model are available at: https://github.com/lazyJzr/UDAtask.
Deep learning based object detection methods have achieved promising performance recently. However, these methods lack sufficient capabilities to handle satellite images owing to the fact that small-sized objects in remote sensing images are difficult to detect. To address this issue, we propose a novel small object detection method based on YOLO X named Attention Cross Stage Transformers Network (ACSTNet). Specifically, a novel backbone network, Multi-scale Cross Fusion Network (MCFNet) is constructed to capture semantic dependencies between pixels over long distances and increase the depth-interaction information at different levels. Meanwhile, a new feature fusion layer is added to the upper feature output layer of dark3, allowing the model to maximize the retention of low-level features of small objects and to locate them more accurately. Furthermore, to address the problem of the inaccurate feature extraction caused by overlapping and occlusion of dense objects, we propose an efficient channel and space normalized fusion attention mechanism (ECSNFAM), which is composed of channel attention, space attention, and batch normalization attention branches, using residual structure to enhance the sensitivity of the attention mechanism for small targets. Experiments are conducted to evaluate the performance of the general remote sensing dataset, and the results show that our proposed method improves the mean Average Precision (mAP) by 1.2% and 1.4% on the DIOR and the RSODDATA datasets compared with the YOLO X. The source code is available at https:github.com/Wei-JL/ACSTNet.git.
Prior knowledge of the medical domain has consistently enriched medical image analysis, yet its full potential remains to be explored. Our Deformable Symmetry Attention (DSA) aims to leverage anatomical symmetry prior, commonly used by clinicians, to enhance the automated analysis of medical images, especially in inherently grainy nuclear medicine images. DSA processes an explicit feature alignment attention. It starts with a prior alignment by reflecting the input image to leverage symmetry, followed by a fine-grained alignment using deformable image registration. Our novel Rotational Spatial Transformation Function (RSTF) enhances the registration by estimating local displacement with rotation in relative coordinates, enforcing local smoothness under anatomical variations. Additionally, the Attention Redistribution (AR) module addresses pixel-level mismatch in grainy nuclear medicine images by a set of learnable Gaussian kernels. By translating the anatomical symmetry prior into a locality prior, DSA seamlessly integrates with and enhances sophisticated convolutional neural networks (CNNs) and transformer-based segmentation models. In our study, the DSA framework, when combined with robust baseline models, demonstrated significant performance enhancements across multiple challenging medical imaging tasks. Specifically, it achieved a 3.8% to 8.48% improvement in segmenting lesions in adolescent Whole-Body Bone Scans (WBS) and a 2.88% improvement in the extended 3D MRI brain tumor segmentation experiment. Additionally, the method was validated in asymmetric scenarios, showing an improvement of 0.63% to 10.57% in chest X-ray anatomical segmentation. These results underscore DSA's effectiveness and versatility across diverse experimental settings. It presents a novel and promising perspective for explicitly leveraging prior knowledge in medical image analysis.
To address the issue of acoustic signals being easily interfered by noise and insufficient information in traditional fault acoustic recognition methods, this paper proposes a fault acoustic recognition method based on crossmodal distillation and semantic calibration. By introducing vibration modality as auxiliary information, a modal interference suppressor and a semantic calibration module are designed, combined with triplet loss and adaptive contrastive loss functions to enhance feature representation and transfer efficiency. Experimental results show that the proposed method achieves recognition accuracies of 98.57% and 95.13% on the UORED-VAFCLS and JUST datasets, respectively, significantly outperforming existing methods and effectively improving the accuracy and robustness of fault recognition.
Computed tomography attenuation correction (CTAC) is commonly used in cardiac SPECT imaging to reduce soft-tissue attenuation artifacts. However, CTAC is prone to inaccuracies due to CT artifacts and SPECT-CT mismatch, along with additional radiation exposure to patients. Thus, these limitations have led to increasing interest in CT-free AC, with deep learning (DL) offering promising solutions. We proposed a new DL-based CT-free AC methods for cardiac SPECT. We developed a feature alignment attenuation correction network (FA-ACNet) based on the 3D U-Net framework to generate predicted DL-based AC SPECT (Deep AC). The network was trained on 167 cardiac SPECT/CT studies using 5-fold cross validation and tested in an independent testing set (n = 35), with CTAC serving as the reference. During training, multi-scale features from non-attenuation-corrected (NAC) SPECT and CT were processed separately and then aligned with the encoded features from NAC SPECT using adversarial learning and distance metric learning techniques. The performance of FA-ACNet was evaluated using mean square error (MSE), structural similarity index (SSIM) and peak signal-to-noise ratio (PSNR). Additionally, semi-quantitative evaluation of Deep AC images was performed and compared to CTAC using Bland-Altman plots. FA-ACNet achieved an MSE of 16.94 ± 2.03 × 10− 6, SSIM of 0.9955 ± 0.0006 and PSNR of 43.73 ± 0.50 after 5-fold cross validation. Compared to U-Net, MSE and PSNR improved by aligning multi-scale features from NAC SPECT and CT with those from NAC SPECT. In the testing set, FA-ACNet achieved an MSE of 11.98 × 10− 6, SSIM of 0.9976 and PSNR of 45.54. The 95
Speech emotion recognition (SER) aims to identify the speaker's emotional states in specific utterances accurately. However, existing methods still face feature confusion when attempting to recognize certain emotions because traditional acoustic feature extraction methods fail to capture dynamic emotional changes, blurring emotional boundaries. Additionally, existing classification networks (CNs) are constrained by fixed learning strategies, hindering their ability to capture subtle emotional nuances and resulting in label confusion. To address these two issues, we introduce 3D multiresolution modulation filtered cochleogram (MMCG) features by computing the deltas and delta-deltas of MMCG features to enhance the dynamic emotional changes and produce distinct emotional boundaries. We then customize a conditional emotion feature diffusion (CEFD) module, which progressively diffuses features based on emotional context to retain emotional nuances effectively and reduce reliance on conditioned information. In addition, a confidence filtering module is used to filter diffused features based on confidence-based posterior probabilities to ensure enhanced feature discrimination. We design a flexible training strategy named the progressive interleaved learning strategy (PILS) to learn further complex emotional nuances, which consists of two alternating stages: fine-tuning the CN parameters and supervising the CEFD output. Testing on the IEMOCAP, CASIA, and EMODB corpora demonstrates significant performance improvements in SER.
In light of the challenges associated with effectively leveraging interactive information across channels and spatial dimensions derived from skeleton key point data, a continuous action recognition method based on multi-branch attention and spatiotemporal graph convolutional network was proposed. Firstly, the interactive information between channel attention and spatial dimension is extracted by embedding multi-branch attention. The features obtained by spatial graph convolution are weighted by multi-branch attention, and then the subsequent temporal graph convolution operation is executed to enhance the representation ability of the model for spatial relations. Finally, the ability of a spatiotemporal graph convolutional network to capture spatiotemporal relationships was used to realize the recognition of continuous actions. The findings from the experiments indicate that the proposed model outperforms current methods in terms of recognition accuracy on the NTU-RGBD and Kinetics datasets.
Speech Emotion Recognition (SER) in noisy environments is challenging due to the overlap between emotional and noise-related signals. We propose a novel emotion-diffusion approach to enhance SER performance in noisy conditions that transfers emotional information from clean to noisy speech through a multi-stage diffusion process. First, Mel Frequency Cepstral Coefficients and their delta features are extracted to capture emotional dynamics. Additionally, an Emotional Bidirectional Mamba Encoder with a Multi-time View Bidirectional State Space Model is designed to capture temporal emotional patterns. Next, the Emotion-Transferring Diffusion Network (ETDN) applies confidence filtering to retain key emotional features, ensuring effective emotion transfer despite noise. Finally, the Confidence-Guided Mutual Learning strategy refines noisy features for the Classification Network (CN), while the CN supervises the ETDN to maintain label consistency. Experiments on the IEMOCAP dataset show significant improvements in weighted accuracy, unweighted accuracy, and weighted F1 score across various signal-to-noise ratios.
Speech emotion recognition (SER) in noisy environments is challenging due to the overlap of emotional cues with background noise. This article proposes a novel approach to transfer emotional information from clean to noisy speech, ensuring robust recognition even in adverse conditions. First, 3-D multiresolution modulated filtered cochleogram features are extracted to capture dynamic emotional information, while delta and delta-delta features enhance emotional dynamics in both time and frequency domains. Additionally, the Emotional BiMamba Encoder, utilizing bidirectional parallel processing through the Multitime View Bidirectional state-space model is designed to capture complex emotional patterns across temporal scales, preserving short- and long-term dependencies. Next, the adaptive emotion denoising diffusion (AEDD) based on a diffusion-denoising probabilistic model, applies confidence filtering to select representative emotional segments and transfer emotional information from clean to noisy speech, addressing noisy emotional data scarcity. Finally, the iterative confidence learning strategy (ICLS) enhances the classification network (CN) through a two-stage learning process, progressively adapting CN to noisy feature distributions while ensuring consistency during diffusion. The experimental results show significant improvements over the state-of-the-art methods on three datasets: on interactive emotion dynamics capture (IEMOCAP), weighted accuracy (WA) increases by 5.07%, unweighted accuracy (UA) by 5.23%, and weighted average F1 (WF1) by 5.44%; on Chinese Academy of Sciences Automation Institute of Automation (CASIA), WA and UA rise by 2.72%, and WF1 by 2.23%; and on Berlin German Emotion Speech Bank (EMODB), WA improves by 3.46%, UA by 3.23%, and WF1 by 3.45%. These consistent gains across different signal-to-noise ratios further validate the effectiveness of the proposed method.
Due to the abstraction and complexity of speech data, accurately recognizing human emotions through speech is a very challenging task. A difficult problem to solve in many previous studies is the severe misclassification of several specific emotions, such as “Happy” and “Sad”. In this paper, we propose a dual-path framework with Multi-Head Bayesian Coattention (MHBA) for speech emotion recognition under human-machine interaction (HMI), which consists of the context emotional feature extraction branch and the long-term prosody information supplement branch. Specifically, in the context emotional feature extraction branch, we firstly select two specific transformer blocks of pre-trained Wav2vec2.0 to extract emotional representation with rich short-term acoustic information as short-term emotional representation. Furthermore, we employ the United Attention Network (UAN) to selectively discover target emotion regions from representation, wherein Context-aware Attention Pooling mechanism (CAAP) capturing the contextual dependencies of spatiotemporal features. Meanwhile, in the long-term prosody information supplement branch, we utilize the long-term prosody feature as supplementary input to complement the missing prosody emotional information in the Wav2vec2.0 feature. Finally, to effectively integrate information from both branches, we enhance the coattention with the Bayesian Attention Module (BAM), which estimates prior distributions using emotion-relevant knowledge. Our approach has been evaluated on the benchmark dataset IEMOCAP, yielding promising experimental results. Specifically, we observed significant improvements in both weighted accuracy(WA) and unweighted accuracy(UA). Our approach achieves an absolute increase of 3.68% in WA and 0.49% in UA, further validating its effectiveness and performance.
In speech emotion recognition, existing models often struggle to accurately classify emotions with high similarity. In this paper, we propose a novel architecture that integrates a multi-view attention network (MVAN) and diffusion joint loss to alleviate confusion by placing a stronger focus on emotions that are challenging to classify accurately. First, we use logarithmic Mel-spectrograms (log-Mels), deltas, and delta-deltas of log-Mels as three-dimensional features to minimize external interference. Then, we design the MVAN to extract effective multi-time scale emotion features, where the channel and spatial attention are used to selectively localize the regions in the input features related to the target emotion. A Multi-time view bidirectional long and short-term memory network is used to extract the shallow edge features and deep semantic features, and multi-scale self-attention fuses these features through cross-scale attention fusion to obtain multi-time scale emotion features. Finally, a diffusion joint loss strategy is introduced to distinguish the emotional embeddings with high similarity by the generated complex emotion triplets in a diffusing fashion. We evaluated our proposed method on the Interactive Emotional Mood Binary Motion Capture (IEMOCAP), Chinese Academy of Sciences Automation Institute of Automation (CASIA), and Berlin German Emotion Speech Bank (EMODB) corpus. The results show significant improvements over existing methods, achieving 86.87% WA, 86.60% UA, and 86.82% WF1 on IEMOCAP; 70.74% WA, 70.74% UA, and 70.25% WF1 on CASIA; and 93.65% WA, 91.13% UA, and 92.26% WF1 on EMODB. These results confirm the superiority of our method. Our code and model are available at https://github.com/Littleznnz/MVAN-DiffSEG.
ABSTRACT:A 57-year-old woman who had persistent symptoms of transthyretin cardiac amyloidosis underwent 99m Tc-pyrophosphate ( 99m Tc-PYP) scintigraphy. The 99m Tc-PYP planar and SPECT/CT fusion image showed diffuse myocardial uptake and multiple fractures of the sternum and ribs. These fractures interfered with semiquantitative scores of 99m Tc-PYP uptake, leading to false positive in 99m Tc-PYP imaging.
Automated tongue segmentation plays a crucial role in the realm of computer-aided tongue diagnosis. The challenge lies in developing algorithms that achieve higher segmentation accuracy and maintain less memory space and swift inference capabilities. To relieve this issue, we propose a novel Pool-unet integrating Pool-former and Multi-task mask learning for tongue image segmentation. First of all, we collected 756 tongue images taken in various shooting environments and from different angles and accurately labeled the tongue under the guidance of a medical professional. Second, we propose the Pool-unet model, combining a hierarchical Pool-former module and a U-shaped symmetric encoder-decoder with skip-connections, which utilizes a patch expanding layer for up-sampling and a patch embedding layer for down-sampling to maintain spatial resolution, to effectively capture global and local information using fewer parameters and faster inference. Finally, a Multi-task mask learning strategy is designed, which improves the generalization and anti-interference ability of the model through the Multi-task pre-training and self-supervised fine-tuning stages. Experimental results on the tongue dataset show that compared to the state-of-the-art method (OET-NET), our method has 25% fewer model parameters, achieves 22% faster inference times, and exhibits 0.91% and 0.55% improvements in Mean Intersection Over Union (MIOU), and Mean Pixel Accuracy (MPA), respectively.