Objective State-of-the-art speaker verification models typically rely on fixed receptive fields, which limits their ability to represent multi-scale acoustic patterns while increasing parameter counts and computational loads. Speech contains layered temporal-spectral structures, yet the use of dynamic receptive fields to characterize these structures is still not well explored. The design principles for effective dynamic receptive field mechanisms also remain unclear. Methods Inspired by the non-linear coupling behavior of tidal surges, a Tide-Ripple Convolution (TR-Conv) layer is proposed to form a more effective receptive field. TR-Conv constructs primary and auxiliary receptive fields within a window by applying power-of-two interpolation. It then employs a scan-pooling mechanism to capture salient information outside the window and an operator mechanism to perceive fine-grained variations within it. The fusion of these components produces a variable receptive field that is multi-scale and dynamic. A Tide-Ripple Convolutional Neural Network (TR-CNN) is developed to validate this design. To mitigate label noise in training datasets, a total loss function is introduced by combining a NoneTarget with Dynamic Normalization (NTDN) loss and a weighted Sub-center AAM Loss variant, improving model robustness and performance. Results and Discussions The TR-CNN is evaluated on the VoxCeleb1-O/E/H benchmarks. The results show that TR-CNN achieves a competitive balance of accuracy, computation, and parameter efficiency (Table 1). Compared with the strong ECAPA-TDNN baseline, the TR-CNN (C=512, n=1) model attains relative EER reductions of 4.95%, 4.03%, and 6.03%, and MinDCF reductions of 31.55%, 17.14%, and 17.42% across the three test sets, while requiring 32.7% fewer parameters and 23.5% less computation (Table 2). The optimal TR-CNN ( C=1 024, n =1) model further improves performance, achieving EERs of 0.85%, 1.10%, and 2.05%. Robustness is strengthened by the proposed total loss function, which yields consistent improvements in EER and MinDCF during fine-tuning (Table 3). Additional evaluations, including ablation studies (Tables 5 and 6), component analyses (Fig. 3 and Table 4), and t-SNE visualizations (Fig. 4), confirm the effectiveness and robustness of each module in the TR-CNN architecture. Conclusions This research proposes a simple and effective TR-Conv layer built on the T-RRF mechanism. Experimental results show that TR-Conv forms a more expressive and effective receptive field, reducing parameter count and computational cost while exceeding conventional one-dimensional convolution in speech feature modeling. It also exhibits strong lightweight characteristics and scalability. Furthermore, a total loss function combining the NTDN loss and a Sub-center AAM loss variant is proposed to enhance the discriminability and robustness of speaker embeddings, particularly under label noise. TR-Conv shows potential as a general-purpose module for integration into deeper and more complex network architectures.
Language identification (LID) is a key component in downstream tasks. Recently, the self-supervised speech representation learned by Wav2Vec 2.0 (W2V2) has been demonstrated to be very effective for various speech-related tasks. In LID, it is commonly used as a feature extractor for frame-level feature extraction. However, there is currently no effective method for extracting temporal information from frame-level features to enhance the performance of LID systems. To deal with this issue, we propose a LID framework based on deep temporal representation (DTR) learning. First, the W2V2 model is used as a front-end feature extractor. This model can capture contextual representations from continuous raw audio in which temporal dependencies are embedded. Then, a temporal network responsible for learning temporal dependencies is proposed to process the output of W2V2. This temporal network comprises a temporal representation extractor for extracting utterance-level representations and a temporal regularization term to impose constraints on temporal dynamics. Finally, the temporal dependencies are used as utterance-level representations for the subsequent classification. The proposed DTR method is evaluated on the OLR2020 database and compared to other state-of-the-art methods. The results show that the proposed method achieves decent experimental performance on all the three tasks of OLR2020 database.
Language identification is a core technology in the field of multilingual information processing, and its goal is to automatically identify the corresponding language by analyzing speech signals. However, existing approaches face three major challenges: (1) the limited representation capacity of single-scale features for complex speech patterns, (2) suboptimal model performance caused by single-objective optimization, and (3) the computational intensity of processing high-dimensional speech data. To effectively address these computational challenges, powerful supercomputing capabilities have become the key support for advancing research. Based on this, this study innovatively proposes a method based on multi-scale features recursive fusion and adaptive loss, which is referred to as MSFRF-AL. Specifically, the wav2vec2.0 pre-trained model is first used for feature extraction. The extracted features are then sent to the multi-scale feature recursive fusion network (MSFRF-Net) for fusion from different time scales. Then, the fused features are sent to the simple back-end architecture composed of the statistics pooling layer and the fully connected layer. Finally, the proposed adaptive loss is employed to optimize the model training process. This objective function consists of cross-entropy and efficient triplet loss, combining classification ability and feature representation ability to simultaneously enhance feature discrimination and classification accuracy. Experiments are verified on three tasks of the OLR2020 dataset. The experimental results show that this method can effectively reduce the interference of the external environment on language identification, especially demonstrating stronger robustness in cross-channel and noisy environments, and significantly improving the performance of language identification.
Synthetic speech detection systems have been proposed to bolster speaker verification systems against misinformation and speech spoofing. At present, the most effective methods are based on neural network models. The features extracted by different models often contain different information. Through the fusion of these models, their respective characteristics can be better combined. To fulfill this requirement, we propose the associative-discriminative fusion networks (ADFN) to fuse embeddings from different models. The ADFN can build the association between different models by the decomposition of a joint scatter matrix. And the decomposed matrices can be employed to compose the initialization transformation matrices of the input embeddings. With such structured initialization, the fusion model can start with a more suitable optimized location and achieve the better performance than each signal pre-fused model. The proposed ADFN method is evaluated on the Automatic Speaker Verification Spoofing and Countermeasures (ASVspoof) 2019 and 2021 datasets. The experimental results show that ADFN method achieves decent performance on the logical access (LA) sub-challenge of the two datasets, proving its practical applicability and effectiveness in fusing different types of embeddings.
Adversarial examples pose a significant challenge to the robustness of deep neural networks (DNNs). Traditional adversarial attacks often lack semantic features and concealment. This paper proposes AdvShadow, an adversarial example generation algorithm based on a conditional diffusion model. AdvShadow dynamically adjusts shadow perturbations to attack sensitive regions of images, generating visually realistic shadow effects that deceive deep models. Our approach ensures that adversarial shadows seamlessly integrate with image features, enhancing attack concealment and effectiveness. Experiments on multiple datasets demonstrate that AdvShadow achieves high attack success rates while maintaining high image quality, as evidenced by improved structural similarity index (SSIM) and peak signal-to-noise ratio (PSNR) scores. This study not only provides a novel adversarial attack method, but also sheds light on the robustness of deep learning models against stealthy attacks. The raw data Oxford-IIIt-Pet required to reproduce the above findings are available to download from https://www.robots.ox.ac.uk/ vgg/data/pets/ . Code will be released in https://github.com/Raineasy/AdvShadow-Camouflaged-Adversarial-Attacks-via-Conditional-Diffusion-Model-Generated-Shadows .
Food quality detection is of great importance for human health and industrial production. Currently, the common detection methods are difficult to achieve the need for fast, accurate, and non-destructive detection. In this work, an electronic nose (E-nose) detection method based on the combination of convolutional neural network combined with wavelet scattering network (CNN-WSN) and improved seahorse optimizes kernel extreme learning machine (ISHO-KELM) is proposed for identifying the quality level of a variety of food products. In the feature extraction part, the abstract features of CNN are fused with the scattering features of WSN, and the obtained CNN-WSN fusion features can characterize the original information of the food quality effectively. In the classifier design and decision-making section, chaotic mapping is used to initialize the population in the seahorse optimisation algorithm (SHO), avoiding the problem that SHO may fall into local optimal solutions. The kernel parameters and regularisation coefficients of the KELM model were then optimized by improving the locomotion, predation, and reproduction behaviors of the hippocampal populations, which solved the problem of the difficult selection of the key parameters in the model, and thus improved the accuracy and generalization of the overall model. To validate the effectiveness of the proposed food quality detection model, the E-nose system was first built and milk quality data were collected independently, and then tested on two publicly available food quality datasets as well as a self-collected milk quality dataset, respectively. The experimental results show that the food quality detection method proposed in this work has good quality assessment effect on different datasets.
In industrial production, complex working conditions often result in poor generalization of fault diagnosis models. Consequently, we propose a method called dual invariant feature domain generalization (DIFDG) for fault diagnosis of rolling bearings. This method is built upon the knowledge distillation framework, wherein phase information is extracted from 1-D signals using the Fourier transform. The phase information remains insensitive to domain changes and serves as the training data for the teacher network, enabling it to learn internally invariant features. Simultaneously, employing the knowledge distillation technique, the student network learns internally invariant features from 1-D time-series signals. To prevent information loss, 2-D time-frequency diagram data is incorporated into the student network training. The student network is trained to acquire mutually invariant features using a class-related loss function for data from various domains. This method trains a student network to extract dual invariant features from the data, enhancing the generalization to applications involving unknown data. Experimental results demonstrate that this approach outperforms other representative methods in both the Case Western Reserve University (CWRU) and NEFU_PHM datasets.
Gas mixtures, as the prevalent form of gases in real atmospheric environments, accurately realizing the detection of gas mixture concentrations is one of the practical problems commonly faced in the process of productizing electronic nose technology. In view of this, this paper proposes a two-stage gas mixture concentration detection method, which can further estimate the concentration of each gas component on the basis of determining the composition of the gas mixture. At the stage of gas mixture composition identification, the multidimensional response signal of the sensor array is reconstructed in one dimension, and the reconstructed signal is decomposed into a set of intrinsic mode functions (IMFs) containing amplitude and frequency information using variational modal decomposition (VMD), and then the gas features contained in each IMF are extracted through the sensitivity of amplitude aware permutation entropy (AAPE) to amplitude and frequency changes, and finally the gas mixture components are identified through the multiclassification relevance vector machine (MRVM) model. In the stage of gas mixture concentration detection, a fast multi-output relevance vector regression method coupled with the sparrow search algorithm (SSA-FMRVR) is proposed for the estimation of mixed gas concentration, which establishes the FMRVR through the steady-state response characteristics of the sensor array, simultaneously realizes the real-time prediction of the concentration of multiple gas components in the gas mixture by using the characteristics of multi-input and multi-output of the FMRVR, and then adopts the SSA to optimize the kernel parameter of the FMRVR in order to further improve the accuracy of the concentration estimation, and to solve the problem that the existing methods cannot take into consideration of both the prediction accuracy and the time efficiency. The experimental results show that the proposed method can effectively identify the components of the gas mixture and accurately estimate its concentration.
The operating temperature of the motor is directly influenced by motor loss, and there is a significant difference in the distribution of loss between high-speed motor and conventional speed motor, so it is extremely important to analyze the temperature field of high-speed motor. The manuscript takes a high-speed induction motor as an example for three-dimensional temperature field analysis, calculates the temperature field of the motor under various loss conditions, and obtains the characteristic relationship between different loss and winding temperature and rotor squirrel cage temperature. The rationality of the simulation model and the correctness of the analysis results are verified by comparing the motor’s full load operation test with the simulation results. Finally, the basic design ideas for high-speed motors are summarized.
In the context of dealing with limited annotated data, this paper introduces a weakly supervised whole slide image (WSI) classification approach based on contrastive learning. The proposed method aims to detect whether cancer cells have metastasized in anterior lymph nodes of breast cancer in whole slide images. Initially, small patches are extracted from whole-slide pathology images, and an unsupervised pretraining is performed on the feature extraction model using the MoCo v2 framework. Subsequently, the feature extraction model is used to extract features from the small patches. Finally, CLAM is employed to aggregate the extracted features to obtain the overall whole slide image (WSI) classification results. Experimental results demonstrate that using MoCo v2 for unsupervised pretraining of the feature extraction model achieves an accuracy of 0.8808 in the small patch classification task. Moreover, under coarse-grained WSI-level labels, the proposed approach achieves area under the receiver operating characteristic curve (AUC) values of 0.957 ± 0.0276 and 0.9442 on different datasets, outperforming typical weakly supervised and partially supervised methods in terms of classification performance.
Synthetic speech is becoming increasingly rampant, and automatic speaker verification (ASV) systems are vulnerable to its attacks. However, most current synthetic speech detection methods focus on the influence of a single feature in the detection. Since different features can represent the difference between real speech and synthetic speech to a certain extent, there must be common information between different types of features. Effectively finding and fully utilizing this information will facilitate the extraction of better discriminative features and achieve improved performance. Based on the above analysis, we propose a deep correlation network (DCN) to learn the latent common information between different embeddings. It consists of two parts, the bi-parallel network and the correlation learning network. Bi-parallel networks consist of different neural models to learn the middle-level representations from front-end acoustical features. The correlation learning network is the core part of the DCN and is proposed to explore the common information between the above middle-level features. The common information obtained after DCN processing have better discriminative ability for synthetic speech detection. Experimental results show that the proposed DCN can significantly improve the performance of synthetic speech detection system on ASVspoof 2019 and ASVspoof 2021 logical access sub-challenge.
Automatic speaker verification (ASV) systems are highly vulnerable to synthetic speech attack. And the artifacts are the key spoofing clue to distinguish real and synthetic speech. In this paper, we focus on the detection of artifacts and proposed the twice attention networks (TA-networks). It is an end-to-end network which consists of feature extraction module and back-end classifier. The feature extraction module is the core of the TA-networks, and it is a twice attention Unet (TA-Unet). It contains two sequential attention modules: (1) a five-layer U-shaped network with attention gate to first obtain the general contour of artifacts and then (2) a softmax-based filter with adaptive coefficient to dynamically highlight the differences between different frequencies, and these differences can be regarded as elaborate artifacts. After the processing of the TA-Unet, the feature maps of real and synthetic speech are more discriminative for the back-end SCG-Res2Net50 classifier. Experimental results show that the TA-networks achieve equal error rates of 1.62% on ASVspoof 2019 logical access sub-challenge, and it is significantly better than most of the other experimental models.
In recent years, deep learning technology has shown great potential in the fault diagnosis of rotating machinery based on vibration signals. However, the feature extraction and noise robustness still need to be improved. To this end, we propose a multi-scale deep neural network fault diagnosis method. Firstly, multi-scale down sampling of time-domain vibration signals. Next, the attention long short-term memory network and the fully convolutional neural network of the multi-scale convolution kernel are used for feature extraction. Then, a fusion module is utilized to fuze the extracted features. The proposed method is evaluated on the public bearing datasets. Experimental results demonstrate that the proposed method can achieve high accuracy and noise robustness.
Histopathological images classification plays a significant role in cancer diagnosis, but current deep learning methods fail to account for the unique characteristics of histopathological images. To address this limitation, we present SSANet, a Spatial and Stain Attention Network focusing on staining information in the cell nucleus and cytoplasm of histopathological images. Our approach first separates the stain channels and generates a stain attention map using a specialized stain attention module, which then activates staining information in the feature map. The experiments on two computational pathology problems, CRC-HE and BreakHis datasets, demonstrate that our method outperforms state-of-the-art methods with test F1 values up to 97.63% and 94.91% for cancer subtypes classification. Our contributions include the novel SSANet architecture, a stain attention module that enhances the focus on crucial cytosolic and cytoplasmic information, and improved classification results for histopathological images.
For the problem that noise has a great impact on the measurement data during the electrical capacitance tomography data acquisition process, a denoising method based on the truncated singular value decomposition combined with the total least squares model is proposed. Soft threshold is performed on the effective value of the truncated singular value decomposition to remove the influence of external noise interference in the measurement data, and as far as possible to retain the original characteristics of the data. For the problem of different errors in the measurement data and the coefficient matrix, a mathematical model is introduced based on total least squares. It is used to improve the total least squares iterative method and reduce both the measurement error and the coefficient matrix error. To solve the problems of slow convergence speed and low efficiency caused by the ill-posedness of the equation during the iteration, adaptive correction parameters are introduced, which effectively avoids the occurrence of local convergence and improves the speed of convergence and imaging accuracy. To solve the ill-condition of the total least squares model, a regularization matrix is added so that the imaging results can achieve the goal of overall optimization. For 12-electrode electrical capacitance tomography system, the simulation experiments are carried out based on four typical flow patterns. The results show that the algorithm effectively increases the robustness of reconstruction and improves the accuracy of the reconstructed images.
The i-vector method is one of the mainstream methods in spoken language identification (SLID). It estimates the total variability space (TVS) to obtain a low-rank representation which can characterize the language, called the i-vector. However, on small-scale datasets, low learning resources can significantly degrade the performance of SLID system. Therefore, it is necessary to improve the performance of SLID system in low-resourced condition. In this paper, we propose a common latent representation learning (CLRL) method to learn the TVS, which introduces prior information to address the lack of information in low-resourced condition. The prior information includes category label and parameter prior hypothesis. The CLRL method is evaluated on the OLR2020 dataset. Compared with other state-of-the-art methods, the CLRL method shows better performance on all datasets of different data scales. Moreover, the CLRL method can effectively improve the performance of the SLID system on low-resourced/small-scale datasets.
Deep learning methods benefit from data sets with comprehensive coverage (e.g., ImageNet, COCO, etc.), which can be regarded as a description of the distribution of real-world data.The models trained on these datasets are considered to be able to extract general features and migrate to a domain not seen in downstream.However, in the open scene, the labeled data of the target data set are often insufficient.The depth models trained under a small amount of sample data have poor generalization ability.The identification of new categories or categories with a very small amount of sample data is still a challenging task.This paper proposes a few-shot fine-grained image recognition method.Feature maps are extracted by a CNN module with an embedded attention network to emphasize the discriminative features.A channel-based feature expression is applied to the base class and novel class followed by an improved cosine similarity-based measurement method to get the similarity score to realize the classification.Experiments are performed on main few-shot benchmark datasets to verify the efficiency and generality of our model, such as Stanford Dogs, CUB-200, and so on.The experimental results show that our method has more advanced performance on fine-grained datasets.
In the research on energy-efficient networking methods for precision agriculture, a hot topic is the energy issue of sensing nodes for individual wireless sensor networks. The sensing nodes of the wireless sensor network should be enabled to provide better services with limited energy to support wide-range and multi-scenario acquisition and transmission of three-dimensional crop information. Further, the life cycle of the sensing nodes should be maximized under limited energy. The transmission direction and node power consumption are considered, and the forward and high-energy nodes are selected as the preferred cluster heads or data-forwarding nodes. Taking the cropland cultivation of ginseng as the background, we put forward a particle swarm optimization-based networking algorithm for wireless sensor networks with excellent performance. This algorithm can be used for precision agriculture and achieve optimal equipment configuration in a network under limited energy, while ensuring reliable communication in the network. The node scale is configured as 50 to 300 nodes in the range of 500 × 500 m2, and simulated testing is conducted with the LEACH, BCDCP, and ECHERP routing protocols. Compared with the existing LEACH, BCDCP, and ECHERP routing protocols, the proposed networking method can achieve the network lifetime prolongation and mitigate the decreased degree and decreasing trend of the distance between the sensing nodes and center nodes of the sensor network, which results in a longer network life cycle and stronger environment suitability. It is an effective method that improves the sensing node lifetime for a wireless sensor network applied to cropland cultivation of ginseng.
Human beings have the ability to quickly recognize novel concepts with the help of scene semantics. This kind of ability is meaningful and full of challenge for the field of machine learning. At present, object recognition methods based on deep learning have achieved excellent results with the use of large-scale labeled data. However, the data scarcity of novel objects significantly affects the performance of these recognition methods. In this work, we investigated utilizing knowledge reasoning with visual information in the training of a novel object detector. We trained a detector to project the image representations of objects into an embedding space. Knowledge subgraphs were extracted to describe the semantic relation of the specified visual scenes. The spatial relationship, function relationship, and the attribute description were defined to realize the reasoning of novel classes. The designed few-shot detector, named KR-FSD, is robust and stable to the variation of shots of novel objects, and it also has advantages when detecting objects in a complex environment due to the flexible extensibility of KGs. Experiments on VOC and COCO datasets showed that the performance of the detector was increased significantly when the novel class was strongly associated with some of the base classes, due to the better knowledge propagation between the novel class and the related groups of classes.