This paper presents the system submitted to Track 1 of the Voice Timbre Attribute Detection (vTAD) 2025 Challenge. The core objective of the vTAD challenge is to address the intensity comparison task, which requires determining the relative strength of timbre attributes between two speech signals in dimensions of human perception. The system utilizes pre-trained speaker representations and gender representations as front-end inputs, and employs a residual neural network to output the intensity comparison results of speech pairs under specific descriptors. The system ultimately secured third place on the Seen track of the vTAD 2025 Challenge, achieving an accuracy of 95. 38
Fully few-shot class-incremental audio classification (FFCAC) addresses real-world scenarios where new classes have only limited data. The Expandable Dual-embedding Extractor (EDE) is an effective baseline for this task, but it requires progressively expanding embeddings as new classes arrive, leading to increased training and inference complexity. To overcome this limitation, we propose Adaptive Embedding Fusion EDE (AEF-EDE), which prevents embedding expansion and improves model efficiency, while further leveraging contrastive learning to achieve significant performance gains. Experiments on LS-100, NSynth-100, and FSC-89, along with ablation studies, demonstrate the effectiveness of our approach.
Internet of Things (IoT) systems often generate time series from multiple related entities, creating a practical need for unified cross-entity forecasting. Although these entities may operate within the same system, they can still exhibit long-term structural differences and short-term temporal variations under different operating regimes, which makes cross-entity forecasting particularly challenging. To address this issue, we propose an Entity- and Regime-conditioned Model (ERM) for multi-entity time series forecasting. Built on a shared forecasting backbone, ERM introduces two complementary conditioning mechanisms: Regime-MoE captures sample-level regime variations, and Entity-FiLM models long-term entity-specific structural shifts through entity-conditioned residual feature modulation. On ME-BTS, a reconstructed multi-entity HVAC benchmark derived from real building management system data, we conduct systematic experiments under source-entity joint training, unseen-entity zero-shot forecasting, and limited target-entity adaptation settings. The results show that ERM delivers stronger unseen-entity generalization and more stable performance across entities, while remaining highly competitive after adaptation with limited target-entity data. These findings suggest that combining shared temporal modeling with regime- and entity-conditioned mechanisms is an effective direction for multi-entity time series forecasting.
Target Speaker Extraction (TSE) uses a reference cue to extract the target speech from a mixture. In TSE systems relying on audio cues, the speaker embedding from the enrolled speech is crucial to performance. However, these embeddings may suffer from speaker identity confusion. Unlike previous studies that focus on improving speaker embedding extraction, we improve TSE performance from the perspective of speaker consistency. In this paper, we propose a speaker consistency-aware target speaker extraction method that incorporates a centroid-based speaker consistency loss. This approach enhances TSE performance by ensuring speaker consistency between the enrolled and extracted speech. In addition, we integrate conditional loss suppression into the training process. The experimental results validate the effectiveness of our proposed methods in advancing the TSE performance. A speech demo is available online:https://sc-tse.netlify.app/
Acoustic Impulse Response (AIR) provides crucial spatial information about the environment, significantly enhancing audio immersion. However, achieving high perceptual quality while computing AIR in real-time for interactive audio-video media (IAVM) presents a challenging problem. This study proposes the Mesh to Parametric AIR (M2PAIR), a method for computing AIR designed for IAVM. M2PAIR integrates neural networks with psychoacoustics. It takes the 3D scene mesh, the listener positions, and the sound source positions as inputs, utilizes perceptual parameters as intermediaries, and computes the desired high-quality AIR signal based on these parameters. Experimental results demonstrate that M2PAIR improves the perceptual quality of AIR output compared to existing methods while reducing the model complexity. Additionally, it meets the requirements of IAVM, including real-time computation, high sampling rates, and flexible duration for the output AIR.
The state-of-the-art (SOTA) Automatic Speech Recognition (ASR) systems are mostly based on the data-driven methods. However, low-resource languages may lack data for training. Articulatory Features (AFs) describe the movements of the vocal organ which can be shared across languages. Thus, this paper investigates AFs-based semi-supervised techniques to share data between languages. First, the traditional acoustic features and the AFs are combined as front-end features to provide articulatory information for cross-lingual knowledge transfer. Then, the dropout-based lattice decoded are used as the pseudo-labels for the unsupervised data to address the problem of data deficiency. In addition, the Lattice-free Maximum Mutual Information (LF-MMI) objective is adopted to better adapt to small datasets. Experiments show that our system can obtain a relative improvement of 58.6
Binaural audio delivers an immersive spatial auditory experience to human listeners, but most existing videos lack binaural audio due to the expertise required for recording environments. Recent studies have been dedicated to converting monaural audio into binaural ones conditioned on the visual inputs. In this paper, we propose a novel audio-visual spatialization network with two added audio decoders, which rely on carefully designed visual features to generate audio outputs for the left and right channels, respectively. In addition, we propose an audio-visual matching loss to further explore the correlation between binaural audio and the scene visual input. Experiment results show that the proposed method outperforms several state-of-the-art binaural audio generation methods on two benchmark datasets FAIR-Play and MUSIC-Stereo. Qualitative results are also presented to demonstrate the effectiveness of the proposed method.
Recently, a novel task called audio-visual segmentation (AVS) has emerged, focusing on pixel-wise segmentation of sounding objects in videos. This task is particularly challenging as it involves segmenting individual pixels based on objects in video frames accompanied by sound. We propose a Motion Based Audio-Visual Segmentation model, which incorporates optical flow maps with motion information into the AVS task for the first time. The Motion-Vision Attention Module (MVA) is proposed to facilitate the fusion of motion and visual features to exploit motion information. Additionally, the Cross-Modal Bilateral-Attention Module (CMBA) is introduced to integrate multimodal features through crossmodal attention. The proposed model is evaluated on two distinct datasets, S4 and MS3, the outperformance of which demonstrates its effectiveness and feasibility in addressing the AVS task.
Sonar images are often affected by reverberation, posing challenges to detecting and recognizing target signals. This paper combines the strong interpretability and controllability of signal processing model with the powerful data mining capability of deep learning technology to propose a novel sonar image dereverberation method. Initially, an unsupervised deep learning, guided by domain knowledge from Independent Component Analysis (ICA), is employed as the first-level network. The first-level called ICANET, enables end-to-end dereverberation without labeled data. Subsequently, the supervised deep learning network called SKRNET (Supervised Knowledge-based Reverberation Network), further enhances performance by incorporating labeled data, along with knowledge obtained from ICANET. Sea trial experiments demonstrate that compared to benchmark methods, SKRNET achieves the highest signal-to-reverberation ratios (SRR) at 35.76 and entropy at 7.39, resulting in improvements of 20.2% and 4.8% over the raw data, respectively. Simulation experiments also confirm the superior signal restoration capability of the proposed method.
As machine learning technologies continue to ad-vance and natural language processing systems become in-creasingly sophisticated, conversational systems are becoming more prevalent across various domains. Simultaneously, the improvement of healthcare systems has made it possible to develop specialized medical task-oriented conversational systems. Against this backdrop, this project builds a big data-driven conversational generation system focused on disease queries, it is based on a Chinese healthcare database and uses the Rasa framework. The system aims to extract information required for constructing a medical conversational system from big data, covering queries related to disease basic descriptions, causes, symptoms, recommended foods, dietary restrictions, treatment methods, medications, examination items, susceptible populations, and preventive measures. The system's effectiveness is validated using a combined subjective and objective evaluation approach, offering users a personalized online querying platform. By leveraging big data to construct a medical conversational generation system, this project elevates the system's intelligence level and establishes a paradigm for constructing conversational systems using big data. It aims to provide users with more comprehensive and precise medical services.
To address the increasing demand for interviews, we have proposed an intelligent interview system. This system has the capability to generate interview questions and design interview processes. This system introduces a ChatGPT-based iterative enhancement algorithm for seed datasets to create a high-quality prompt-question dataset specific to the interview domain. Through large-scale language model fine-tuning, the system can generate more accurate and expected interview questions, demonstrating significant performance improvements across multiple evaluation metrics. Additionally, the study introduces an innovative multi-agent interaction algorithm to enhance interview process efficiency and comprehensive information gathering. Experimental results indicate improvements in text coherence and contextual understanding with the fine-tuned model, and subjective testing confirms the system's fluency and practicality.
With the rapid development of deep learning techniques, the applications have become increasingly widespread in various domains. However, traditional deep learning methods are often referred to as "black box" models with low interpretability of their results, posing challenges for their application in certain critical domains. In this study, we propose a comprehensive method for the interpretability analysis of sentiment models. The proposed method encompasses two main aspects: attention-based analysis and external knowledge integration. First, we train the model within sentiment classification and generation tasks to capture attention scores from multiple perspectives. This multi-angle approach reduces bias and provides a more comprehensive understanding of the underlying sentiment. Second, we incorporate an external knowledge base to improve evidence extraction. By leveraging character scores, we retrieve complete sentiment evidence phrases, addressing the challenge of incomplete evidence extraction in Chinese texts. Experimental results on a sentiment interpretability evaluation dataset demonstrate the effectiveness of our method. We observe a notable increase in accuracy by 1.3%, Macro-F1 by 13%, and MAP by 23%. Overall, our approach offers a robust solution for enhancing the interpretability of sentiment models by combining attention-based analysis and the integration of external knowledge.
Unsupervised Anomalous Sound Detection (ASD) aims to design a generalizable method that can be used to detect anomalies when only normal sounds are given. In this paper, Anomalous Sound Detection based on Diffusion Models (ASD-Diffusion) is proposed for ASD in real-world factories. In our pipeline, the anomalies in acoustic features are reconstructed from their noisy corrupted features into their approximate normal pattern. Secondly, a post-processing anomalies filter algorithm is proposed to detect anomalies that exhibit significant deviation from the original input after reconstruction. Furthermore, denoising diffusion implicit model is introduced to accelerate the inference speed by a longer sampling interval of the denoising process. The proposed method is innovative in the application of diffusion models as a new scheme. Experimental results on the development set of DCASE 2023 challenge task 2 outperform the baseline by 7.75%, demonstrating the effectiveness of the proposed method.
With the growing significance of non-intrusive speech quality assessment in speech systems, existing methods predominantly rely on neural networks to extract low-order features. Typically, these features undergo a low-dimensional linear transformation, yielding the network’s output. However, the intercorrelation between feature points is often overlooked. In this paper, we explore the concept of kernel method, which maps features into high dimensional space through dot product, in order to enhance the extraction of relationships among all feature points. Considering the unique advantages of tensors in complex data representation, we extend the utilization of tensor network and propose a novel framework that incorporates a matrix product state (MPS) layer to predict mean opinion score (MOS). By integrating the MPS layer, our model can transform low-order features into higher-order representations, facilitating linear transformation in a high dimensional space without increasing the number of parameters. Furthermore, we propose a loss function that concurrently assesses regression and classification biases, along with correlation with real MOS labels. Experimental results demonstrate that our proposed model consistently outperforms the baseline system across all evaluation metrics and surpasses state-of-the-art models on the test set.
It is a common method to quantify the latent in audio generation and then use diffusion models to estimate noise or data from the corrupted data to generate the quantized latent. Unlike the method, we consider that the targets estimated by the diffusion model include both noise and data, rather than just one of them. Based on this idea and multi-task learning methods, we design the network Dual-Unet, which is simply modified by U-net and can estimate both noise and data simultaneously. Combining Dual-Unet and Variational AutoEncoders with Residual Vector Quantizer, we propose Multitask diffusion model(MTDiffusion), which can generate foley sound audio with a given label. We validate our proposed model on the DCASE task7B dataset. The experimental results show the effectiveness of our proposed model, and both subjective and objective metrics of the generated audio significantly exceed the baseline and the first place ranked on 2023 DCASE task7B.
The audio codec is one of the core modules in audio communication for real-time transmission. With the development of neural networks, end-to-end audio codecs have emerged and demonstrated effects beyond conventional codecs. However, current neural network-based codecs have the weakness of high computational complexity, and the performance of these methods decreases rapidly after decreasing the complexity, which is not conducive to deployment under low computational resources. In this paper, a low-complexity audio codec is proposed. To realize the low complexity of the model with high quality, a structure based on frequency band division is designed, which is implemented using a within bandacross band interaction (WBABI) module to learn the features across and within the subband. Further, we propose a new quantization-compensation module, which reduces the quantization error by 90%. The experimental results show that for audio with a sample rate of 24kHz, the model shows excellent performance at 3~6kbps compared to other codecs, and the complexity is only 0.8 Giga Multiply-Add Operations per Second(GMACs).
The goal of unsupervised anomalous sound detection (ASD) for industrial machines is to identify anomalous sounds using only normal sounds for training. A common method is to use attribute information as auxiliary labels to train ASD models. However, this approach often faces the challenge of obtaining auxiliary attribute information in complex industrial environments. This paper proposes a novel unsupervised anomalous sound detection method that fine-tunes pre-trained models without using attribute information, achieved by training a machine type classifier. The proposed method is called Annotation-Free Fine-tuning (AFF) for unsupervised anomalous sound detection. In addition, we propose an anomaly score calculation method that combines the machine type classifier with an unsupervised anomalous sound estimator, further improving the anomalous detection performance of AFF. Experiments on DCASE 2024 Task 2 development dataset indicate that our method outperforms other typical ASD methods that do not utilize attribute information.
In speaker verification, performance degradation caused by domain mismatch has been a common problem as the test domain lies outside the training distribution. In this paper, we present a novel domain transfer network called Adversarial Diffusion Probabilistic Model (ADPM), to better alleviate this problem. More specifically, ADPM is used to transfer melspec-trogram from the source domain into the target domain. To generate the melspectrogram, we propose to regard the diffusion model as the generator and a discriminator is employed for adversarial training. We also explore the contrastive learning objective to retain the context information of source domain. The generated and the original feature maps from the source domain are fed into the ResNet34 network jointly to construct cross-domain speaker verification. We evaluate the proposed techniques on VOiCES dataset, and our best model achieves a relative 8.94% Equal Error Rate (EER) drop compared to the previous adaption methods.
This article studies consumer online review behaviors based on recommendation reward programs. It mainly studies the influence of different types of rewards in the recommendation reward plan on consumers’ online review behavior. Through the questionnaire survey, the results show that consumers are more inclined to choose monetary rewards as their return in the recommendation reward program. At the same time, different rewards in the recommendation reward plan, consumer product satisfaction, and the interaction of the two also have a certain impact on consumers’ online review behavior. This article divides the amount of rewards into high rewards and low rewards. When the consumer’s satisfaction reaches a certain level, low rewards in the referral reward program will bring more benefits to the merchants than high rewards. High rewards will only have a greater effect on consumers who are less satisfied with the product. When consumer product satisfaction is not high, and the reward limit set by the recommendation reward plan is low, consumers’ perception of fairness will cause them not to conduct online evaluation behaviors, even if the reward limit is high, the effect is not significant.
In advanced transportation-management systems, variable speed limits are a crucial application. Deep reinforcement learning methods have been shown to have superior performance in many applications, as they are an effective approach to learning environment dynamics for decision-making and control. However, they face two significant difficulties in traffic-control applications: reward engineering with delayed reward and brittle convergence properties with gradient descent. To address these challenges, evolutionary strategies are well suited as a class of black-box optimization techniques inspired by natural evolution. Additionally, the traditional deep reinforcement learning framework struggles to handle the delayed reward setting. This paper proposes a novel approach using covariance matrix adaptation evolution strategy (CMA-ES), a gradient-free global optimization method, to handle the task of multi-lane differential variable speed limit control. The proposed method uses a deep-learning-based method to dynamically learn optimal and distinct speed limits among lanes. The parameters of the neural network are sampled using a multivariate normal distribution, and the dependencies between the variables are represented by a covariance matrix that is optimized dynamically by CMA-ES based on the freeway's throughput. The proposed approach is tested on a freeway with simulated recurrent bottlenecks, and the experimental results show that it outperforms deep reinforcement learning-based approaches, traditional evolutionary search methods, and the no-control scenario. Our proposed method demonstrates a 23% improvement in average travel time and an average of a 4% improvement in CO, HC, and NOx emission.Furthermore, the proposed method produces explainable speed limits and has desirable generalization power.