
In this study, we introduce ExpressiveSinger, an end-to-end ex-pressive singing voices synthesis model, which accurately re-flect users' musical expression by analyzing real-played MIDI sequences and lyrics. We propose a novel method to auto-matically annotate velocity labels for MIDI sequences in SVS datasets, as these sequences do not inherently contain velocity information compared to real-played MIDI sequences. More-over, we separately model expressive features and modify the vocoder to enhance controllability and quality of the synthetic singing voices. Finally, we adopt a soft-vc like approach for end-to-end training to effectively preserve more linguistic content features. Our experiments on the professional Mandarin singing corpus validate our data annotation method and demonstrate the effectiveness of ExpressiveSinger in terms of naturalness and a strong correlation between the synthetic singing voice and the MIDI input.
Auditory attention represents a key neural process in the brain's handling of auditory information, essential for effective human communication. Electroencephalography (EEG) has proven to be effective in auditory attention decoding (AAD). However, current methods primarily concentrate on individual subject settings, which fail to achieve robust models across subjects due to notable variations between subjects and noisy labels. To overcome these issues, this paper investigates a novel domain adaptation method using prototypical representation(DApro) to extract both specific and invariant representations across subjects. Specifically, we align the features from the source and target domains to a shared feature space, while considering the relationships with prototypical representations both within and between classes. Features of the same category are clustered together, while those of different categories are separated. The proposed method was evaluated on the publicly available AHU dataset using a leave-one-subject-out cross-validation protocol. The experimental results demonstrate the effectiveness and the superiority of our model in terms of generalization and robustness in the cross-subject auditory attention classification using EEG signals.
The exponential growth of academic papers presents a huge challenge for researchers, exacerbating the already tedious literature review process. Current tools like Google Scholar and Connected Papers offer solutions for text-based and citation-based searches but fail to address the need for finding semantically similar yet terminologically different papers efficiently. This paper proposes an innovative approach to paper discovery using semantic search to create a knowledge graph of topics and papers. By generating a tree of topics using ChatGPT 4o and calculating semantic similarity with SciBERT, this method aims to uncover relevant papers overlooked by traditional citation-based searches. The solution, validated through quantitative evaluation, demonstrates the potential to improve the efficiency and comprehensiveness of paper discovery.
Mandarin seven-sound test was introduced to bridge a frequency gap from 6000 Hz to 8000 Hz identified in the standard Mandarin version of the Ling six-sound test. However, until the present research, there are few relevant studies to confirm the feasibility of this test or to investigate its potential for use across various pronunciation communities. In this study, eight speakers were invited to record the audio material for Mandarin seven-sound test. Subsequently, 31 adult participants took the Mandarin seven-sound test using these recordings. The results showed that individuals with hearing loss performed significantly poorer than those with normal hearing in terms of both reaction time and accuracy rate. Specifically, within the hearing-loss group, the accuracy rates for the /s/ and /(sic)/ sounds were notably lower compared to other phonemes. Additionally, this study found that the demographic characteristics of the speakers, such as gender and age, did not significantly influence the accuracy rate for either group. These results indicate that Mandarin seven-sound test not only effectively addresses the frequency gap of the Ling six-sound test but also maintains a high level of consistency across different speakers. Moreover, the online testing system has proven to be strong user-friendliness and provides significant value for clinical speech audiometry, speech therapy, and the development of hearing aid technology.
In addition to affecting the segmental levels, speech rate can also influence suprasegmental acoustic features, such as rhythm. As an essential element of prosody, the rhythmic properties of Mandarin at different speech rates have been underexplored. This study aims to explore the impact of speech rate on the rhythmic characteristics of Mandarin Chinese in semi-spontaneous speech. Using speech materials from the CCTV news broadcast, the study employs a linear mixed-effects model to compare the differences in rhythmic patterns between normal and slow speech rates. The results indicate that speech rate has a statistically significant effect on Mandarin rhythm, both for vowels and consonants, with the values of rhythmic parameters being significantly higher under slow speech conditions compared to normal speech conditions. It suggests that Mandarin Chinese exhibits greater intra-syllabic and inter-syllabic variability in rhythm at slower speech rates. Reduced speech rate may enhance the expressiveness and complexity of Mandarin rhythm. This finding highlights the importance of considering speech rate in studies of Mandarin prosody.
Data augmentation is recognized as a powerful technique for enhancing the robustness of speaker verification systems. However, this method can lead to discrepancies between the training and inference phases in the feature domain. A common solution is to simply combine the original and augmented data for the network optimization. Unfortunately, this approach can not fully utilize the information across different feature domains. In this paper, we introduce a novel Dynamic Cross Triplet (DCT) method, and this method initially measures intra-domain and inter-domain similarities using speaker embeddings, and then transfers them across domains to address the mismatch issue. Notably, training with DCT is both efficient and effective, making it applicable to various networks. Experimental results on the benchmark Voxceleb dataset demonstrate the superiority of the proposed DCT method.
Speaker diarization is the task of determining “who spoke when” in an audio recording, which is of practical importance in various multi-talker scenarios such as inquiries of patients and meetings. Current speaker diarization methods are usually supervised clustering algorithms based on artificial neural networks, which usually fail to tackle an unlimited number of speakers. In this paper, we proposed Dynamic Linear, which can automatically add output nodes and generate new vectors. By switching the output layer in Discriminative Neural Clustering (DNC) to dynamic linear, we proposed a speaker diarization method based on Transformer, which can tackle an unlimited number of speakers. We also modified the training strategy, loss function and decode method to facilitate our proposed method. Experiments show that our proposed speaker diarization method not only enables DNC to handle an unlimited number of speakers, but also achieves a 1.83% performance improvement on the AMI dataset relative to baseline method DNC. In addition, dynamic linear can substitute the output linear layer in any neural clustering algorithm, enabling it to deal with an unlimited number of clusters, closer to general clustering algorithms.
Adapting an existing well-trained system to a new domain using only unlabeled data is a highly sought-after yet challenging task for speaker verification in real-world scenarios. In this paper, we study two different domain adaptation methods, the adversarial domain adaptation (ADA) and the self-supervised learning-based domain adaptation (SSDA). To facilitate the deployment of unsupervised adaptation methods in applications, we conduct a detailed analysis of the characteristics of both the ADA and SSDA adaptation strategies. Our findings indicate that the SSDA strategy's performance is highly influenced by the amount of target domain data, whereas the ADA strategy is relatively insensitive to data quantity. Furthermore, augmenting target domain data enhances SSDA system performance but diminishes ADA performance. To further enhance system performance, we explore the complementarity between ADA and SSDA. Our results demonstrate that ADA and SSDA complement each other. When both strategies are applied jointly, the best system achieves over 20.0% relative Equal Error Rate (EER) improvement on the Cnceleb evaluation set and over 35.0% relative average EER improvement on the SRE16 Cantonese and Tagalog evaluation set under domain mismatched conditions.
The perturbation effect of fundamental frequency at vowel onset (i.e., onset f0) caused by the voicing distinction of preceding consonants was extensively explored across languages (hereafter as “CF0”, consonant-related f0 perturbations). The aspiration contrast mainly distinguished by vot is an important distinctive feature in Mandarin, which divides stops and affricates into two categories: aspirated and unaspirated. Although CF0 due to the aspiration contrast has been demonstrated in Mandarin, the role of f0 in the perception of aspiration contrast has received limited attention. The present study focused on the relative contribution of vot and onset f0 to the perception of aspirated consonants. A series of synthesized syllables that were orthogonally covarying in vot and onset f0 were created by manipulating a set of minimal pair that were identical in all features except for the aspiration. The results based on mixed-effects logistic regression model showed that the beta coefficients indicating the relative cue weighting of acoustic cues, were 3.41 for vot and 0.42 for onset f0, which suggested that although onset f0 contributed to the perception of aspirated sounds, vot was still a robust determinant. Moreover, the weighting of onset f0 was associated with the transition duration. When the onset f0 changed linearly from the onset to the steady-state f0 of voicing over the first 50ms, the weighting of onset f0 increased from 0.42 to 0.74.
With the widespread adoption of internet technology and the rapid development of information technology, an increasing number of viewpoints are being conveyed on online platforms, making emotion recognition of these viewpoints very important. Information is typically conveyed in multimodal forms, such as text, audio, and video modalities. Combining information from these modalities can enhance emotion recognition. This study integrates information from two modalities: the BERT model and the RoBERTa model are introduced for text modality to extract features, and the data2vec2.0 model is introduced for audio modality to extract features, incorporating contextual information from previous dialogues for training. All experiments in this paper are conducted on the IEMOCAP dataset, using a five-fold cross-validation method. The final multimodal emotion recognition achieves a weighted accuracy of 78.49%, an unweighted accuracy of 79.38%, and an F1-score of 78.43%. The models and code can be found on https://github.com/xiazhengshun/multimodal-emotion-recognition.
Non-intrusive audio quality assessment, particularly for subjective MOS prediction for music signal, is crucial in real-time audio communication and playback systems. While network-based methods have been extensively used for objective speech quality assessment, evaluating audio quality presents a greater challenge due to higher sampling rates and more complex signal spectrum. In this paper, we design a non-intrusive audio quality assessment system based on deep neural network for subjective MOS prediction of distorted audio signals. Mixed perceptual features are extracted for signal analysis, and both objective and subjective indicators are utilized as labels for two-step training on simulated data. Besides, we apply improved convolution layers, attention layers, and a type of new loss to improve the performance of our model. The experimental results show that the proposed system performs better than conventional assessment methods in correlation.
This paper introduces an expressive speech synthesis system submitted to Track 1 of ICAGC 2024. The objective of this track is to clone the voices of the target speakers using provided speech data and modulate them to convey appropriate emotions for various themes, such as novel chapters and ancient Chinese poems. Our system primarily employs a pretrained GPT-SoVITS, a two-stage large-scale speech synthesis system. In addition, we have developed a theme-oriented few-shot learning strategy tailored to specific themes. This strategy involves fine-tuning the pre-trained models with sentences spoken by different speakers but on the same theme. This approach aims to refine the models to focus on both the specific themes and individual speaker characteristics. The competition results underscore the efficacy of our approach, culminating in a fourth-place finish among all participating teams.
Early detection is crucial for timely intervention aimed at pre-venting and slowing the progression of neurocognitive disorder (NCD), a common and significant health problem among the aging population. Recent evidence has suggested that language-related functional magnetic resonance imaging (fMRI) may be a promising approach for detecting cognitive decline and early NCD. In this paper, we proposed a novel, naturalistic language-related fMRI task for this purpose. We examined the effectiveness of this task among 97 non-demented Chinese older adults from Hong Kong. The results showed that machine-learning classification models based on fMRI features extracted from the task and demographics (age, gender, and education year) achieved an average area under the curve of 0.86 when clas-sifying participants' cognitive status (labeled as NORMAL vs DECLINE based on their scores on a standard neurcognitive test). Feature localization revealed that the fMRI features most frequently selected by the data-driven approach came primarily from brain regions associated with language processing, such as the superior temporal gyrus, middle temporal gyrus, and right cerebellum. The study demonstrated the potential of the naturalistic language-related fMRI task for early detection of aging-related cognitive decline and NCD.
This study aimed to explore how listeners selectively attend to speech in a noisy environment with complex sound sources. Participants were firstly instructed to attend to a specified speaker among 5. Subsequently, they were exposed to white noise played from all speakers for 500 ms followed by a speech signal played from the target location while ignoring speech from another speaker. EEG responses were recorded and an-alyzed using Inter-Trial Phase Coherence and Event-Related Spectral Perturbation. Our results demonstrated that the theta and alpha frequencies are crucial in reflecting auditory attention allocation. Specifically, ITPC was sensitive to the left-right division of attention, whereas ERSP, particularly in alpha power, may be more effective in differentiating adjacent locations dynamically.
Early objective identification and assessment of dysarthria due to neurological deficits are essential for neurorehabilitation. Developing a system to achieve this requires a large-scale database of pathological information with detailed labeling. In the present study, a high-quality Chinese multimodal audio-visual database, consisting of 64 subacute stroke patients and 25 healthy participants, named the "Mandarin Subacute Stroke Dysarthria Multimodal (MSDM) database", was established. The materials of MSDM include a series of speech tasks such as syllables, characters, words, sentences, and spontaneous speech. All audio-visual data in this database were manually annotated and simultaneously verified by experienced researchers. Additionally, comprehensive clinical assessments of speech-motor function (e.g., Frenchay Dysarthria Assessment) and cognitive function (e.g., Montreal Cognitive Assessment) for each individual were included in the database. In conclusion, the MSDM database is believed to provide sufficient data resources for developing automatic assessment and speech recognition methods and contribute to understanding the pathological mechanisms of dysarthria.
Ultrasound technology, capable of capturing the tongue's contour, has received significant attention in speech visualization. There is an increasing interest in utilizing tongue motion data from ultrasound imaging to drive 3D tongue models. However, traditional driving methods have not fully utilized all the contour information of tongue, typically using a limited number of contour points to drive the tongue model. These approaches often lead to pathological shapes that deviate from natural speech articulation. To address this issue, we propose an innovative method that drives the tongue model by utilizing the entire tongue contour captured from ultrasound images. Initially, the complete tongue contour is extracted from the ultrasound images. Subsequently, a mapping model is developed to establish the relationship between ultrasound tongue contour and model control parameters. Finally, Root mean squared error is used to evaluate the reconstructed model control parameters, and the curve similarity index is used to assess the resemblance between the ultrasound tongue contour and the model midsagittal shape. This evaluation determines the accuracy of the driven tongue model. The results demonstrate that the reconstruction error of the control parameters is within 3%, and the average contour curve similarity between the tongue model of each phoneme and the ultrasound tongue contour is approximately 95%. These findings indicate the feasibility of driving tongue models using the entire tongue contour, effectively generating 3D tongue models that match 2D ultrasound images and avoiding issues with pathological shapes in the driven tongue model.
In recent years, speech as an easily accessible biomarker, has gained significant attention due to its numerous advantages and is currently a focus of active research and development in healthcare and medical fields. Due to the advancements in signal processing and deep learning technology, many studies have reported the effectiveness of using speech to detect clinical depression. However, these studies lack clinical interpretability, hindering the linkage of acoustic features to disease biology and reducing their utility for clinicians and patients. In this study, we proposed two explainable features common temporal pattern (CTP) and multi-frequency band Hurst (MFB-Hurst). These features are designed to measure clinically common symptoms of depression in speech, specifically reduced speech inflection, and noise during vocal tract closure in speech production. Based on the pronunciation data of 10 phrases, including sus-tained vowels and common expressions from healthy individuals and patients, CTP and MFB-Hurst achieved 70% and 64% accuracy in depression detection, respectively. The combination of CTP and MFB-Hurst achieved an average accuracy of 72%. These results demonstrate the potential of the proposed features in depression detection. Moreover, since these inter-pretable features are derived from clinical experience, they facilitate future applications in clinical settings.
In recent years, deep neural network based methods for speaker verification have made remarkable progress in clean environments. However, background noise significantly reduces the accuracy and reliability of speaker verification systems by masking or changing the voice characteristics of the speaker. In this paper, we propose a cascaded framework optimized with multiobjective loss to mitigate the interference of different levels and types of noise on the speaker verification task. The proposed architecture consists of two components: a speech enhancement module based on improved 2D-UNet, which reduces the structural limitations of directly using classical UNet for noise reduction, and a back-end speaker embedding extraction module. We carry our experiments on the VoxCeleb1 and VOiCES datasets, as well as in the presence of out-of-domain noise conditions. The evaluations have demonstrated this method shows great potential for speaker verification in noisy environments.
Large language models have achieved great success in natural language processing tasks. It has recently become a new research hot spot. For example, in tasks such as mathematical reasoning and story writing, large models have emerged with extremely strong capabilities. However, their huge size and computing requirements have brought great challenges to actual deployment. In terms of reasoning speed, as the model size increases significantly, the model's reasoning speed will drop a lot. Therefore, it is necessary to prune and accelerate large models. Existing structured and unstructured pruning methods have problems in compatibility and are not fully applicable to large models after pruning. Although these pruning methods are effective in theory, they usually show different applicability and effects when applied to complex models. For example, structured pruning methods may be more suitable for achieving model compression through sparse word embedding matrices or reducing the number of attention heads, while unstructured pruning methods focus more on pruning redundant parameter connections. However, these methods often lack sufficient compatibility and general applicability in practice. We mainly explore the research on fusion algorithms of pruning methods, including fusion acceleration solutions that combine structured pruning and unstructured pruning, as well as fusion acceleration solutions that combine pruning and other acceleration methods.
Optimization is a fundamental problem in AI research. Widely used first-order optimization methods like SGD and Adam have drawbacks, such as difficulty in hyper-parameter selection, poor performance on "ill-conditioned" loss landscapes, and bottleneck of efficiency improvement. This has led to interest in "second-order" optimization methods, which, however, face high space and time complexity due to the storage and inversion of second-order matrices. Designing practical secondorder methods to overcome these complexities is crucial. The Hessian-free method based on the conjugate gradient method (HFCG) addresses these issues by approximating the "second-order" updates iteratively. This paper extends HFCG to the Conformer model, implementing directional derivative calculations for multi-head self-attention, convolution, LayerNorm, and BatchNorm modules for the first time. A two-stage distributed training procedure (TDTP) is introduced to enhance optimization performance and reduce training time, including an improved update strategy and a new CG batch extraction method. TDTP was tested on the Penn Treebank and AMI datasets for language modelling tasks. Ablation experiments demonstrated the efficiency of HFCG and the new TDTP strategies. Additionally, the impact of the choice of HL in the generalized Gauss-Newton matrix on performance was explored and the hypothesis that HFCG might reduce performance differences caused by normalization was proposed.