Dysphonia encompasses a broad spectrum of vocal disorders with diverse etiologies, among which adductor spasmodic dysphonia (ADSD) and primary muscle tension dysphonia (pMTD) are particularly challenging to diagnose. Currently, the primary diagnostic method relies on subjective auditory perception by highly experienced clinicians. To alleviate the scarcity of diagnostic resources, this study develops a deep learning-based approach for automatically diagnosing ADSD and pMTD using patients' speech data. Our contributions are: (1) designing a convolutional neural network (CNN)-based diagnostic model that leverages handcrafted features derived from expert knowledge and (2) incorporating self-supervised learning (SSL) to extract more discriminative representations as input from raw waveforms adaptively. This marks the first application of deep learning techniques to ADSD and pMTD diagnostic modeling, achieving a classification accuracy of 83.3% on our newly constructed dataset.
We have developed an innovative speech enhancement (SE) model backbone that utilizes cross-attention among spectrum, waveform and self-supervised learned representations (CA-SW-SSL) to integrate knowledge from diverse feature domains. The CA-SW-SSL model integrates the cross spectrum and waveform attention (CSWA) model to connect the spectrum and waveform branches, along with a dual-path cross-attention module to select outputs from different layers of the self-supervised learning (SSL) model. To handle the increased complexity of SSL integration, we introduce a bidirectional knowledge distillation (BiKD) framework for model compression. The proposed adaptive layered distance measure (ALDM) maximizes the Gaussian likelihood between clean and enhanced multi-level SSL features during the backward knowledge distillation (BKD) process. Meanwhile, in the forward process, the CA-SW-SSL model acts as a teacher, using the novel teacher–student Barlow Twins (TSBT) loss to guide the training of the CSWA student models, including both lite and tiny versions. Experiments on the DNS-Challenge and Voicebank+Demand datasets demonstrate that the CSWA-Lite+BiKD model outperforms existing joint spectrum-waveform methods and surpasses the state-of-the-art on the DNS-Challenge non-blind test set with half the computational load. Further, the CA-SW-SSL+BiKD model outperforms all CSWA models and current SSL-based methods.
Depression is a major psychological disorder with a growing impact worldwide. Traditional methods for detecting the risk of depression, predominantly reliant on psychiatric evaluations and self-assessment questionnaires, are often criticized for their inefficiency and lack of objectivity. Advancements in deep learning have paved the way for innovations in depression risk detection methods that fuse multimodal data. This paper introduces a novel framework, the Audio, Video, and Text Fusion-Three Branch Network (AVTF-TBN), designed to amalgamate auditory, visual, and textual cues for a comprehensive analysis of depression risk. Our approach encompasses three dedicated branches—Audio Branch, Video Branch, and Text Branch—each responsible for extracting salient features from the corresponding modality. These features are subsequently fused through a multimodal fusion (MMF) module, yielding a robust feature vector that feeds into a predictive modeling layer. To further our research, we devised an emotion elicitation paradigm based on two distinct tasks—reading and interviewing—implemented to gather a rich, sensor-based depression risk detection dataset. The sensory equipment, such as cameras, captures subtle facial expressions and vocal characteristics essential for our analysis. The research thoroughly investigates the data generated by varying emotional stimuli and evaluates the contribution of different tasks to emotion evocation. During the experiment, the AVTF-TBN model has the best performance when the data from the two tasks are simultaneously used for detection, where the F1 Score is 0.78, Precision is 0.76, and Recall is 0.81. Our experimental results confirm the validity of the paradigm and demonstrate the efficacy of the AVTF-TBN model in detecting depression risk, showcasing the crucial role of sensor-based data in mental health detection.
Depression is a prevalent mental health problem across the globe, presenting significant social and economic challenges. Early detection and treatment are pivotal in reducing these impacts and improving patient outcomes. Traditional diagnostic methods largely rely on subjective assessments by psychiatrists, underscoring the importance of developing automated and objective diagnostic tools. This paper presents IntervoxNet, a novel computeraided detection system designed specifically for analyzing interview audio. IntervoxNet incorporates a dual-modal approach, utilizing both the Audio Mel-Spectrogram Transformer (AMST) for audio processing and a hybrid model combining Bidirectional Encoder Representations from Transformers with a Convolutional Neural Network (BERT-CNN) for text analysis. Evaluated on the DAIC-WOZ database, IntervoxNet demonstrates excellent performance, achieving F1 score, recall, precision, and accuracy of 0.90, 0.92, 0.88, and 0.86 respectively, thereby surpassing existing state of the art methods. These results demonstrate IntervoxNet’s potential as a highly effective and efficient tool for rapid depression screening in interview settings.
Due to the excellent biocompatible physicochemical performance, luminogens with aggregation-induced emission (AIEgens) characteristics have played a significant role in biomedical fluorescence imaging recently. However, screening AIEgens for special applications takes a lot of time and efforts by using conventional chemical synthesis route. Fortunately, artificial intelligence techniques that could predict the properties of AIEgen molecules would be helpful and valuable for novel AIEgens design and synthesis. In this work, we applied machine learning (ML) techniques to screen AIEgens with expected excitation and emission wavelength for biomedical deep fluorescence imaging. First, a database of various AIEgens collected from the literature was established. Then, by extracting key features using molecular descriptors and training various state-of-the-art ML models, a multi-modal molecular descriptors strategy has been proposed to extract the structure-property relationships of AIEgens and predict molecular absorption and emission wavelength peaks. Compared to the first principles calculations, the proposed strategy provided greater accuracy at a lower computational cost. Finally, three newly predicted AIEgens with desired absorption and emission wavelength peaks were synthesized successfully and applied for cellular fluorescence imaging and deep penetration imaging. All the results were consistent successfully with our expectations, which demonstrated the above ML has a great potential for screening AIEgens with suitable wavelengths, which could boost the design and development of novel organic fluorescent materials.
With the rapid growth of virtual reality technology in the Metaverse, where a digital virtual world is integrated with the real world, users are exposed to a variety of content in different scenarios. Similar to the Web2 scenario, accurately profiling users in Metaverse applications is critical for improving user satisfaction. This involves collecting and inferring information such as users’ basic attributes, interests, and consumption habits. In this study, we propose a new approach for profiling users in virtual worlds based on gaze-tracking. We developed a Metaverse shopping experiment platform and recruited volunteers to record their attention time on different commodities in 3-D scenes by wearing eye-tracking devices. We treat this gaze-tracking data as implicit feedback to capture users’ interests, thereby predicting users’ attributes. We tried different ways of extracting features from collected gaze-tracking data. To our knowledge, we are the first to mine user portraits using gaze-tracking sequences. Our experimental results demonstrate that the fusion of sequence and image features significantly improves the accuracy of user gender prediction, reaching 88.89%, which is better than using sequence or image features only. These findings demonstrated the practical and effective advantages of using gaze-tracking data to build user portraits in the Metaverse.
The g-C 3 N 4 /BiOI/CdS double Z-scheme heterojunction photocatalyst with I 3 - /I - redox pairs is prepared using simple calcination, solvothermal, and solution chemical deposition methods. The photocatalyst comprised mesoporous, thin g-C 3 N 4 nanosheets loaded on flower-like microspheres of BiOI with CdS quantum dots. The g-C 3 N 4 /BiOI/CdS double Z-scheme heterojunction has abundant active sites and in situ redox I 3 - /I - mediators and shows quantum size effects, which are all conducive to enhancing the separation of photoinduced charges and increasing the photocatalytic degradation efficiency for bisphenol A, a model pollutant. Specifically, the heterojunction photocatalyst achieves a photocatalytic degradation efficiency for bisphenol A of 98.62% in 120 min and photocatalytic hydrogen production of 863.44 μmol h -1 g -1 on exposure to visible light. The excellent visible-light photocatalytic performance is as a result of the Z-scheme heterojunction, which extends absorption to the visible light region, as well as the I 3 - /I - pairs, which accelerate photoinduced charge carrier transfer and separation, thus dramatically boosting the photocatalytic performance. In addition, the key role of the charge transfer across the indirect Z-scheme heterojunction has been elucidated and the transfer mechanism is confirmed based on the detection of intermediate I 3 - ions. Thus, this study provides guidelines for the design of indirect Z-scheme heterojunction photocatalysts.
PURPOSE:This study aims to translate the English version of the Trans Woman Voice Questionnaire (TWVQ) to simplified Chinese (TWVQ-SC) and to examine its reliability and validity.METHOD:Standardized translation procedures were strictly followed for the translation of the TWVQ. Two hundred sixty trans woman and 128 cis woman subjects completed sociodemographic investigation, the TWVQ-SC, and the Voice Handicap Index-10 (VHI-10) online. Internal consistency was examined by Cronbach reliability coefficient (Cronbach α). Test-retest reliability was quantified by intraclass correlation coefficient (ICC). Content validity, structural validity, and discriminant validity were examined by expert panel's judgment, factor analysis, Spearman's rank correlation coefficient, and the Mann-Whitney U test.RESULTS:The Cronbach α of the TWVQ-SC was .969 and the ICC was .841, indicating excellent internal consistency and good test-retest reliability. The four principal factors explained 21.345%, 18.592%, 13.551%, and 12.027% of the variance respectively with the cumulative contribution rate 65.514%. There was a strong correlation between the total score of the TWVQ-SC and that of the VHI-10 (r = .858, p < .001), indicating good structural validity. The total score of the TWVQ-SC of the trans woman subjects was significantly higher than that of the cis woman subjects (z = 14.590, p < .05), indicating good discriminant validity.CONCLUSION:The TWVQ-SC exhibits overall high reliability and validity, qualified to be applied as a reliable clinical tool to evaluate trans women's voice in mainland China.
The hollow core–shell Co9S8 @ZnIn2S4/CdS nanoreactors are fabricated by growing ZnIn2S4 nanosheets and CdS quantum dots on hollow Co9S8 nanocages for photocatalytic CO2 reduction and H2 generation. Because of unique structural and compositional advantages, Co9S8 @ZnIn2S4/CdS exhibits a significant photocatalytic BPA degradation of 90.62%, CO production of 82.10 μmol g−1 h−1, and H2 evolution rate up to 1419.14 μmol g−1 h−1. Excellent photocatalytic performances are attributed to unique core-shell architecture, which can improve the scattering and refraction efficiency of incident light, and the separation and migration of photoinduced carriers. Furthermore, Co9S8 @ZnIn2S4/CdS possesses a broad absorption spectrum facilitating outstanding photothermal performance. The photocatalytic mechanism for the nanoreactor has been discussed in detail.
The glottis's morphology not only reflects vocal and respiratory information, but also plays an important role in the diagnosis of laryngeal diseases. The glottis segmentation is a primary step in computer-aided diagnostic system, however is challenging due to various shapes of glottis, low contrast with surrounding tissues, the existence of laryngeal diseases and so on. In this paper, a deep attention network based on U-Net with color normalization operation (CN-DA-Unet) is proposed to achieve an end-to-end segmentation of the glottal area for the first time. The original images are first processed by color normalization to reduce the adverse effects of low contrast and large differences in colors between different images. The normalized images are then sent to the proposed DA-Unet for feature extraction. In this network, residual structure is incorporated to extract rich features from deep neural networks. After extracting features, a feature pyramid attention (FPA) module is applied to enhance the semantic information of the glottal area. These features are up-sampled and added to the features from the corresponding encoding layer for several times to obtain the final segmented image. The proposed approach is tested on laryngeal images of an in–house dataset including images from healthy subjects and pathologic subjects. Its performance is evaluated by several reliable and popular evaluation metrics, achieving the dice coefficient of 92.9%, sensitivity of 93.5% and precision of 92.6%. These results demonstrate the effectiveness of our proposed approach and the better performance comparing with several popular networks.
Flexible acoustic sensors with high sensitivity, excellent mechanical strength, and easy integration are urgently needed for wearable electronics. MXene holds great promise as a sensing material for this application. However, low flexibility and stability limit the performance of MXene‐based composites. To alleviate the aforementioned issue, a flexible pressure sensor based on MXene/poly(3,4‐ethylenediox‐ythiophene)‐poly(styrenesulfonate) (PEDOT:PSS) is fabricated and used as an acoustic sensor inhibiting high sensitivity, fast response time (57 ms), ultra‐thin thickness (30 μm), and remarkable stability. Excellent performance enables the sensor to detect and identify weak muscle movements and skin vibrations, such as word pronunciation and carotid artery pulse. Furthermore, by combining the proposed deep learning model based on number recognition convolutional neural network (NR‐CNN), speech recognition toward different pronunciations of numbers that appear frequently in daily conversations can be realized. High recognition accuracy (91%) is achieved by training and testing the proposed NR‐CNN with large amounts of data recorded by the sensor. Results demonstrate that the flexible and wearable MXene/PEDOT:PSS acoustic sensor accelerates intelligent artificial acoustics and possesses great potential for applications involving speech recognition and health monitoring.
Antibiotics usage in animal production is considered a primary driver of the occurrence, supply and spread of antibiotic resistance genes (ARGs) in the environment. Pig farms and fish ponds are important breeding systems in food animal production. In this study, we compared and analyzed broad ARGs profiles, mobile genetic elements (MGEs) and bacterial communities in a representative pig farm and neighboring fish ponds around Poyang Lake, the largest freshwater lake in China. The factors influencing the distribution of ARGs were also explored. The results showed widespread detection of ARGs (from 57 to 110) among 283 targeted ARGs in the collected water samples. The differences in the number and relative abundance of ARGs observed from the pig farm and neighboring fish ponds revealed that ARG contamination was more serious on the pig farm than in the fish ponds and that the water treatment plant on the pig farm was not very effective. Based on the variance partition analysis (VPA), MGEs, bacterial communities and water quality indicators (WIs) codrive the relative abundance of ARGs. Based on network analysis, we found that total phosphorus and Tp614 were the most important WIs and MGEs affecting ARG abundance, respectively. Our findings provide fundamental data on farms in lakeside districts and provide insights into establishing standards for the discharge of aquaculture wastewater.
The automatic diagnosis method based on speech signal analysis is able to realize the detection and classification of pathological voices. It plays an important role in the early diagnosis and auxiliary treatment of voice pathology, which effectively relief the discomfort of patients and reduce the workload of doctors. Therefore, the automatic diagnosis method based on speech signal analysis is of great research value. Meanwhile, high accuracy, high precision and stability are the pursuit goals. In this paper, a novel computer-aided assessment based on speech signal analysis for pathological voice classification (CS-PVC) system is proposed. This model focuses on the areas with large differences between different pathological voices and healthy voices, while ignore the negative impact of insignificant information on the performance of the model. Two databases were used in the experiments, one is the Saarbruecken Voice database (SVD), and the other is the self-built Shenzhen People’s Hospital voice database (SZUPD). The pathological voice detection accuracy of the proposed system on the above two databases are 81.6% and 82.2% respectively. The experimental results show that the proposed framework is not data-dependence. In other words, it has the potential to be universally applicable in medical framework in the future.
Background Microvascular invasion (MVI) has a significant effect on the prognosis of hepatocellular carcinoma (HCC), but its preoperative identification is challenging. Radiomics features extracted from medical images, such as magnetic resonance (MR) images, can be used to predict MVI. In this study, we explored the effects of different imaging sequences, feature extraction and selection methods, and classifiers on the performance of HCC MVI predictive models. Methods After screening against the inclusion criteria, 69 patients with HCC and preoperative gadoxetic acid-enhanced MR images were enrolled. In total, 167 features were extracted from the MR images of each sequence for each patient. Experiments were designed to investigate the effects of imaging sequence, number of gray levels (Ng), quantization algorithm, feature selection method, and classifiers on the performance of radiomics biomarkers in the prediction of HCC MVI. We trained and tested these models using leave-one-out cross-validation (LOOCV). Results The radiomics model based on the images of the hepatobiliary phase (HBP) had better predictive performance than those based on the arterial phase (AP), portal venous phase (PVP), and pre-enhanced T1-weighted images [area under the receiver operating characteristic (ROC) curve (AUC) =0.792 vs. 0.641/0.634/0.620, P=0.041/0.021/0.010, respectively]. Compared with the equal-probability and Lloyd-Max algorithms, the radiomics features obtained using the Uniform quantization algorithm had a better performance (AUC =0.643/0.666 vs. 0.792, P=0.002/0.003, respectively). Among the values of 8, 16, 32, 64, and 128, the best predictive performance was achieved when the Ng was 64 (AUC =0.792 vs. 0.584/0.697/0.677/0.734, P<0.001/P=0.039/0.001/0.137, respectively). We used a two-stage feature selection method which combined the least absolute shrinkage and selection operator (LASSO) and recursive feature elimination (RFE) gradient boosting decision tree (GBDT), which achieved better stability than and outperformed LASSO, minimum redundancy maximum relevance (mRMR), and support vector machine (SVM)-RFE (stability =0.967 vs. 0.837/0.623/0.390, respectively; AUC =0.850 vs. 0.792/0.713/0.699, P=0.142/0.007/0.003, respectively). The model based on the radiomics features of HBP images using the GBDT classifier showed a better performance for the preoperative prediction of MVI compared with logistic regression (LR), SVM, and random forest (RF) classifiers (AUC =0.895 vs. 0.850/0.834/0.884, P=0.558/0.229/0.058, respectively). With the optimal combination of these factors, we established the best model, which had an AUC of 0.895, accuracy of 87.0%, specificity of 82.5%, and sensitivity of 93.1%. Conclusions Imaging sequences, feature extraction and selection methods, and classifiers can have a considerable effect on the predictive performance of radiomics models for HCC MVI.
Medical image synthesis receives much popularity in recent years, and ample medical images can be synthesized by diverse deep learning models to alleviate the problem of lack of data in many medical imaging utilizations. However, most medical image synthesis methods still incorporate the well-known pooling operation in their convolutional neural networks-based / generative adversarial networks-based models, from which image details will be inevitably lost due to the pooling operation. In order to tackle the above problem, improved capsule-based networks, in which no pooling operation is executed and spatial details of images can be effectively preserved thanks to the equivariance characteristics of capsule models, are proposed in this paper to synthesize arterial spin labeling images, for the first time. Technically, three important issues in constructing improved capsule-based networks, including the depth of basic convolutions, the layer of capsules, and the capacity of capsules, are thoroughly investigated. Comprehensive experiments made up of region-based / voxel-based partial volume corrections and dementia diseases diagnosis based on two different datasets are conducted. The superiority of improved capsule-based networks introduced in this paper is substantiated from the statistical point of view.
Wearable sound detectors require strain sensors that are stretchable, sensitive, and capable of adhering conformably to the skin, and toward this end, 2D materials hold great promise. However, the vibration of vocal cords and muscle contraction are complex and changeable, which can compromise the sensing performance of devices. By combining deep learning and 2D MXenes, an MXene-based sound detector is prepared successfully with improved recognition and sensitive response to pressure and vibration, which facilitate the production of a high-recognition and resolution sound detector. By training and testing the deep learning network model with large amounts of data obtained by the MXene-based sound detector, the long vowels and short vowels of human pronunciation are successfully recognized. The proposed scheme accelerates the application of artificial throat devices in biomedical fields and opens up practical applications in voice control, motion monitoring, and many other fields.
Acoustic devices are widely applied in telephone communication, human-computer voice interaction systems, medical ultrasound examination, and other applications. However, traditional acoustic devices are hard to integrate into a flexible system and therefore it is necessary to fabricate light weight and flexible acoustic devices for audible sound generation and detection. Recent advances in acoustic devices have greatly overcome the limitations of conventional acoustic sensors in terms of sensitivity, tunability, photostability, and in vivo applicability by employing nanomaterials. In this review, light weight and flexible nanomaterial-enabled acoustic devices (NEADs) including sound generators and sound detectors are covered. Additionally, the fundamental concepts of acoustic as well as the working principle of the NEAD are introduced in detail. Also, the structures of future acoustic devices, such as flexible earphones and microphones, are forecasted. Further exploration of flexible acoustic devices is a key priority and will have a great impact on the advancement of intelligent robot-human interaction and flexible electronics.
Laryngeal tumor is a typical head and neck disease that may be cancerous, causing harm to human health. Automatic laryngeal tumor detection in laryngeal endoscopic images is beneficial to the further analysis of tumor characteristics to aid in treatment, such as computer assisted surgery. However, there have been very few attempts to automatically detect the tumor of larynx. In this paper, three commonly used object detection models based on convolutional neural networks (CNNs) are used to achieve automatic detection of the tumors on a dataset of laryngeal endoscopic images. As far as we know, this is the first time that object detection models based on CNNs have been used to create end-to-end detection of laryngeal tumors. The experimental results show that all of the three methods have good performances and single shot multibox detector (SSD) is more suitable for laryngeal tumor detection in terms of our evaluation metrics.