dysfunction is a common complication following thyroid surgery. However, the application of explainable machine learning for predicting postoperative voice recovery remains largely unexplored. Therefore, an investigation was done to examine voice recovery based on acoustic, objective, and glottal features. Voice recordings were collected from female patients before surgery and one month after surgery. Acoustic and glottal and others, were automatically extracted from the recordings. sion with Sequential Feature Selection were applied to examine model behavior and identify feature importance. Model stability and interpretability were evaluated across cross-validation folds. Performance metrics varied over folds, highlighting the exploratory and statistically fragile nature of predictions in small variability in feature contributions, emphasizing the need for cautious interpretation and detailed methodological reporting. Our findings provide preliminary guidance for applying explainable machine learning to small biomedical datasets. They demonstrate the importance of careful methodological design.
Detecting offensive speech poses a challenge due to the absence of a universally accepted definition delineating its boundaries. However, the scarcity of labeled data often poses a significant challenge for training robust offensive speech detection models. In this paper, we propose an approach to handle data scarcity through data augmentation techniques tailored for offensive speech detection tasks. By augmenting the existing labeled data with speech samples generated through noise injection, our method effectively expands the training dataset, enabling more comprehensive model training. We evaluate our approach on Vera Am Mittag (VAM) corpus and demonstrate significant improvements in offensive speech detection performance compared to that without data augmentation. Our findings highlight the efficacy of data augmentation in mitigating data scarcity challenges and enhancing the reliability of offensive speech detection systems in a real-world scenario.
OBJECTIVE:Vocal assessment post thyroidectomy is mainly based on laryngeal endoscopy. The purpose of our study is to assess vocal outcomes following thyroidectomy objectively based on the visual analysis of the glottal area and on acoustic parameters. STUDY DESIGN:Prospective analytical study. METHODS:We included patients who underwent thyroidectomy between September 2021 and February 2022. We performed vocal assessment preoperatively and postoperatively on Day 1 and Month 1 based on acoustic and visual glottal area analysis. The analyzed video-endoscopic glottal features included Maximal and Minimal glottal areas (MaxGA and MinGA), Amplitude Ratio (Ramp), and Phase Symmetry Index (PSI). The analyzed acoustic parameters were fundamental frequency (F0), Jitter, shimmer, harmonic-to-noise ratio (HNR), and glottal-to-noise excitation (GNE). RESULTS:We included 51 patients. At Day 1, we noted vocal fold palsy in nine patients. It persisted in three patients at Month1. MaxGA decreased significantly at Day 1 and Month 1 expressing the limitation of vocal folds (VFs) opening. PSI increased significantly on Day 1 and recovered at Month 1. Ramp decreased significantly on Day 1 then recovered partially at Month 1. F0 decreased significantly on Day 1 without a significant change of the other acoustic parameters. MaxGA correlated significantly with shimmer, GNE, and HNR. CONCLUSION:In our study, thyroid surgery resulted in MaxGA decrease reflecting a dysfunction of VFs abduction. It also resulted in transient asymmetrical VFs movements without causing acoustic and vocal alteration. Glottal area analysis can present an accurate modality to assess VFs movements post thyroidectomy even without associated clinical manifestations.
This paper aims to explore the usefulness on hidden Markovian model for cough detection from abdominal and thoracic respiratory signals. Time and frequency domains based features are extracted and relevant ones are identified. Hidden Markovian Model (HMM) parameters (initial, transition and emission probabilities) are estimated. Classification is carried in parallel with frame duration optimization. Compared to some conventional machine learning methods, competitive performance were obtained when considering tradeoff between evaluation metrics.
This research studies the effects of mind wandering (introspection) and attentional control which alternate between focused attention and unfocused mental states during a task of force production. More specifically, our aim is to identify features which allow investigating the mechanical force modifications which result from i) an isometric force task performed at three level of force (low, medium and high), ii) a mental training of attentional control (using endogenous and exogenous attention)iii) repetition of the task over time allowing to modulate the concentration. Behavior (Reaction time) and biomechanical (normalized force peak) indices are measured and selected as features in our study. We conclude: 1) the high force level, close to maximum voluntary force is better reproduced, 2) the low force level is sensitive to stimulating conditions 3) the medium force level is difficult to master but mental training reinforced it.
Offensive speech covers many different ways someone may harm another such as humiliation, criticism, yelling, or verbal abuse. Such behavior results in negative consequences with possible dangerous impacts like biasing people’s thoughts, spreading racism, discrimination and even physical violence among citizens. While offensive speech takes place online and offline and in all forms of social interaction, our work contributes at filtering the alarming diffusion of such behaviors in televisiondebate programs by proposing a speech-based offensive behavior detection system. To do so, we investigated the use of the Vera Am Mittag (VAM) corpus which is a collection of recordings taken from a German TV talk show. From these, a mixture of Mel Frequency Cepstral Coefficients (MFCC) and Stationary Wavelet Transform (SWT) based features was extracted and several feature selection techniques were applied for capturing the most relevant features. Finally, both of K-Nearest Neighbors (KNN) and deep learning Convolutional Neural Networks (CNNs) were employed for classification. Results highlight that the best performance is associated with CNNs whith feature selection, reaching a classification accuracy of $97.21 \%$.
ObjectivesVocal dysfunction is a frequent complication following thyroidectomy that can be associated with a negative impact on patients’ quality of life. Although the effect of thyroidectomy on acoustic features has been widely studied, the examination of glottal flow characteristics to assess vocal outcomes following thyroid gland surgery has not been included in empirical research, to date. The goal of our study was to evaluate early and short-term vocal outcomes following thyroidectomy based on the analysis of glottal acoustic features during speech production.Study designProspective analytical study.Materials and methodsWe evaluated vocal outcomes in patients who underwent thyroidectomy between September 2021 and March 2022. We extracted glottal flow features from their vocal recordings preceding surgery and postoperatively at Day1 and Month1 postoperatively. The extraction of glottal features was performed using a signal processing-based approach. We extracted the following features: Open quotients (OQ1 and OQ2), Quasi-open quotient (QoQ), Closing quotient (ClQ), Amplitude quotient (AQ), Normalized Amplitude quotient (NAQ) and Speed quotients (SQ1 and SQ2).We included 39 patients. OQ2 and QoQ decreased significantly at Day1 and Month1. OQ1 and NAQ decreased significantly at Month1. ClQ remained stable at both postoperative assessments. AQ decrease was not significant at both dates. SQ1 increased at Day1 and Month1 but the change was not significant. SQ2 decreased significantly at both Day1 and Month1. OQ, QoQ, AQ, NAQ, and SQ2 did not recover at Month1. We noted that the decrease of SQ1 and SQ2 correlated significantly with the increase of the Voice Handicap Index-10 (VHI-10) at Month1.ConclusionThe analysis of glottal acoustic features can be a reliable modality to detect vocal changes following thyroidectomy. Thyroidectomy was associated with a vocal dysfunction that was manifested by the decrease of open, amplitude, and speed quotients. Glottal features can present a potential tool to objectively assess the effect of thyroidectomy on vocal folds movements.
The goal of this study is to contribute on the development of human-machine interface that helps to identify individuals with suspicious behavior (of violence, terrorism,…) from speech.To do it, it was neceassry to define suspicious behavior through emotions according to valence and activation. Next, a scenario including background noise, security agent speech and vistor speech was defined. A dataset was constructed, speech processing and machine learning based algorithms are designed: after features extraction, audio stream is segmented into relevant content (noise, agent speech and visitor speech). The visitor speech is selected and a classification is carried to detect suspicious behavior. Simulation results show the effectiveness of the proposed solution.
While temporal preparation has frequently been examined through the manipulation of foreperiods, the role of force level during temporal preparation remains underexplored. In our study, we propose to manipulate mental training of attentional control in order to shed light on the role of the force level and autonomic nervous system in the temporal preparation of an action. Forty subjects, divided into mental training group (n = 20) and without mental training group (n = 20), participated in this study. The influence of the attentional control and force levels on the autonomic nervous system were measured using the skin conductance response and the heart rate variability; the accuracy of the motor responses was measured using a method derived from machine learning. Behaviorally, only the mental training group reinforced its motor and attentional control. When using short foreperiod durations and high force level, motor and attentional control decreased, consistent with the dominant sympathetic system. This resulted in an increased anticipation rate of responses with a higher reaction time compared to the long foreperiods duration and low force level, in which the reaction time significantly decreased, with enhancement of the expected force level, showing consistency with the dominant parasympathetic system. Interestingly, results revealed a predictive relationship between the sympathovagal balance and motor and attentional control during the long foreperiods and low force level. Finally, results demonstrate that attentional mental training leads to the reinforcement of interactions between the autonomic nervous system and attentional processes which are involved in the temporal preparation of a force task.
The main goal of this research is to develop a a machine learning based method in order to detect cough from acceleration signals. In this study, two different methods are proposed: a conventional one that uses XGBoost as a classifier and a deep learning which uses CNN-1D as an architecture. We found that these models were able to distinguish between acceleration signals caused by coughing and acceleration signals caused by other activities such as clearing throat, talking, laughing and movements in different directions with high accuracy. This study affirms that cough monitoring based on accelerometer measurements generated by the Hexoskin device is possible, making it a new user-independent tool for cough detection.
Searchable abstracts of presentations at key conferences in endocrinology ISSN 1470-3947 (print) | ISSN 1479-6848 (online)
Searchable abstracts of presentations at key conferences in endocrinology ISSN 1470-3947 (print) | ISSN 1479-6848 (online)
After thyroid surgery, some patients complain of voice disorders, recurrent laryngeal nerve palsy, hoarseness, and breathing problems, which can significantly impact their quality of life. In this paper, we sought to assess and identify the impact of thyroidectomy on voice quality by following some patients before and after surgery. To accomplish this, we collected a dataset of voice recordings from patients one day before, one day after, and one month after thyroid surgery. Then, we extracted acoustic features related to harmonicity and pitch variations in amplitude and duration and used statistical tests to study the discrimination abilities of these features in preoperative and postoperative. The results revealed that the (mean, max, min, and median) fundamental frequency values were not significantly decreased in the first postoperative stage. However, they changed significantly in the second postoperative stage. For females, the Max (F0) was not changed after one day and one month. The other parameters did not differ between the preoperative and postoperative. The analysis and results clarify that thyroidectomy can affect voice quality for up to one month. Further studies are needed to fully understand the factors behind these alterations.
OBJECTIVES:Voice changes are a common complication after a thyroidectomy, which is a surgical procedure involving partial or total removal of the thyroid gland. The main objective of this work is to examine the possible voice disorders after thyroid surgery. More precisely, it is an investigation of partial and total thyroidectomy, as well as the effects that cancerous and noncancerous thyroid glands can have regarding postsurgical vocal and their association with age and gender. METHODS:Patients were evaluated using acoustic voice parameters, including harmonics-to-noise ratio (HNR), fundamental frequency (F0), jitter, speaker phonation frequency (SPF) range, cepstral peak prominence (CPP), maximum phonational frequency range (MPFR), and shimmer at the preoperative stage and postoperatively at the 1 day, and first-month stages. RESULTS:Results demonstrated a significant change in F0 parameters, SPF range, and CPP feature 1 month after surgery, depending on the type of thyroidectomy and thyroid pathology. No significant changes were observed in the HNR, shimmer, and jitter features. Age was associated with the CPP parameter in the entire sample. In contrast, the MPFR parameter was also related to the type of thyroidectomy in the entire sample. However, maximum F0 was significantly associated with the type of thyroidectomy, specifically in the female sample. CONCLUSIONS:Results indicated that a thyroidectomy can have a negative impact on voice quality. The age and type of thyroidectomy performed are not responsible for this change. Potentially this change can be due to factors such as nerve damage or the subjects' experience, such as job, anxiety, and their physical condition, as well as treatments they may have undergone before thyroidectomy. Further efforts are needed to fully understand the background of voice changes after thyroidectomy.
Identification of voice disorders plays a major role in our life nowadays. In this context, voice analysis can be used, as a complementary technique with other traditional invasive methods, such as laryngoscopy, to identify voice disorders. This paper explores a set of acoustic features extracted from the vocal folds signals namely the pitch, to this purpose. First, 49 pitch-based features are extracted from speech recordings. Then, relevant ones are selected according to their discriminatory power between normal and pathological voice classes. To do so, different feature selection techniques have been used and compared. Afterward, KNN, SVM, Random Tree and Naive Bayes classifiers are applied to decide on the existence of a voice disorder or not and which pathological class is detected. The experimental results denote the usefulness of retained features and despite, the simplicity of classification technique (KNN for instance), the best performance in term of accuracy rate reached 91.5%.
Research on psychological comfort and serenity in social media becomes a necessity because of the excess of negative waves generated by the users. In this paper, we aim to distinguish between offensive and ordinary speech based on machine learning classification techniques. For this purpose, the VAM emotional audio database has been restructured and adjusted to the context of offensive speech detection. Besides, a feature fusion based on Mel Frequency Cepstral Coefficients (MFCCs) as well as Stationary Wavelet Transform (SWT) has been employed and K-nearest neighbors (KNN) algorithm has been used as a classification tool. Results show that the considered feature set has relatively great power in recognizing suspicious behavior, reaching 96.2% as the highest accuracy rate.
This paper has two objectives: the first is to generate two binary flags to indicate useful frames permitting the measurement of cardiac and respiratory rates from Ballistocardiogram (BCG) signals-in fact, human body activities during measurements can disturb the BCG signal content, leading to difficulties in vital sign measurement; the second objective is to achieve refined BCG signal segmentation according to these activities. The proposed framework makes use of two approaches: an unsupervised classification based on the Gaussian Mixture Model (GMM) and a supervised classification based on K-Nearest Neighbors (KNN). Both of these approaches consider two spectral features, namely the Spectral Flatness Measure (SFM) and Spectral Centroid (SC), determined during the feature extraction step. Unsupervised classification is used to explore the content of the BCG signals, justifying the existence of different classes and permitting the definition of useful hyper-parameters for effective segmentation. In contrast, the considered supervised classification approach aims to determine if the BCG signal content allows the measurement of the heart rate (HR) and the respiratory rate (RR) or not. Furthermore, two levels of supervised classification are used to classify human-body activities into many realistic classes from the BCG signal (e.g., coughing, holding breath, air expiration, movement, et al.). The first one considers frame-by-frame classification, while the second one, aiming to boost the segmentation performance, transforms the frame-by-frame SFM and SC features into temporal series which track the temporal variation of the measures of the BCG signal. The proposed approach constitutes a novelty in this field and represents a powerful method to segment BCG signals according to human body activities, resulting in an accuracy of 94.6%.
Automatic deception detection is an important task that has gained a huge interest in different fields due to its potential applications. Particularly, it can improve justice and security in society by helping in detecting deceivers in high-stakes situations across jurisprudence, law enforcement, and national security domains, among others. However, the existing deception detection systems until today are not as accurate as it is expected, which makes their use very risky especially in critical fields. This article outlines an approach for automatically distinguishing between deceit and truth based on audio, video and text modalities and explores the possibility of combining them together in order to detect deception more accurately. First, each modality has been evaluated separately and then a feature and decision-level fusion approaches have been proposed to combine the considered modalities. The proposed feature level fusion approach investigates a diversified feature selection techniques to select the most relevant ones among the whole used feature set, while the decision level fusion approach is based on the belief theory considering information about the certainty degree of each modality. To do so, we used a real-life video dataset of people communicating truthfully or deceptively collected from public american court trials. Unimodal models trained on audio, video and text separately achieved an accuracy rate of 60%, 94% and 58% respectively. When using the feature level fusion approach, the best accuracy deception detection result reaches 93% using only 19 combined features, while a 100% deception recognition rate has been obtained with the decision-level fusion proposed approach, outperforming the results obtained in the literature.
Amyotrophic Lateral Sclerosis (ALS) is a specific disease that causes the death of neurons controlling voluntary muscles. The death affects Upper Motor Neurons (UMN) and Lower Motor Neurons (LMN). In this paper, we aim to identify UMN, LMN ALS patients and healthy subjects by studying their Electromyography (EMG) signals. More specifically, the Right Anterior Tibialis (RAT) muscle, responsible of dorsiflexing and inverting the foot during walking is studied. The solution is mainly composed of two blocks: features extraction and classification. A large feature vector, combining statistics, time, frequency and time-frequency domains features is extracted and their relevance is checked. Then, many classification techniques are tested in order to identify the most powerful ones. Finally, a dimensionality reduction, using both Principal Component Analysis and Sequential Forward Selection technique is carried out in order to select the most relevant features. Simulation results achieved 97% of overall accuracy for healthy/LMN/UMN separation and more than 99% for healthy/ALS and LMN/UMN classification.
The present study sought to evaluate how mental effort modulates premotor activity within forearm muscles in the context of an isometric grasping task. Muscle activity of the flexor digitorum superficialis (FDS) and extensor digitorum communis (EDC) was recorded during the application of maximum grip forces in nineteen healthy adult subjects. Each subject was examined under two experimental conditions: 1) spontaneous initiation of grasp (SI) and 2) focused concentration preceding the initiation of grasp (CA). Two novel parameters, the mean premotor duration (MPD) and the mean premotor power (MPP) were used to distinguish patterns of muscle activity. Here we tested the hypothesis was maximal grip strength is primed by muscle activity during the premotor phase. Our results demonstrate that MPD for each muscle group was significantly longer in the CA condition than for the SI condition (BF10 = 491497) and that MPP was significantly greater in EDC than in FDS (BF10 = 4305). Furthermore, both the MPD and MPP of the EDC were significantly correlated with maximum grip force. These results suggest that the increase of premotor activity consequent to the mental effort (focused concentration) may support internal biomechanical and physiological mechanisms which serve to enhance patterns of neuromuscular synergies.
Bessam Abdulrazak合作论文数Université de Sherbrooke4