Early detection of Alzheimer's Dementia (AD) and Mild Cognitive Impairment (MCI) is critical for timely intervention, yet current diagnostic approaches remain resource-intensive and invasive. Speech, encompassing both acoustic and linguistic dimensions, offers a promising non-invasive biomarker for cognitive decline. In this study, we present a machine learning framework for the PROCESS Challenge, leveraging both audio embeddings and linguistic features derived from spontaneous speech recordings. Audio representations were extracted using Whisper embeddings from the Cookie Theft description task, while linguistic features-spanning pronoun usage, syntactic complexity, filler words, and clause structure-were obtained from transcriptions across Semantic Fluency, Phonemic Fluency, and Cookie Theft picture description. Classification models aimed to distinguish between Healthy Controls (HC), MCI, and AD participants, while regression models predicted Mini-Mental State Examination (MMSE) scores. Results demonstrated that voted ensemble models trained on concatenated linguistic features achieved the best classification performance (F1 = 0.497), while Whisper embedding-based ensemble regressors yielded the lowest MMSE prediction error (RMSE = 2.843). Comparative evaluation within the PROCESS Challenge placed our models among the top submissions in regression task, and mid-range for classification, highlighting the complementary strengths of linguistic and audio embeddings. These findings reinforce the potential of multimodal speech-based approaches for scalable, non-invasive cognitive assessment and underline the importance of integrating task-specific linguistic and acoustic markers in dementia detection.
Assessing blending of instruments is important in music performance and perception research, but remains underexplored due to its complex multi-dimensional nature. Despite extensive research on source-level blending, the influence of room acoustics on this process is rarely examined. This study proposes a computational modelling approach to evaluate the perceived overall blending between instruments examining the blending at the source-level and its alteration brought by room acoustics. Three audio stimuli, each showcasing different degrees of source-level blending between two violins, were auralized in 25 simulated room acoustic environments, with expert listeners assessing their overall perceived blending. The correlation analysis of room acoustic parameters revealed that their influence on overall blending is contingent upon source-level blending. A random forest regression model is proposed to predict perceived overall blending ratings using source-level blending ratings and room acoustic parameters. Its viability was confirmed through twofold evaluation, including Leave-one-out-cross-validation and separate training and test data, with a mean absolute error of 6% in each case. Feature importance analysis revealed that source-level blending contributes 60%, while room acoustics contribute 40% of the overall perceived blending ratings, with perceived reverberance being the primary contributor. Overall, this investigation contributes to a more holistic understanding of blending perception.
Accurate placement of nasogastric (NG) tubes is crucial to ensure effective treatment and avoid potential complications. Conventional methods of confirming NG tube placement, such as auscultation and pH testing, can be prone to errors. This study explores the use of auscultation signals in conjunction with machine learning acoustic feature modelling to verify NG tube placement. Study participants were recruited from hospital inpatients who had NG tubes inserted for relevant medical indications, and underwent radiographic imaging for confirmation of NG tube placement. Eligible patients were recruited into the study in either the gastric (NG tube in stomach) or esophageal (NG tube removed with tip lying in esophagus at 30cm mark) phase, or both. Audio signals were obtained from two locations: left chest and epigastrium, during insufflation of 50 millilitres of air for each phase. Building upon on a pre-trained model, we extracted acoustic signatures and transfer-learned VGGish pretrained network model to predict the NG tube placement. A total of 103 audio file pairs (left chest and epigastrium) were obtained from 61 patients, of which 74 audio pairs (44 gastric and 30 esophageal phase) were utilised for training and validation of the prediction model. An accuracy of 89.3
We present a method for reconstructing the gold-standard electrocardiogram (ECG) from the more easily obtainable phonocardiogram (PCG). Both ECG and PCG are commonly employed in the primary diagnosis of various cardiovascular diseases (CVD). By leveraging a Bidirectional Long Short-Term Memory (BiLSTM) network capturing bidirectional temporal dependencies in ECG-PCG pairs, our approach effectively reconstructs key morphological features of the ECG, evidenced by strong performance metrics in healthy subjects. Although handling diverse pathological ECG patterns remains challenging, the model successfully reconstructs essential ECG structures in most cases; noise in PCG signal was found to potentially compromise reconstruction accuracy. Importantly, the combination of ECG and PCG data offers a more complementary perspective, aiding informed decision-making for a variety of CVDs. This transformation opens new clinical avenues offering a simple and convenient approach for obtaining critical ECG data with a smartphone's onboard microphone, significantly enhancing diagnostic capabilities and patient care 'in the field'.
This study evaluates the use of machine learning, specifically the Random Forest Classifier, to differentiate normal and pathological swallowing sounds. Employing a commercially available wearable stethoscope, we recorded swallows from both healthy adults and patients with dysphagia. The analysis revealed statistically significant differences in acoustic features, such as spectral crest, and zero-crossing rate between normal and pathological swallows, while no discriminating differences were demonstrated between different fluid and diet consistencies. The system demonstrated fair sensitivity (mean +/- SD: 74% +/- 8%) and specificity (89% +/- 6%) for dysphagic swallows. The model attained an overall accuracy of 83% +/- 3%, and F1 score of 78% +/- 5%. These results demonstrate that machine learning can be a valuable tool in non-invasive dysphagia assessment, although challenges such as sampling rate limitations and variability in sensitivity and specificity in discriminating between normal and pathological sounds are noted. The study underscores the need for further research to optimize these techniques for clinical use.
Large language models (LLMs) find increasing applications in many fields. Here, three LLM chatbots (ChatGPT-3.5, ChatGPT-4, and Bard) are assessed in their current form, as publicly available, for their ability to recognize Alzheimer’s dementia (AD) and Cognitively Normal (CN) individuals using textual input derived from spontaneous speech recordings. A zero-shot learning approach is used at two levels of independent queries, with the second query (chain-of-thought prompting) eliciting more detailed information than the first. Each LLM chatbot’s performance is evaluated on the prediction generated in terms of accuracy, sensitivity, specificity, precision, and F1 score. LLM chatbots generated a three-class outcome (“AD”, “CN”, or “Unsure”). When positively identifying AD, Bard produced the highest true-positives (89% recall) and highest F1 score (71%), but tended to misidentify CN as AD, with high confidence (low “Unsure” rates); for positively identifying CN, GPT-4 resulted in the highest true-negatives at 56% and highest F1 score (62%), adopting a diplomatic stance (moderate “Unsure” rates). Overall, the three LLM chatbots can identify AD vs. CN, surpassing chance-levels, but do not currently satisfy the requirements for clinical application.
Globally prevalent, cardiovascular diseases (CVD) result in high annual mortality rates. Heart auscultation, though accessible and non-invasive, is $a$ difficult skill to acquire for accurate CVD diagnosis, and requires years of extensive training. To tackle this challenge, our study explores automatic heart sound classification to assist in the early detection of CVD. Our study collected cardiac sounds from 20 healthy individuals and 30 individuals with pathological heart conditions, compiling the dataset within a clinical environment using a smartphone. Employing transfer learning, we adapted the VGGish model for binary classification (healthy vs. pathological), achieving an accuracy of 95.0%. Feature extraction involved Mel spectrograms processed by a 26-layer VGGish model. The Grad-CAM method further enhanced interpretability by highlighting frequency regions deemed influential in the decision-making process.
Acoustic simulation tools, although common in the acoustic evaluation of buildings with large atrium space, are costly both in terms of time and resources. This study aims to develop an efficient prediction tool for Reverberation Time (RT) and Sound Pressure Level (SPL) within the atriums using deep learning, with Building Information Modelling (BIM) files as the only input. Initially, a 3-D acoustic simulation model was benchmarked against experimental measurements of RT for an existing building’s atrium, demonstrating an agreement within 15%. Subsequently, BIM files of 60 buildings were used to simulate RT and SPL at various listener locations within their atrium spaces. The BIM files provided essential geometry information, including atrium shape, dimensions, door placements, floor plans, material properties, and sound sources. This dataset was then utilized to train deep learning algorithms, enabling rapid and convenient predictions of RT and SPL for any new building's atrium space. The proposed prediction model serves as a valuable tool for architectural planning and noise regulation during the early design stages. By relying solely on basic building information from the BIM file, this tool obviates the need for time-consuming and computationally expensive simulation software typically used for acoustic evaluations in large atrium spaces.
Automated techniques to detect Alzheimer’s Dementia through the use of audio recordings of spontaneous speech are now available with varying degrees of reliability. Here, we present a systematic comparison across different modalities, granularities and machine learning models to guide in choosing the most effective tools. Specifically, we present a multi-modal approach (audio and text) for the automatic detection of Alzheimer’s Dementia from recordings of spontaneous speech. Sixteen features, including four feature extraction methods (Energy–Time plots, Keg of Text Analytics, Keg of Text Analytics-Extended and Speech to Silence ratio) not previously applied in this context were tested to determine their relative performance. These features encompass two modalities (audio vs. text) at two resolution scales (frame-level vs. file-level). We compared the accuracy resulting from these features and found that text-based classification outperformed audio-based classification with the best performance attaining 88.7%, surpassing other reports to-date relying on the same dataset. For text-based classification in particular, the best file-level feature performed 9.8% better than the frame-level feature. However, when comparing audio-based classification, the best frame-level feature performed 1.4% better than the best file-level feature. This multi-modal multi-model comparison at high- and low-resolution offers insights into which approach is most efficacious, depending on the sampling context. Such a comparison of the accuracy of Alzheimer’s Dementia classification using both frame-level and file-level granularities on audio and text modalities of different machine learning models on the same dataset has not been previously addressed. We also demonstrate that the subject’s speech captured in short time frames and their dynamics may contain enough inherent information to indicate the presence of dementia. Overall, such a systematic analysis facilitates the identification of Alzheimer’s Dementia quickly and non-invasively, potentially leading to more timely interventions and improved patient outcomes.
Broadband excitation introduced at the speaker's lips and the evaluation of its corresponding relative acoustic impedance spectrum allow for fast, accurate and non-invasive estimations of vocal tract resonances during speech and singing. However, due to radiation impedance interactions at the lips at low frequencies, it is challenging to make reliable measurements of resonances lower than 500 Hz due to poor signal to noise ratios, limiting investigations of the first vocal tract resonance using such a method. In this paper, various physical configurations which may optimize the acoustic coupling between transducers and the vocal tract are investigated and the practical arrangement which yields the optimal vocal tract resonance detection sensitivity at low frequencies is identified. To support the investigation, two quantitative analysis methods are proposed to facilitate comparison of the sensitivity and quality of resonances identified. Accordingly, the optimal configuration identified has better acoustic coupling and low-frequency response compared with existing arrangements and is shown to reliably detect resonances down to 350 Hz (and possibly lower), thereby allowing the first resonance of a wide range of vowel articulations to be estimated with confidence.
Quantifying auditory perception of blending between sound sources is a relevant topic in music perception, but remains poorly explored due to its complex and multidimensional nature. Previous studies were able to explain the source-level blending in musically constrained sound samples, but comprehensive modelling of blending perception that involves musically realistic samples was beyond their scope. Combining the methods of Music Information Retrieval (MIR) and Machine Learning (ML), this investigation attempts to classify sound samples from real musical scenarios having different musical excerpts according to their overall source-level blending impression. Monophonically rendered samples of 2 violins in unison, extracted from in-situ close-mic recordings of ensemble performance, were perceptually evaluated and labeled into blended and non-blended classes by a group of expert listeners. Mel Frequency Cepstral Coefficients (MFCCs) were extracted, and a classification model was developed using linear and non-linear feature transformation techniques adapted from the dimensionality reduction strategies such as Principal Component Analysis (PCA), Linear Discriminant Analysis (LDA), and t-Stochastic Neighbourhood Embedding (t-SNE), paired with Euclidean distance measure as a metric to evaluate the similarity of transformed feature clusters. Results showed that LDA transformed raw MFCCs trained and validated using a separate train-test data set and Leave-One-Out Cross-Validation (LOOCV) resulted in an accuracy of 87.5%, and 87.1% respectively in correctly classifying the samples into blended and non-blended classes. In this regard, the proposed classification model which incorporates “ecological” score-independent sound samples without requiring access to individual source recordings advances the holistic modeling of blending.
In this study, we develop a noise annoyance prediction tool using deep learning in Singapore, a densely populated city-state where a significant portion of the population resides in public housing near various noise sources like MRT, highways, bus routes, and construction sites. To investigate short-term annoyance caused by surrounding noises, we created an easily accessible web-based subjective listening test. We created a noise dataset featuring typical Singaporean noise sources, including traffic (buses, highways), trains, aviation, neighbourhood activities (playgrounds, schools, hawker centers, funerals, wildlife, home renovations), and construction. Participants were exposed to 35 noise stimuli, lasting between 15 to 30 seconds, and asked to rate perceived annoyance on a 5-point scale, with 5 being the most annoying. Using these reported annoyance levels, we ranked the stimuli relative to each other and categorized them into three classes: low, medium, and high noise annoyance. Using these three labels, Long-Short Term Memory (LSTM) networks were trained to predict the perceived annoyance of new audio samples. Such a perceived annoyance assessment tool for new audio samples can help in addressing specific noise concerns of residents, leading to more effective and targeted noise control strategies, and consequently creating a quieter and more pleasant urban living environment.
Nanofiber-porous systems comprising a porous substrate overlaid with nanofiber weave offer the potential for higher acoustic absorption than the substrate alone with negligible increase in thickness. The characterization of nanofibers from acoustic measurements is investigated in this work, and a regression model for predicting their acoustic properties from a single physical parameter is proposed to enable the design of nanofiber-porous systems directly from fabrication parameters. Characterization as a resistive screen via Johnson-Champoux-Allard and lumped element models for transfer matrix computations of absorption coefficient for nanofiber-porous systems exhibited good agreement with the measured spectra. The lumped element model was chosen as it was defined by fewer parameters and did not require nanofiber layer thickness measurements, eliminating the associated uncertainty. A regression model for lumped element parameters vs areal density established a design tool based on a single, easily measured physical property for optimized absorption at target frequencies without prior acoustic characterization of the nanofiber layer, enabling the analysis of complex acoustic networks incorporating nanofiber-porous systems. Practical considerations of applying adhesives at the nanofiber-porous interface were studied to evaluate possible enhancement of acoustic performance. For comparison with prior work by others, flow resistances from physical measurement and acoustic characterization were compared.
A reconfigurable fully integrated automatic gain control (AGC) amplifier is presented which is based on a dual mode continuous gain adjustment variable-gain amplifier (VGA). The VGA is realized based on the current-steering structure, achieving a linear-in-dB gain tuned by AGC’s feedback voltage. A current division active load is proposed, which provides additional gain control range and realizes gain robustness to process and supply variations without extra calibration. Simulated with Cadence IC, the maximum gain variations due to process and supply variations are 3.5 dB and 2.6 dB, respectively. By using dual mode gain tuning technique, the VGA achieves a total gain range of 68.2 dB. Both bandwidth and gain of the VGA are adjusted independently. The AGC is realized in UMC 55-nm CMOS technology with 0.026 mm 2 core area. With the current division ratio K of 0.5, the proposed VGA achieves a linear-in-dB gain of 42.2 dB (−30.0 to 12.2 dB) with less than 0.79 dB error. The −3 dB bandwidth can be adjusted from 80 MHz to 140 MHz and is insensitive to gain variations. The power dissipation is from 2.9 mW to 4.5 mW. This design features bandwidth scalability and the gain independent of bandwidth to achieve various applications.
A data-driven approach using artificial neural networks is proposed to address the classic inverse area function problem, i.e., to determine the vocal tract geometry (modelled as a tube of nonuniform cylindrical cross-sections) from the vocal tract acoustic impedance spectrum. The predicted cylindrical radii and the actual radii were found to have high correlation in the three- and four-cylinder model (Pearson coefficient (ρ) and Lin concordance coefficient (ρc) exceeded 95%); however, for the six-cylinder model, the correlation was low (ρ around 75% and ρc around 69%). Upon standardizing the impedance value, the correlation improved significantly for all cases (ρ and ρc exceeded 90%).
In this article, we demonstrate the use of a simple pendulum to explore the concepts of kinematics and dynamics. A simple homemade pendulum and a phone-based accelerometer are used to determine, at various points in time, the acceleration of a moving train. The dynamical and kinematics data from the homemade pendulum and the accelerator can then be compared, and this in turn highlights the relationship between dynamical and kinematics quantities. This project is originally part of a coursework requirement assigned to the first four authors for their introductory classical mechanics course at a university in the summer of 2020. Similar ideas have been explored earlier independent of this work, but this work will focus on the detailed description of our approach and using updated technologies as our benchmark. The project is one that can be implemented readily as an activity in high school physics classes.
This database presents acoustic impedance spectra measured on a G-key Dízi (Chinese transverse flute): a total of 35 fingerings (both standard and alternative fingerings) for 24 semi-chromatic notes ranging from D5 to G7 (i.e. the instrument’s typical playing range), for both with the membrane attached and without. Accompanying each fingering entry is the sound recording and corresponding audio spectra. An interactive graphical version of this database is also presented at https://acoustics.sutd.edu.sg/dizi-impedance/ The impedance measurements (using the “three-microphone two-calibration” technique) were conducted at the Acoustics Lab, School of Physics, UNSW, Sydney Australia. The authors are grateful to John Smith and Joe Wolfe for hosting us. The acoustic impedance measurement hardware and software was developed by Paul Dickens and we used a version of the software subsequently modified by Noel Hanna.
This paper puts forth a method that is easily transposable to a realtime environment by utilising a "Physics Approximating Neural Network" to predict 1D output signal. The technique described in this paper is inspired by Physics Informed Neural Networks put forth by Raissi et al. The model demonstrated in this paper makes use of a recurrent input. This is passed to a "preprocessing" stage of 2 layers by 8 neurons wide. The output of the preprocessing stage is passed to the "approximation layer". Lastly, the output of the "approximation layer" is passed to a "postprocessing" layer of 5 layers by 8 neurons wide. The architecture of the model is explained and tested on a specially developed dataset. The results show closer similarity to the ground truth compared to a linear model or a dense-layer neural network.
We investigated the relationship between optical BCG and ECG signals measured simultaneously for the same heartbeat cycles. Despite the long history of BCG, earlier studies compared BCG and ECG features across large time cycles (inter-heartbeat), but not within the heartbeat cycle (intra-heartbeat). The non-invasively derived BCG signal was found to have a remarkable relationship with the arterial pressure signal, which has not been previously reported. We achieved synchronization of the two disparate modalities to within an estimated uncertainty of 50-70 ms, which allowed us to compare features within the heart cycle (which may be related to the arterial pressure) for one pathological and four healthy subjects lying supine, and found it consistent regardless of their breathing condition, gender and health status. Although not a one-to-one correlation, we show optical BCG proves to be a convenient and an unobtrusive and complementary modality to monitor cardiac activity alongside the well-established ECG.
BACKGROUND:Uroflowmetry remains an important tool for the assessment of patients with lower urinary tract symptoms (LUTS), but accuracy can be limited by within-subject variation of urinary flow rates. Voiding acoustics appear to correlate well with conventional uroflowmetry and show promise as a convenient home-based alternative for the monitoring of urinary flows. OBJECTIVE:To evaluate the ability of a sound-based deep learning algorithm (Audioflow) to predict uroflowmetry parameters and identify abnormal urinary flow patterns. DESIGN, SETTING, AND PARTICIPANTS:In this prospective open-label study, 534 male participants recruited at Singapore General Hospital between December 1, 2017 and July 1, 2019 voided into a uroflowmetry machine, and voiding acoustics were recorded using a smartphone in close proximity. The Audioflow algorithm consisted of two models-the first model for the prediction of flow parameters including maximum flow rate (Qmax), average flow rate (Qave), and voided volume (VV) was trained and validated using leave-one-out cross-validation procedures; the second model for discrimination of normal and abnormal urinary flows was trained based on a reference standard created by three senior urologists. OUTCOME MEASUREMENTS AND STATISTICAL ANALYSIS:Lin's correlation coefficient was used to evaluate the agreement between Audioflow predictions and conventional uroflowmetry for Qmax, Qave, and VV. Accuracy of the Audioflow algorithm in the identification of abnormal urinary flows was assessed with sensitivity analyses and the area under the receiver operating curve (AUC); this algorithm was compared with an external panel of graders comprising six urology residents/general practitioners who separately graded flow patterns in the validation dataset. RESULTS AND LIMITATIONS:A total of 331 patients were included for analysis. Agreement between Audioflow and conventional uroflowmetry for Qmax, Qave, and VV was 0.77 (95% confidence interval [CI], 0.72-0.80), 0.85 (95% CI, 0.82-0.88) and 0.84 (95% CI, 0.80-0.87), respectively. For the identification of abnormal flows, Audioflow achieved a high rate of agreement of 83.8% (95% CI, 77.5-90.1%) with the reference standard, and was comparable with an external panel of six residents/general practitioners. AUC was 0.892 (95% CI, 0.834-0.951), with high sensitivity of 87.3% (95% CI, 76.8-93.7%) and specificity of 77.5% (95% CI, 61.1-88.6%). CONCLUSIONS:The results of this study suggest that a deep learning algorithm can predict uroflowmetry parameters and identify abnormal urinary voids based on voiding sounds, and shows promise as a simple home-based alternative to uroflowmetry in the management of patients with LUTS. PATIENT SUMMARY:In this study, we trained a deep learning-based algorithm to measure urinary flow rates and identify abnormal flow patterns based on voiding sounds. This may provide a convenient, home-based alternative to conventional uroflowmetry for the assessment and monitoring of patients with lower urinary tract symptoms.