Agentic systems offer a potential path to solve complex clinical tasks through collaboration among specialized agents, augmented by tool use and external knowledge bases. Nevertheless, for chest X-ray (CXR) interpretation, prevailing methods remain limited: (i) reasoning is frequently neither clinically interpretable nor aligned with guidelines, reflecting mere aggregation of tool outputs; (ii) multimodal evidence is insufficiently fused, yielding text-only rationales that are not visually grounded; and (iii) systems rarely detect or resolve cross-tool inconsistencies and provide no principled verification mechanisms. To bridge the above gaps, we present RadAgents, a multi-agent framework for CXR interpretation that couples clinical priors with task-aware multimodal reasoning. In addition, we integrate grounding and multimodal retrieval-augmentation to verify and resolve context conflicts, resulting in outputs that are more reliable, transparent, and consistent with clinical practice.
Identifying the underlying relationship between visual movements in tagged-MRI and intelligible speech is a vital problem to better understand speech production in health and disease. Due to their heterogeneous representations, however, direct mapping between the two modalities is challenging. We develop a deep learning framework that can synthesize a sequence of tagged MRI data to its corresponding mel-spectrogram, and then convert back into the audio waveform. Our network adopts a parallel encoder-decoder structure to take as input a pair of tagged MRI sequences. The 3D CNN-based encoders learn to extract the feature of spatiotemporally varying motions. The decoder then learns to generate the corresponding spectrograms conditioned on the latent space feature. For the pair of the same utterance, we further make the latent space feature as close as possible with the Kullback-Leibler divergence. To demonstrate the performance of our framework, we used a leave-one-out evaluation strategy on a total of 63 tagged MRI sequences from two utterances, including 43 “ageese” and 20 “asouk.” Our framework enabled the generation of clear audio given a sequence of tagged MRI unseen in training, which could potentially aid in better understanding speech production and improving treatment strategies for patients with speech-related disorders.
Multimodal representation learning using visual movements from cine magnetic resonance imaging (MRI) and their acoustics has shown great potential to learn shared representation and to predict one modality from another. Here, we propose a new synthesis framework to translate from cine MRI sequences to spectrograms with a limited dataset size. Our framework hinges on a novel fully convolutional heterogeneous translator, with a 3D CNN encoder for efficient sequence encoding and a 2D transpose convolution decoder. In addition, a pairwise correlation of the samples with the same speech word is utilized with a latent space representation disentanglement scheme. Furthermore, an adversarial training approach with generative adversarial networks is incorporated to provide enhanced realism on our generated spectrograms. Our experimental results, carried out with a total of 63 cine MRI sequences alongside speech acoustics, show that our framework improves synthesis accuracy, compared with competing methods. Our framework thereby has shown the potential to aid in better understanding the relationship between the two modalities.
Identifying sections is one of the critical components of understanding medical information from unstructured clinical notes and developing assistive technologies for clinical note-writing tasks. Most state-of-the-art text classification systems require thousands of in-domain text data to achieve high performance. However, collecting in-domain and recent clinical note data with section labels is challenging given the high level of privacy and sensitivity. The present paper proposes an algorithmic way to improve the task transferability of meta-learning-based text classification in order to address the issue of low-resource target data. Specifically, we explore how to make the best use of the source dataset and propose a unique task transferability measure named Normalized Negative Conditional Entropy (NNCE). Leveraging the NNCE, we develop strategies for selecting clinical categories and sections from source task data to boost cross-domain meta-learning accuracy. Experimental results show that our task selection strategies improve section classification accuracy significantly compared to meta-learning algorithms.
The glossectomy procedure, involving surgical resection of cancerous lingual tissue, has long been observed to affect speech production. This study aims to quantitatively index and compare complexity of vocal tract shaping due to lingual movement in individuals who have undergone glossectomy and typical speakers using real-time magnetic resonance imaging data and Principal Component Analysis. The data reveal that (i) the type of glossectomy undergone largely predicts the patterns in vocal tract shaping observed, (ii) gross forward and backward motion of the tongue body accounts for more change in vocal tract shaping than do subtler movements of the tongue (e.g., tongue tip constrictions) in patient data, and (iii) fewer vocal tract shaping components are required to account for the patients' speech data than typical speech data, suggesting that the patient data at hand exhibit less complex vocal tract shaping in the midsagittal plane than do the data from the typical speakers observed.
Emotional speech production has been previously studied using fleshpoint tracking data in speaker-specific experiment setups. The present study introduces a real-time magnetic resonance imaging database of emotional speech production from 10 speakers and presents articulatory analysis results of speech emotional expression using the database. Midsagittal vocal tract parameters (midsagittal distances and the vocal tract length) were parameterized based on a two-dimensional grid-line system, using image segmentation software. The principal feature analysis technique was applied to the grid-line system in order to find the major movement locations. Results reveal both speaker-dependent and speaker-independent variation patterns. For example, sad speech, a low arousal emotion, tends to show smaller opening for low vowels in the front cavity than the high arousal emotions more consistently than the other regions of the vocal tract. Happiness shows significantly shorter vocal tract length than anger and sadness in most speakers. Further details of speaker-dependent and speaker-independent speech articulation variation in emotional expression and their implications are described.
네이버, 다음, 구글 등의 웹 상에서 제공하는 지도 서비스와 KML, GML, GeoRSS와 같은 기술들을 이용하여 하나의 위치 정보로 통합한 지리정보를 사용자에게 제공해 줄 수 있는 연구들은 현재까지 활발히 진행되어 왔다. 그러나 이러한 연구들은 위치 정보만 통하여 줄 뿐 의미 정보까지 통합하여 사 용자들에게 다양한 정보들을 제공해주지 못한다. 이 논문에서는 KML, GML, GeoRSS 등으로 표현된 풍부한 지리정보들을 통합하여 웹 기반 지도 서비스에 제공해주는 시스템을 제안한다. 또한 지리정보 들의 스키마 통합을 위해 어댑터 기반 의미 처리 방법과 정적/동적 의미 관리 기반 접근 방법을 혼합 한 하이브리드 스키마 매칭(Hybrid Schema Matching, HSM) 방법을 제안하고, 제안 시스템의 평가를 위해 스키마 매칭을 위한 4가지 접근 방법과 비교 평가를 수행한다. 평가의 결과로 제안 시스템은 의 미 해석에 대한 신뢰성이 보장되고 시스템 구축 비용과 데이터 통합 비용이 상대적으로 낮다는 특징 을 지닌다.
We present the USC Speech and Vocal Tract Morphology MRI Database, a 17-speaker magnetic resonance imaging database for speech research. The database consists of real-time magnetic resonance images (rtMRI) of dynamic vocal tract shaping, denoised audio recorded simultaneously with rtMRI, and 3D volumetric MRI of vocal tract shapes during sustained speech sounds. We acquired 2D real-time MRI of vocal tract shaping during consonant-vowel-consonant sequences, vowel consonant-vowel sequences, read passages, and spontaneous speech. We acquired 3D volumetric MRI of the full set of vowels and continuant consonants of American English. Each 3D volumetric MRI was acquired in one 7-second scan in which the participant sustained the sound. This is the first database to combine rtMRI of dynamic vocal tract shaping and 3D volumetric MRI of the entire vocal tract. The database provides a unique resource with which to examine the relationship between vocal tract morphology and vocal tract function. The USC Speech and Vocal Tract Morphology MRI Database is provided free for research use at http://sail.usc.edu/span/morphdb.
The study of speech pathology involves evaluation and treatment of speech production related disorders affecting phonation, fluency, intonation and aeromechanical components of respiration. Recently, speech pathology has garnered special interest amongst machine learning and signal processing (ML-SP) scientists. This growth in interest is led by advances in novel data collection technology, data science, speech processing and computational modeling. These in turn have enabled scientists in better understanding both the causes and effects of pathological speech conditions. In this paper, we review the application of machine learning and signal processing techniques to speech pathology and specifically focus on three different aspects. First, we list challenges such as controlling subjectivity in pathological speech assessments and patient variability in the application of ML-SP tools to the domain. Second, we discuss feature design methods and machine learning algorithms using a combination of domain knowledge and data driven methods. Finally, we present some case studies related to analysis of pathological speech and discuss their design.
Recent advances in real-time magnetic resonance imaging (rtMRI) of the upper airway for acquiring speech production data provide unparalleled views of the dynamics of a speaker's vocal tract at very high frame rates (83 frames per second and even higher). This paper introduces an effort to collect and make available on-line rtMRI data corresponding to a large subset of the sounds of the world's languages as encoded in the International Phonetic Alphabet, with supplementary English words and phonetically-balanced texts, produced by four prominent phoneticians, using the latest rtMRI technology. The technique images oral as well as laryngeal articulator movements in the production of each sound category. This resource is envisioned as a teaching tool in pronunciation training, second language acquisition, and speech therapy.
This study investigates the relations between the degree of prominence and articulatory-prosodic cues in emotional speech.In particular, this study considers articulatory parameters driven from the Converter/Distributor (C/D) model.The goal is to obtain a better understanding of the link among syllable magnitude in the C/D model, the empirical way to measure it in literature, and syllable-level prominence, and to examine emotional variations appearing in this relation.Since prosodic variations are important cues for prominence and emotion in speech, relations with prosodic parameters (f0, energy, duration) are also considered.Electromagnetic articulography data of two speakers were used for analysis.The degree of prominence was computed on crowd-sourcing annotation data, using the Rapid Prosody Transcription.Results indicate that movements of linguistically critical articulator, energy, syllable magnitude measure are highly correlated with prominence; f0 is relatively less correlated.The movements of linguistically critical articulator tend to be more correlated than syllable magnitude measure.Inter-speaker variability and emotion-dependent variations are also reported.These results suggest complex relations between prominence and articulatory-prosodic cues.They also suggest that incorporating more articulatory and prosodic behaviors than the conventional way can better relate to perception of prominence.
The need for reliable, scalable and efficient diagnosis of Parkinson’s Disease (PD) is a major clinical need. Automating the diagnosis can lead to more accurate and objective predictions as well as provide insights regarding the nature of Parkinson’s condition. This paper proposes a fully automated system to rate the severity (UPDRS-III scale) of PD from patients’ speech. Specifically, the system captures atypicalities in an individual’s voice when performing multiple diverse speaking tasks and makes a unified prediction of the PD severity. The performance is tested in a cross-data setting, with different subjects and dissimilar recording conditions. Results indicate that (i) effective features vary depending on the nature of the specific speech task, (ii) additional novel feature sets to detect distortions in Parkinson’s speech significantly improve the prediction accuracy from the Interspeech15 Challenge baseline system and (iii) our fusion system based on an unsupervised clustering technique also improves the accuracy. Our system incorporates ivector and functionals for segmental features, non-linear time series features, speech rhythm and automatic speech recognition decoding based features. By its application on the Interspeech15 eating condition challenge, the system also shows its potential for detecting other sources of speech variability.
We propose a practical, feature-level and score-level fusion approach by combining acoustic and estimated articulatory information for both text independent and text dependent speaker verification. From a practical point of view, we study how to improve speaker verification performance by combining dynamic articulatory information with the conventional acoustic features. On text independent speaker verification, we find that concatenating articulatory features obtained from measured speech production data with conventional Mel-frequency cepstral coefficients (MFCCs) improves the performance dramatically. However, since directly measuring articulatory data is not feasible in many real world applications, we also experiment with estimated articulatory features obtained through acoustic-to-articulatory inversion. We explore both feature level and score level fusion methods and find that the overall system performance is significantly enhanced even with estimated articulatory features. Such a performance boost could be due to the inter-speaker variation information embedded in the estimated articulatory features. Since the dynamics of articulation contain important information, we included inverted articulatory trajectories in text dependent speaker verification. We demonstrate that the articulatory constraints introduced by inverted articulatory features help to reject wrong password trials and improve the performance after score level fusion. We evaluate the proposed methods on the X-ray Microbeam database and the RSR 2015 database, respectively, for the aforementioned two tasks. Experimental results show that we achieve more than 15% relative equal error rate reduction for both speaker verification tasks.
This paper provides the overview of a part of my doctoral thesis work on emotional speech production. The goal is two-folds: (i) to discover the rules in phonological-level and surface-level emotional speech production processes from the direct measurements of articulatory movements, and (ii) to develop computational models for the production processes and their applications. The achievements towards this goal comprise resource creation (data collection and tool development), scientific investigation and computational modeling. Also, this paper discusses on-going research onrich inversion modeling.
This paper compares prominence that listeners perceive with actual articulatory prominence. We calculated phrasal boundaries from articulatory patterns using an algorithm of the C/D model, and compared those calculated boundaries with perceived boundaries. The jaw displacements, measures of prominence, were measured using EMA; articulatory boundaries were derived from a C/D model algorithm. The data is a set of English sentences that vary in the placement of contrastive emphasis. Perception data were obtained from listeners who were asked to evaluate syllable prominence and syllable boundaries for these sentences. The results indicate that perception of syllable prominence shows strong correlations with articulatory prominence, showing that jaw displacement can be a strong perceptual cue for syllable prominence. Further, perception of syllable prominence is also correlated with algorithmicallycalculated articulatory syllable boundaries. These results encourage us to explore the relation between articulation and perception of language prosody in terms of the C/D model framework.
This study explores one aspect of the articulatory mechanism that underlies emotional speech production, namely, the behavior of linguistically critical and non-critical articulators in the encoding of emotional information. The hypothesis is that the possible larger kinematic variability in the behavior of non-critical articulators enables revealing underlying emotional expression goal more explicitly than that of the critical articulators; the critical articulators are strictly controlled in service of achieving linguistic goals and exhibit smaller kinematic variability. This hypothesis is examined by kinematic analysis of the movements of critical and non-critical speech articulators gathered using eletromagnetic articulography during spoken expressions of five categorical emotions. Analysis results at the level of consonant-vowel-consonant segments reveal that critical articulators for the consonants show more (less) peripheral articulations during production of the consonant-vowel-consonant syllables for high (low) arousal emotions, while non-critical articulators show less sensitive emotional variation of articulatory position to the linguistic gestures. Analysis results at the individual phonetic targets show that overall, between- and within-emotion variability in articulatory positions is larger for non-critical cases than for critical cases. Finally, the results of simulation experiments suggest that the postural variation of non-critical articulators depending on emotion is significantly associated with the controls of critical articulators.
Automatically evaluating pronunciation quality of non-native speech has seen tremendous success in both research and commercial settings, with applications in L2 learning. In this paper, submitted for the INTERSPEECH 2015 Degree of Nativeness Sub-Challenge, this problem is posed under a challenging cross corpora setting using speech data drawn from multiple speakers from a variety of language backgrounds (L1) reading different English sentences. Since the perception of non-nativeness is realized at the segmental and suprasegmental linguistic levels, we explore a number of acoustic cues at multiple time scales. We experiment with both data-driven and knowledge-inspired features that capture degree of nativeness from pauses in speech, speaking rate, rhythm/stress, and goodness of phone pronunciation. One promising finding is that highly accurate automated assessment can be attained using a small diverse set of intuitive and interpretable features. Performance is further boosted by smoothing scores across utterances from the same speaker; our best system significantly outperforms the challenge baseline.
USC-TIMIT is an extensive database of multimodal speech production data, developed to complement existing resources available to the speech research community and with the intention of being continuously refined and augmented. The database currently includes real-time magnetic resonance imaging data from five male and five female speakers of American English. Electromagnetic articulography data have also been presently collected from four of these speakers. The two modalities were recorded in two independent sessions while the subjects produced the same 460 sentence corpus used previously in the MOCHA-TIMIT database. In both cases the audio signal was recorded and synchronized with the articulatory data. The database and companion software are freely available to the research community.