Segmenting vocal tract articulators in real-time MRI (rtMRI) is a challenging dynamic image segmentation problem characterized by low contrast, rapid motion, and limited spatial resolution. However, while rtMRI acquisitions may provide synchronized acoustic signals, existing methods discard this information, and the few multimodal approaches that incorporate audio cannot be deployed when audio is unavailable. We propose a three-stage framework that leverages acoustic and phonological supervision during training while requiring only the rtMRI image at inference: phonological representations are converted into spatial bounding-box priors for articulator localization, visual and acoustic encoders are aligned via dual-level cross-modal contrastive pretraining, and the learned representations are fused through a cross-attention decoder, effectively transferring multimodal knowledge into a single-modality inference pipeline. Evaluated on 75-Speaker~Annot-16 and USC-TIMIT datasets, our method outperforms existing unimodal and multimodal methods, demonstrating that multimodal supervision provides transferable benefits for precise and clinically deployable vocal tract segmentation.
Understanding the relationship between vocal tract motion during speech and the resulting acoustic signal is crucial for aided clinical assessment and developing personalized treatment and rehabilitation strategies. Toward this goal, we introduce an audio-to-video generation framework for creating Real Time/cine-Magnetic Resonance Imaging (RT-/cine-MRI) visuals of the vocal tract from speech signals. Our framework first preprocesses RT-/cine-MRI sequences and speech samples to achieve temporal alignment, ensuring synchronization between visual and audio data. We then employ a modified stable diffusion model, integrating structural and temporal blocks, to effectively capture movement characteristics and temporal dynamics in the synchronized data. This process enables the generation of MRI sequences from new speech inputs, improving the conversion of audio into visual data. We evaluated our framework on healthy controls and tongue cancer patients by analyzing and comparing the vocal tract movements in synthesized videos. Our framework demonstrated adaptability to new speech inputs and effective generalization. In addition, positive human evaluations confirmed its effectiveness, with realistic and accurate visualizations, suggesting its potential for outpatient therapy and personalized simulation of vocal tract visualizations.
PURPOSE:Segmentation of velopharyngeal and vocal tract anatomy is crucial for quantifying structures and analyzing dynamics during speech. Automatic methods are needed to replace labor-intensive and poorly reproducible manual segmentation. To address this, we compared four multi-atlas-based methods to determine which provides the most accurate segmentation of the tongue, velum, and adenoid and how segmentation accuracy varies with the number of temporal frames. METHOD:Five spatiotemporal atlases of speech tasks were built, each with around 80 frames. We performed diffeomorphic registration of 10-40 frames of each atlas, then propagated labels to evaluate how the number of frames impacts performance. Four label fusion techniques-majority voting label fusion (MVLF), simultaneous truth and performance level estimation (STAPLE), multi-atlas label fusion (MALF), and corrective learning (CL)-were applied to evaluate segmentation accuracy. Performance was assessed using dice similarity coefficient (DSC) values across increasing numbers of input frames (10-40). RESULTS:CL consistently produced the highest segmentation accuracy, with average DSC improving from 0.89 at 10 frames to 0.92 at 40 frames. In contrast, MVLF, STAPLE, and MALF showed relatively flat trends, with DSC values ranging from 0.82 to 0.87. Statistical analysis confirmed that only CL exhibited a significant positive association between frame count and DSC. CONCLUSIONS:CL demonstrates superior performance in segmenting pediatric vocal tract structures in dynamic magnetic resonance imaging, particularly as the number of temporal frames increases. These findings support the use of CL in analyzing speech-related anatomy and highlight its potential to replace manual segmentation in small specialized data sets. SUPPLEMENTAL MATERIAL:https://doi.org/10.23641/asha.33107006.
Early diagnosis of attention-deficit/hyperactivity disorder (ADHD) in children plays a crucial role in improving outcomes in education and mental health. Diagnosing ADHD using neuroimaging data, however, remains challenging due to heterogeneous presentations and overlapping symptoms with other conditions. To address this, we propose a novel parameter-efficient transfer learning approach that adapts a large-scale 3D convolutional foundation model, pre-trained on CT images, to an MRI-based ADHD classification task. Our method introduces Low-Rank Adaptation (LoRA) in 3D by factorizing 3D convolutional kernels into 2D low-rank updates, dramatically reducing trainable parameters while achieving superior performance. In a five-fold cross-validated evaluation on a public diffusion MRI database, our 3D LoRA fine-tuning strategy achieved state-of-the-art results, with one model variant reaching 71.9% accuracy and another attaining an AUC of 0.716. Both variants use only 1.64 million trainable parameters (over 113× fewer than a fully fine-tuned foundation model). Our results represent one of the first successful cross-modal (CT-to-MRI) adaptations of a foundation model in neuroimaging, establishing a new benchmark for ADHD classification while greatly improving efficiency.
Fully automated myocardial segmentation from cardiac magnetic resonance imaging (MRI) is vital for efficient diagnosis and treatment planning. Although numerous automated methods have been proposed, they typically focus on single MRI sequences and therefore have difficulties in generalizing across vendors and across cardiac MRI protocols. Simultaneous analysis of complementary cardiac MRI sequences, such as cine, T1 mapping, and late gadolinium enhancement (LGE) MRI, remains challenging due to their distinct image characteristics and scanner-specific variations. To address these issues, we propose an unsupervised domain adaptation approach that allows robust myocardial segmentation across multi-vendor cine, T1, and LGE MRI data. In particular, we introduce a class-imbalance self-training framework to transfer information learned from a source domain with labels to any unlabeled target domain, while maintaining consistent performance across different MRI sequences. Our framework iteratively refines segmentation accuracy by generating pseudo-labels for target data using a hardness-aware strategy, thus effectively addressing the problem of class imbalance in cardiac MRI segmentation. To mitigate data scarcity following pseudo-label selection, we employ a variance-guided vicinal feature extrapolation, which expands data points in the feature space into a probabilistic distribution. This, in turn, facilitates joint source-target training by generating a larger intersection in the feature space. Experimental results demonstrate that our framework outperforms existing methods when assessed using the Dice coefficient and Hausdorff distance. Our framework enables cardiac evaluation across MRI protocols without sequence-specific manual annotations.
Tagged magnetic resonance imaging (tMRI) is a valuable tool for visualizing and quantifying tissue deformation in vivo. Its use is often hampered, however, by tag fading, long computation times, and the challenge of ensuring diffeomorphic, incompressible motion fields. In this paper, we describe a novel integration of the harmonic phase (HARP) approach to tMRI analysis with an unsupervised deep learning-based registration framework to estimate 2D and 3D motion fields that are diffeomorphic and nearly incompressible. The resulting method, called deep sinusoidally transformed HARP, or DSHARP, enables end-to-end network training by implementing a transformation of the harmonic phase to remove phase-wrapping discontinuities. It produces diffeomorphic motion by estimating a stationary velocity field from which motion is computed using the scaling and squaring technique. Finally, it encourages incompressibility using a novel Jacobian determinant loss term during network training. We evaluated DSHARP on 2D and 3D phantom data with simulated incompressible motions, real 3D human tongue data acquired during speech from both healthy and glossectomy subjects, and cardiac tagged MRI from the public STACOM 2011 benchmark. Our approach outperforms HARP, SinMod, SyN, PVIRA, VoxelMorph, and DeepTag in tracking accuracy, computation speed, and preservation of incompressibility.
BACKGROUND:Accurate delineation of the clinical target volume (CTV) is essential in the radiotherapy treatment of soft tissue sarcomas. However, this process is subject to inter-reader variability due to the need for clinical assessment of risk and extent of potential microscopic spread. This can lead to inconsistencies in treatment planning, potentially impacting treatment outcomes. Most existing automatic CTV delineation methods do not account for this variability and can only generate a single CTV for each case. PURPOSE:This study aims to develop a deep learning-based technique to generate multiple CTV contours for each case, simulating the inter-reader variability in the clinical practice. METHODS:We employed a publicly available dataset consisting of fluorodeoxyglucose positron emission tomography (FDG-PET), x-ray computed tomography (CT), and pre-contrast T1-weighted magnetic resonance imaging (MRI) scans from 51 patients with soft tissue sarcoma, along with an independent validation set containing five additional patients. An experienced reader drew a contour of the gross tumor volume (GTV) for each patient based on multi-modality images. Subsequently, two additional readers, together with the first one, were responsible for contouring three CTVs in total based on the GTV. We developed a diffusion model-based deep learning method that is capable of generating arbitrary number of different and plausible CTVs to mimic the inter-reader variability in CTV delineation. The proposed model incorporates a separate encoder to extract features from the GTV masks, leveraging the critical role of GTV information in accurate CTV delineation. RESULTS:The proposed diffusion model demonstrated superior performance with the highest Dice Index (0.902 compared to values below 0.881 for state-of-the-art models) and the best generalized energy distance (GED) (0.209 compared to values exceeding 0.221 for state-of-the-art models). It also achieved the second-highest recall and precision metrics among the compared ambiguous image segmentation models. Results from both datasets exhibited consistent trends, reinforcing the reliability of our findings. Additionally, ablation studies exploring different model structures and input configurations highlighted the significance of incorporating prior GTV information for accurate CTV delineation. CONCLUSIONS:The proposed diffusion model successfully generates multiple plausible CTV contours for soft tissue sarcomas, effectively capturing inter-reader variability in CTV delineation.
The human tongue is a muscular hydrostat that performs critical roles in speech, swallowing, and mastication through highly coordinated muscle activity. Previous research applied Granger causality analysis to time-series strain values from individual tongue muscles, revealing predictive relationships that shed light on their sequential interactions during protrusive and speech-related tasks. Building on these findings, this study shifts the focus from single muscles to functional units of tongue motion—groups of cohesive local muscle regions—to better understand how they collaborate to produce articulate speech. We collected diffusion MRI to capture muscle fiber orientation and tagged MRI to capture motion dynamics from four participants while they articulated the phrase “a kouk.” After validating the stationarity of the strains along fiber orientations with statistical tests, we identified optimal time lags and applied Granger causality analysis to evaluate predictive relationships among these functional units. Our pipeline provides new insights into tongue-movement coordination, enabling improved predictions of motion patterns and articulatory behaviors. These findings not only enhance our understanding of the biomechanics of speech but also hold promise for informing rehabilitative strategies for individuals with speech disorders.
The analysis of human speech with magnetic resonance imaging (MRI) provides essential information on the dynamic processes involved in speech production, allowing unobtrusive monitoring of the complete vocal tract during speech production. In clinical applications, personalized monitoring and increased speed of speech rehabilitation can be achieved through targeted phonological therapy i.e., by breaking down spoken words into their linguistic units. While MRI provides detailed visualization of the anatomical structures involved in speech production, the acoustic information provided by synchronized audio signals allows higher temporal resolution to process speech sounds.
Myocardial segmentation is crucial for accurate assessment of cardiac functions and pathology. Automatic segmentation of myocardial tissue from magnetic resonance (MR) images is actively researched especially with cine MR data that offer high spatial and temporal resolution. Challenges arise when processing cine images acquired from different scanner systems, as segmentation networks trained using one manufacturer's images may perform poorly on another's due to variations between training and test data distributions. Unsupervised domain adaptation addresses the domain shift issue by transferring knowledge from a labeled source domain to an unlabeled target domain, thereby offering a general solution framework. In addition, another major challenge lies in the pixel inconsistencies among different segmented label regions, unbalancing the reliability of segmentation across various image structures. This study proposes a domain-adaptive framework for myocardial segmentation using cine MR data from various manufacturers, employing a class-imbalance self-training structure. The framework iteratively refines the segmentation model using pseudo labels generated in the target domain by adopting a class-specific label threshold that re-balances pixel inconsistencies between different label regions. Embedded in a convolutional UNet architecture tailored for segmenting cine cardiac MR images, the proposed method was tested to adapt a UNet trained with MR data from a Siemens scanner to effectively segment MR data from a Philips scanner. Results show improved segmentation quality over two other comparison methods with respect to both the Dice similarity score and the Hausdorff distance.
Magnetic resonance (MR) tagging is an imaging technique for noninvasively tracking tissue motion in vivo by creating a visible pattern of magnetization saturation (tags) that deforms with the tissue. Due to longitudinal relaxation and progression to steady-state, the tags and tissue brightnesses change over time, which makes tracking with optical flow methods error-prone. Although Fourier methods can alleviate these problems, they are also sensitive to brightness changes as well as spectral spreading due to motion. To address these problems, we introduce the brightness-invariant tracking estimation (BRITE) technique for tagged MRI. BRITE disentangles the anatomy from the tag pattern in the observed tagged image sequence and simultaneously estimates the Lagrangian motion. The inherent ill-posedness of this problem is addressed by leveraging the expressive power of denoising diffusion probabilistic models to represent the probabilistic distribution of the underlying anatomy and the flexibility of physics-informed neural networks to estimate biologically-plausible motion. A set of tagged MR images of a gel phantom was acquired with various tag periods and imaging flip angles to demonstrate the impact of brightness variations and to validate our method. The results show that BRITE achieves more accurate motion and strain estimates as compared to other state of the art methods, while also being resistant to tag fading.
Accurate prediction of glioblastoma patient survival can significantly aid in personalized treatment planning. While pre-operative multimodal magnetic resonance imaging (MRI) offers complementary information, current methods are constrained by relatively limited data and largely rely on hand-crafted features extracted from segmentation results. To address these issues, in this work, we propose a data-efficient multi-task framework to take advantage of hierarchical segmentation features within advanced Swin UNETR for survival prediction. By integrating multi-scale features, we are able to capture detailed spatial information and global context, while employing the shifted window mechanism to maintain computational efficiency and scalability for 3D volumes. We further alleviate survival data scarcity through segmentation pre-training, while the features are fine-tuned to align with the survival prediction task and refined by statistical F-values. In addition, age information is incorporated alongside the extracted features to enhance survival prediction performance. Through comprehensive evaluations on the BraTS dataset, we demonstrate that our model achieves superior segmentation accuracy and state-of-the-art survival prediction performance, offering a robust solution for clinical prognosis in glioblastoma patients.
Accurate segmentation of the left ventricle (LV) in cardiac CT images is crucial for assessing ventricular function and diagnosing cardiovascular diseases. Creating a sufficiently large training set with accurate manual labels of LV can be cumbersome. More efficient semi-automatic segmentation, however, often includes unwanted structures, such as papillary muscles, due to low contrast between the LV wall and surrounding tissues. This study introduces a two-input-channel method within a Hybrid-Fusion Transformer deep-learning framework to produce refined LV labels from a combination of CT images and semi-automatic rough labels, effectively removing papillary muscles. By leveraging the efficiency of semi-automatic LV segmentation, we train an automatic refined segmentation model on a small set of images with both refined manual and rough semi-automatic labels. Evaluated through quantitative cross-validation, our method outperformed models that used only either CT images or rough masks as input.
Current semantic segmentation models typically require a substantial amount of manually annotated data, a process that is both time-consuming and resource-intensive. Alternatively, leveraging advanced text-to-image models such as Midjourney and Stable Diffusion has emerged as an efficient strategy, enabling the automatic generation of synthetic data in place of manual annotations. However, previous methods have been limited to generating singleinstance images, as the generation of multiple instances with Stable Diffusion has proven unstable and masks can be significantly affected by occlusion between different objects. To overcome this limitation and broaden the variety of synthetic datasets, we propose a novel framework, Free-Mask. It combines a Diffusion Model for segmentation with advanced image editing capabilities, allowing the insertion of multiple objects into images through text-to-image models. In addition, we introduce a new active learning paradigm that benefits both model generalization and data optimization. Our method enables the creation of realistic datasets that closely reflect open-world environments while generating accurate segmentation masks. Our code is released on GitHub.
Background and purpose:: Accurate delineation of the gross tumor volume (GTV) is essential for radiotherapy of soft tissue sarcomas. However, manual GTV delineation from multi-modality images is time-consuming. Furthermore, GTV delineation is subject to inter- and intra-reader variability, which reduces the reproducibility of treatment planning. To address these issues, this work aims to develop a highly accurate automatic delineation technique modeling reader variability for soft tissue sarcomas using deep learning. Materials and methods:: We employed a publicly available soft tissue sarcoma dataset consisting of Fluorodeoxyglucose Positron Emission Tomography (FDG-PET), X-ray Computed Tomography (CT), and pre-contrast T1-weighted Magnetic Resonance Imaging (MRI) scans for 51 patients, of which 49 were selected for analysis. The GTVs were delineated by six experienced readers, each reader performing GTV contouring multiple times for every patient. The confidence maps were calculated by averaging the labels provided by all readers, resulting in values ranging from 0 to 1. We developed and trained a diffusion model-based neural network to predict confidence maps of GTV for soft tissue sarcomas from multi-modality medical images. Results:: Quantitative analysis showed that the proposed diffusion model performed competitively with U-Net-based models, frequently ranking first or second across five evaluation metrics: Dice Index, Hausdorff Distance, Recall, Precision, and Brier Score. Additionally, experiments evaluating the impact of different imaging modalities demonstrated that incorporating multi-modality image inputs provided improved performance compared to single-modality and dual-modality inputs. Conclusion:: The proposed diffusion model is capable of predicting accurate confidence maps of GTV for soft tissue sarcomas from multi-modality inputs.
Bone marrow lesions (BMLs) are critical indicators of knee osteoarthritis (OA). Since they often appear as small, irregular structures with indistinguishable edges in knee magnetic resonance images (MRIs), effective detection of BMLs in MRI is vital for OA diagnosis and treatment. This paper proposes a semi-supervised local anomaly detection method using mask inpainting models for identification of BMLs in high-resolution knee MRI, effectively integrating a 3D femur bone segmentation model, a large mask inpainting model, and a series of post-processing techniques. The method was evaluated using MRIs at various resolutions from a subset of the public Osteoarthritis Initiative database. Dice score, Intersection over Union (IoU), and pixel-level sensitivity, specificity, and accuracy showed an advantage over the multiresolution knowledge distillation method-a state-of-the-art global anomaly detection method. Especially, segmentation performance is enhanced on higher-resolution images, achieving an over two times performance increase on the Dice score and the IoU score at a 448x448 resolution level. We also demonstrate that with increasing size of the BML region, both the Dice and IoU scores improve as the proportion of distinguishable boundary decreases. The identified BML masks can serve as markers for downstream tasks such as segmentation and classification. The proposed method has shown a potential in improving BML detection, laying a foundation for further advances in imaging-based OA research.
Accurate survival prediction using multimodal magnetic resonance imaging (MRI) plays a crucial role in clinical decision-making for patients with glioblastoma (GBM). In this work, we propose a multimodal framework, GlioSurvNet, that integrates deep learning features extracted from Swin UNETR and clinical variables to predict patient survival. Our framework makes use of multiple MRI sequences, including T1, T1 with contrast enhancement, T2-weighted, and FLAIR MRI, to capture diverse tumor characteristics. The Swin UNETR architecture simultaneously carries out tumor segmentation and extracts hierarchical features from multimodal MRI data. These deep learning features are then combined with clinical variables, which are input into a multi-layer perceptron network to yield survival probabilities. We evaluated our framework on a cohort of 287 patients from two independent databases, UPENN-GBM and UCSF-PDGM, demonstrating superior survival prediction performance when compared with existing methods. Our framework achieved a time-dependent concordance index of 0.693 and an integrated brier score of 0.14 with improved risk stratification. GlioSurvNet offers a robust tool for personalized prognosis and treatment planning in GBM patients.
Dynamic Magnetic Resonance Imaging (MRI) of the vocal tract has become an increasingly adopted imaging modality for speech motor studies. Beyond image signals, systematic data loss, noise pollution, and audio file corruption can occur due to the unpredictability of the MRI acquisition environment. In such cases, generating audio from images is critical for data recovery in both clinical and research applications. However, this remains challenging due to hardware constraints, acoustic interference, and data corruption. Existing solutions, such as denoising and multi-stage synthesis methods, face limitations in audio fidelity and generalizability. To address these challenges, we propose a Knowledge Enhanced Conditional Variational Autoencoder (KE-CVAE), a novel two-step “knowledge enhancement + variational inference” framework for generating speech audio signals from cine dynamic MRI sequences. This approach introduces two key innovations: (1) integration of unlabeled MRI data for knowledge enhancement, and (2) a variational inference architecture to improve generative modeling capacity. To the best of our knowledge, this is one of the first attempts at synthesizing speech audio directly from dynamic MRI video sequences. The proposed method was trained and evaluated on an open-source dynamic vocal tract MRI dataset recorded during speech. Experimental results demonstrate its effectiveness in generating natural speech waveforms while addressing MRI-specific acoustic challenges, outperforming conventional deep learning-based synthesis approaches ( https://github.com/YaxuanLi-cn/KE-CVAE ).
Accurate segmentation of articulatory structures in real-time MRI (rtMRI) remains challenging, as existing methods rely primarily on visual cues and overlook complementary information from synchronized speech signals. We propose VocSegMRI, a multimodal framework integrating video, audio, and phonological inputs via cross-attention fusion and a contrastive learning objective that improves cross-modal alignment and segmentation precision. Evaluated on USC-75 and further validated via zero-shot transfer on USC-TIMIT, VocSegMRI outperforms unimodal and multimodal baselines, with ablations confirming the contribution of each component.