We propose a multimodal latent diffusion model that jointly synthesizes volumetric magnetic resonance imaging (MRI) and tabular clinical data within a shared latent space via cross-attention. This approach enables coherent joint representation learning of MRI and tabular modalities for generative modeling. Our model utilizes a variational autoencoder to fuse the two modalities before diffusion-based synthesis, allowing modality-appropriate reconstruction with separate decoders for MRI and tabular data. We evaluated the framework on data from the German National Cohort (NAKO Gesundheitsstudie), comprising over 10,000 participants with MRI scans and clinical tabular features such as age, sex, body measurements, and ethnicity. The generated MRI volumes exhibited anatomical plausibility and body composition consistent with the synthesized tabular attributes. Quantitative evaluation using Frechet distance and precision-recall metrics confirmed high-fidelity image generation. In the tabular modality, our model outperformed CTGAN across standard evaluation metrics and achieved results comparable to TVAE, demonstrating competitive performance relative to established unimodal baselines. This work is, to our knowledge, the first to demonstrate the feasibility of jointly modeling MRI and mixed-type tabular data in a single latent diffusion framework, offering a proof-of-concept for generating coherent synthetic multimodal patient data and aligning with the broader goal of developing digital twins in healthcare.
Segmentation models in medical imaging that are able to segment a wide variety of structures and generalize on different image data is a relevant and recent research topic. With more and more universal segmentation models for both MR and CT being released recently a trend towards generalizing segmentation models for large amounts of structures can be observed. While universal MR segmentation models provide segmentations for a wide variety of structures, other structures are limited to models trained on CT data. One such model is TotalSegmentator which is able to segment up to 117 structures. In our work we present and evaluate a method to leverage models trained on CT data like the TotalSegmentator model for MRI data by training a structure-consistent CycleGAN on unpaired and unregistered data. We demonstrate the feasibility of using domain transfer by leveraging unlabeled and unpaired MR and CT datasets from various scanners and sites, with different sequences and protocols from the public AMOS22 abdomen dataset. This approach translates MR to CT contrast, allowing the synthetic CT image to be used as input for the TotalSegmentator model. Furthermore, we evaluate the segmentation accuracy of our approach on different structure types such as organs, muscles and bones on internal MR and CT datasets and compare them to the recently released TotalSegmentatorMRI and MRSegmentator models.
The evaluation of synthetic images has become a crucial aspect of research, especially with the rapid advancement of generative models. Recent methodologies have introduced various metrics that not only quantify the quality of synthetic images but also assess their realism, the extent to which they cover the source distribution, and whether the models produce duplicated samples from the original datasets. This is particularly significant in medical imaging, where the utility of synthetic images hinges on their realism and proximity to the source distribution. Traditionally, synthetic images are evaluated using lower-dimensional feature spaces derived from intermediate layers of trained neural networks. The Frechet Inception Distance (FID) remains the dominant metric, utilizing activations from an InceptionV3 model trained on the ImageNet dataset. In addition to FID, various other metrics have been proposed to address its limitations, yet most remain reliant on features from trained models. Despite its prevalence, FID's alignment with human evaluations has been questioned, especially in medical imaging contexts where the applicability of features learned from natural images is debated. The introduction of models trained on the RadImageNet dataset offers a promising alternative for extracting domain-specific features in medical imaging. This work builds on previous investigations by translating these approaches to the medical imaging domain and examining the sensitivity of different models trained on natural and medical image datasets to different augmentations. We propose to use systematic test-time data augmentations altering the evaluation data in controlled dimensions (contrast, shape) to characterize which of the pretrained models reacts in which way to the alterations. Our findings show which image features are captured by distinct representation spaces of these models, thereby facilitating the selection of models and advancing the evaluation of synthetic medical images.
Previous work on methods for cross domain generalization in medical imaging found a simple but very effective method called "global intensity non-linear" (GIN) augmentation. Our goal in this study is to use the GIN approach to train a model as powerful as TotalSegmentator for MRI data, despite having neither sufficient amounts of MRI data nor ground truth organ contours. Instead, we employ the GIN augmentation approach to show qualitatively and quantitatively that this is indeed feasible for a diverse set of anatomical structures including abdominal and thoracic organs as well as bones. The models are trained on the TotalSegmentator and AMOS22 datasets. For evaluation we apply them to whole body MRI scans from the German National Cohort (NAKO) study with a set of in-house reference masks. With GIN augmentation the mean Dice score of the model increases from 0.18 to 0.52 on Dixon water images, when using TotalSegmentator data for training. The improvements can be further split into 0.47 to 0.66 for abdominal organs, 0.55 to 0.79 for thoracic organs and 0.00 to 0.40 for bones.
The presented neural network with 3D-FiLM-cGAN architecture synthesizes cerebral blood flow maps from T1-weighted input images. Acquisition- and subject-specific metadata such as sex, arterial spin labeling (ASL) method and readout techniques were fed into the neural network as auxiliary input. The multi-vendor database including different ASL sequence types was created from ADNI data which were preprocessed in ExploreASL and transformed to MNI standard space. A subset of data from a single vendor (GE) were used for supervised training exemplarily and compared to CBF from acquired ASL data.
Varicose veins are classified as a chronic venous disease of which almost a quarter of the population of the U.S suffers from.1 Although most cases only develop mild symptoms, 6% of the affected women and men between 40 and 80 years develop signs of chronic vein insufficiency like venous ulceration.2 The number of these patients is two million in the U.S. alone. Treatment of varicose veins was mostly composed of surgical interventions until thermal endovenous ablation was introduced3 which resulted in lower cost and faster recovery of the patient.2 A new completely non-invasive method is High-Intensity Focused Ultrasound (HIFU) in which an ultrasound pulse is applied from outside the skin surface in order to thermally ablate the vein and close it permanently.3 This method relies heavily on diagnostic imaging through ultrasound to detect the target vein for ablation and to guide and monitor the procedure. An automated approach to detect and localize the vein during the treatment is rational because of the tedious work to follow the vessel in transversal direction. Previous works in the field of vessel segmentation in ultrasound images with deep learning focus on the frame-wise segmentation of the vessel.4 The possibility of further improvement of this method can be achieved by leveraging the temporal information about the location of the vessel. A previous work proposed by Mathai et. al.5 also features a U-net which implements LSTM-layers in the decoder part of the network and is used for the segmentation of vessels in ultrasound images. The segmentation of ultrasound image sequences can be combined with the prediction of segmentations of future frames to improve the predictive capacity of the model. Zhao et. al. proposed to use a ConvLSTM to predict future frames of ultrasound images for tongue movement,6 which was successful in predicting the next ultrasound image for a sequence of eight frames. In this work we propose a deep learning method for the localization and segmentation of veins in ultrasound sequences in combination with the prediction of future vessel segmentations for the automation of HIFU ablation treatments.
Synthesis of images has recently seen many works that produce high-quality real world images. In the domain of medical imaging the application of deep generative models especially Generative Adversarial Networks (GANs) can be applied to many different tasks. Under the premise of the generation of high-quality images that match the distribution of the original data, the synthesized data can be used to increase the size of small datasets, or in combination with conditioning on meta data, to increase the size of underrepresented classes in the dataset. In this work we propose a model that generates 3D medical images. The model can easily be conditioned on meta data, for example available patient information. We evaluate the quality of the generated images and compare our model against the 3D-StyleGAN model which is also designed for 3D medical image synthesis.
Magnetic resonance imaging (MRI) is the primary clinical tool to examine inflammatory brain lesions in Multiple Sclerosis (MS). Disease progression and inflammatory activities are examined by longitudinal image analysis to support diagnosis and treatment decision. Automated lesion segmentation methods based on deep convolutional neural networks (CNN) have been proposed, but are not yet applied in the clinical setting. Typical CNNs working on cross-sectional single time-point data have several limitations: changes to the image characteristics between single examinations due to scanner and protocol variations have an impact on the segmentation output, while at the same time the additional temporal correlation using pre-examinations is disregarded. In this work, we investigate approaches to overcome these limitations. Within a CNN architectural design, we propose convolutional Long Short-Term Memory (C-LSTM) networks to incorporate the temporal dimension. To reduce scanner- and protocol dependent variations between single MRI exams, we propose a histogram normalization technique as pre-processing step. The ISBI 2015 challenge data was used for network training and cross-validation. We demonstrate that the combination of the longitudinal normalization and CNN architecture increases the performance and the inter-time-point stability of the lesion segmentation. In the combined solution, the dice coefficient was increased and made more consistent for each subject. The proposed methods can therefore be used to increase the performance and stability of fully automated lesion segmentation applications in the clinical routine or in clinical trials.