PURPOSE:To generate perfusion parameter maps from Time-of-flight magnetic resonance angiography (TOF-MRA) images using artificial intelligence to provide an alternative to traditional perfusion imaging techniques. MATERIALS AND METHODS:This retrospective study included a total of 272 patients with cerebrovascular diseases; 200 with acute stroke (from 2010 to 2018), and 72 with steno-occlusive disease (from 2011 to 2014). For each patient the TOF MRA image and the corresponding Dynamic susceptibility contrast magnetic resonance imaging (DSC-MRI) were retrieved from the datasets. The authors propose an adapted generative adversarial network (GAN) architecture, 3D pix2pix GAN, that generates common perfusion maps (CBF, CBV, MTT, TTP, Tmax) from TOF-MRA images. The performance was evaluated by the structural similarity index measure (SSIM). For a subset of 20 patients from the acute stroke dataset, the Dice coefficient was calculated to measure the overlap between the generated and real hypoperfused lesions with a time-to-maximum (Tmax) > 6 s. RESULTS:The GAN model exhibited high visual overlap and performance for all perfusion maps in both datasets: acute stroke (mean SSIM 0.88-0.92, mean PSNR 28.48-30.89, mean MAE 0.02-0.04 and mean NRMSE 0.14-0.37) and steno-occlusive disease patients (mean SSIM 0.83-0.98, mean PSNR 23.62-38.21, mean MAE 0.01-0.05 and mean NRMSE 0.03-0.15). For the overlap analysis for lesions with Tmax>6 s, the median Dice coefficient was 0.49. CONCLUSION:Our AI model can successfully generate perfusion parameter maps from TOF-MRA images, paving the way for a non-invasive alternative for assessing cerebral hemodynamics in cerebrovascular disease patients. This method could impact the stratification of patients with cerebrovascular diseases. Our results warrant more extensive refinement and validation of the method.
OBJECTIVES:To evaluate the transferability of deep learning (DL) models for the early detection of adverse events to previously unseen hospitals.DESIGN:Retrospective observational cohort study utilizing harmonized intensive care data from four public datasets.SETTING:ICUs across Europe and the United States.PATIENTS:Adult patients admitted to the ICU for at least 6 hours who had good data quality.INTERVENTIONS:None.MEASUREMENTS AND MAIN RESULTS:Using carefully harmonized data from a total of 334,812 ICU stays, we systematically assessed the transferability of DL models for three common adverse events: death, acute kidney injury (AKI), and sepsis. We tested whether using more than one data source and/or algorithmically optimizing for generalizability during training improves model performance at new hospitals. We found that models achieved high area under the receiver operating characteristic (AUROC) for mortality (0.838-0.869), AKI (0.823-0.866), and sepsis (0.749-0.824) at the training hospital. As expected, AUROC dropped when models were applied at other hospitals, sometimes by as much as -0.200. Using more than one dataset for training mitigated the performance drop, with multicenter models performing roughly on par with the best single-center model. Dedicated methods promoting generalizability did not noticeably improve performance in our experiments.CONCLUSIONS:Our results emphasize the importance of diverse training data for DL-based risk prediction. They suggest that as data from more hospitals become available for training, models may become increasingly generalizable. Even so, good performance at a new hospital still depended on the inclusion of compatible hospitals during training.
Stroke is a major cause of death or disability. As imaging-based patient stratification improves acute stroke therapy, dynamic susceptibility contrast magnetic resonance imaging (DSC-MRI) is of major interest in image brain perfusion. However, expert-level perfusion maps require a manual or semi-manual post-processing by a medical expert making the procedure time-consuming and less-standardized. Modern machine learning methods such as generative adversarial networks (GANs) have the potential to automate the perfusion map generation on an expert level without manual validation. We propose a modified pix2pix GAN with a temporal component (temp-pix2pix-GAN) that generates perfusion maps in an end-to-end fashion. We train our model on perfusion maps infused with expert knowledge to encode it into the GANs. The performance was trained and evaluated using the structural similarity index measure (SSIM) on two datasets including patients with acute stroke and the steno-occlusive disease. Our temp-pix2pix architecture showed high performance on the acute stroke dataset for all perfusion maps (mean SSIM 0.92-0.99) and good performance on data including patients with the steno-occlusive disease (mean SSIM 0.84-0.99). While clinical validation is still necessary for future studies, our results mark an important step toward automated expert-level perfusion maps and thus fast patient stratification.
Intracranial atherosclerotic disease (ICAD) poses a significant risk of subsequent stroke but current prevention strategies are limited. Mechanistic simulations of brain hemodynamics offer an alternative precision medicine approach by utilising individual patient characteristics. For clinical use, however, current simulation frameworks have insufficient validation. In this study, we performed the first quantitative validation of a simulation-based precision medicine framework to assess cerebral hemodynamics in patients with ICAD against clinical standard perfusion imaging. In a retrospective analysis, we used a 0-dimensional simulation model to detect brain areas that are hemodynamically vulnerable to subsequent stroke. The main outcome measures were sensitivity, specificity, and area under the receiver operating characteristics curve (ROC AUC) of the simulation to identify brain areas vulnerable to subsequent stroke as defined by quantitative measurements of relative mean transit time (relMTT) from dynamic susceptibility contrast MRI (DSC-MRI). In 68 subjects with unilateral stenosis >70% of the internal carotid artery (ICA) or middle cerebral artery (MCA), the sensitivity and specificity of the simulation were 0.65 and 0.67, respectively. The ROC AUC was 0.68. The low-to-moderate accuracy of the simulation may be attributed to assumptions of Newtonian blood flow, rigid vessel walls, and the use of time-of-flight MRI for geometric representation of subject vasculature. Future simulation approaches should focus on integrating additional patient data, increasing accessibility of precision medicine tools to clinicians, addressing disease burden disparities amongst different populations, and quantifying patient benefit. Our results underscore the need for further improvement of mechanistic simulations of brain hemodynamics to foster the translation of the technology to clinical practice.
Background Perfusion assessment in cerebrovascular disease is essential for evaluating cerebral hemodynamics and guides many current treatment decisions. Dynamic susceptibility contrast (DSC) magnetic resonance imaging (MRI) is of great utility to generate perfusion parameter maps, but its reliance on a contrast agent with associated health risks and technical challenges limit its usability. We hypothesized that native Time-of-flight magnetic resonance angiography (TOF-MRA) can be used to generate perfusion parameter maps with an artificial intelligence (AI) method, called generative adversarial network (GAN), offering a contrast-free alternative to DSC-MRI.Methods We propose an adapted 3D pix2pix GAN that generates common perfusion maps from TOF-MRA images (CBF, CBV, MTT, Tmax). The models are trained on two datasets consisting of 272 patients with acute stroke and steno-occlusive disease. The performance was evaluated by the structural similarity index measure (SSIM), for the acute dataset we calculated the Dice coefficient for lesions with a time-to-maximum (Tmax) >6s.Findings Our GAN model showed high visual overlap and high performance for all perfusion maps on both the acute stroke dataset (mean SSIM 0.88-092) and data including steno-occlusive disease patients (mean SSIM 0.83–0.98). For lesions of Tmax>6, the median Dice coefficient was 0.49.Interpretation Our study shows that our AI model can accurately generate perfusion parameter maps from TOF-MRA images, paving the way for clinical utility. We present a non-invasive alternative to contrast agent-based imaging for the assessment of cerebral hemodynamics in patients with cerebrovascular disease. Leveraging TOF-MRA data for the generation of perfusion maps represents a groundbreaking approach in cerebrovascular disease imaging. This method could greatly impact the stratification of patients with cerebrovascular diseases by providing an alternative to contrast agent-based perfusion assessment.Funding This work has received funding from the European Commission (Horizon2020 grant: PRECISE4Q No. 777107, coordinator: DF) and the German Federal Ministry of Education and Research (Go-Bio grant: PREDICTioN2020 No. 031B0154 lead: DF).### Competing Interest StatementDr. Madai reported receiving personal fees from ai4medicine outside the submitted work. Dr. Frey reported receiving grants from the European Commission, reported receiving personal fees from and holding an equity interest in ai4medicine outside the submitted work. While not related to this work, Dr Sobesky reports receipt of speakers honoraria from Pfizer, Boehringer Ingelheim, and Daiichi Sankyo.### Funding StatementThis work has received funding from the European Commission (Horizon2020 grant: PRECISE4Q No. 777107, coordinator: DF) and the German Federal Ministry of Education and Research (Go-Bio grant: PREDICTioN2020 No. 031B0154 lead: DF).### Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesThe details of the IRB/oversight body that provided approval or exemption for the research described are given below:For the data from Heidelberg, the ethics committee of Heidelberg University gave ethical approval for this work. For the PEGASUS study, the ethics committee of Charite - Universitatsmedizin Berlin gave ethical approval for this work.I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.YesThe datasets presented in this article are not readily available because data protection laws prohibit sharing of the PEGASUS and acute stroke datasets at the current time point. Requests to access these datasets should be directed to the Ethical Review Committee of Charite Universitatsmedizin Berlin, ethikkommission{at}charite.de.
Early and reliable prediction of shunt-dependent hydrocephalus (SDHC) after aneurysmal subarachnoid hemorrhage (aSAH) may decrease the duration of in-hospital stay and reduce the risk of catheter-associated meningitis. Machine learning (ML) may improve predictions of SDHC in comparison to traditional non-ML methods. ML models were trained for CHESS and SDASH and two combined individual feature sets with clinical, radiographic, and laboratory variables. Seven different algorithms were used including three types of generalized linear models (GLM) as well as a tree boosting (CatBoost) algorithm, a Naive Bayes (NB) classifier, and a multilayer perceptron (MLP) artificial neural net. The discrimination of the area under the curve (AUC) was classified (0.7 ≤ AUC < 0.8, acceptable; 0.8 ≤ AUC < 0.9, excellent; AUC ≥ 0.9, outstanding). Of the 292 patients included with aSAH, 28.8% ( n = 84) developed SDHC. Non-ML-based prediction of SDHC produced an acceptable performance with AUC values of 0.77 (CHESS) and 0.78 (SDASH). Using combined feature sets with more complex variables included than those incorporated in the scores, the ML models NB and MLP reached excellent performances, with an AUC of 0.80, respectively. After adding the amount of CSF drained within the first 14 days as a late feature to ML-based prediction, excellent performances were reached in the MLP (AUC 0.81), NB (AUC 0.80), and tree boosting model (AUC 0.81). ML models may enable clinicians to reliably predict the risk of SDHC after aSAH based exclusively on admission data. Future ML models may help optimize the management of SDHC in aSAH by avoiding delays in clinical decision-making.
Deep learning (DL) can aid doctors in detecting worsening patient states early, affording them time to react and prevent bad outcomes. While DL-based early warning models usually work well in the hospitals they were trained for, they tend to be less reliable when applied at new hospitals. This makes it difficult to deploy them at scale. Using carefully harmonised intensive care data from four data sources across Europe and the US (totalling 334,812 stays), we systematically assessed the reliability of DL models for three common adverse events: death, acute kidney injury (AKI), and sepsis. We tested whether using more than one data source and/or explicitly optimising for generalisability during training improves model performance at new hospitals. We found that models achieved high AUROC for mortality (0.838-0.869), AKI (0.823-0.866), and sepsis (0.749-0.824) at the training hospital. As expected, performance dropped at new hospitals, sometimes by as much as -0.200. Using more than one data source for training mitigated the performance drop, with multi-source models performing roughly on par with the best single-source model. This suggests that as data from more hospitals become available for training, model robustness is likely to increase, lower-bounding robustness with the performance of the most applicable data source in the training data. Dedicated methods promoting generalisability did not noticeably improve performance in our experiments.
Deep learning requires large labeled datasets that are difficult to gather in medical imaging due to data privacy issues and time-consuming manual labeling. Generative Adversarial Networks (GANs) can alleviate these challenges enabling synthesis of shareable data. While 2D GANs have been used to generate 2D images with their corresponding labels, they cannot capture the volumetric information of 3D medical imaging. 3D GANs are more suitable for this and have been used to generate 3D volumes but not their corresponding labels. One reason might be that synthesizing 3D volumes is challenging owing to computational limitations. In this work, we present 3D GANs for the generation of 3D medical image volumes with corresponding labels applying mixed precision to alleviate computational constraints. We generated 3D Time-of-Flight Magnetic Resonance Angiography (TOF-MRA) patches with their corresponding brain blood vessel segmentation labels. We used four variants of 3D Wasserstein GAN (WGAN) with: 1) gradient penalty (GP), 2) GP with spectral normalization (SN), 3) SN with mixed precision (SN-MP), and 4) SN-MP with double filters per layer (c-SN-MP). The generated patches were quantitatively evaluated using the Fréchet Inception Distance (FID) and Precision and Recall of Distributions (PRD). Further, 3D U-Nets were trained with patch-label pairs from different WGAN models and their performance was compared to the performance of a benchmark U-Net trained on real data. The segmentation performance of all U-Net models was assessed using Dice Similarity Coefficient (DSC) and balanced Average Hausdorff Distance (bAVD) for a) all vessels, and b) intracranial vessels only. Our results show that patches generated with WGAN models using mixed precision (SN-MP and c-SN-MP) yielded the lowest FID scores and the best PRD curves. Among the 3D U-Nets trained with synthetic patch-label pairs, c-SN-MP pairs achieved the highest DSC (0.841) and lowest bAVD (0.508) compared to the benchmark U-Net trained on real data (DSC 0.901; bAVD 0.294) for intracranial vessels. In conclusion, our solution generates realistic 3D TOF-MRA patches and labels for brain vessel segmentation. We demonstrate the benefit of using mixed precision for computational efficiency resulting in the best-performing GAN-architecture. Our work paves the way towards sharing of labeled 3D medical data which would increase generalizability of deep learning models for clinical use.
Splenomegaly is a common cross-sectional imaging finding with a variety of differential diagnoses. This study aimed to evaluate whether a deep learning model could automatically segment the spleen and identify the cause of splenomegaly in patients with cirrhotic portal hypertension versus patients with lymphoma disease. This retrospective study included 149 patients with splenomegaly on computed tomography (CT) images (77 patients with cirrhotic portal hypertension, 72 patients with lymphoma) who underwent a CT scan between October 2020 and July 2021. The dataset was divided into a training (n = 99), a validation (n = 25) and a test cohort (n = 25). In the first stage, the spleen was automatically segmented using a modified U-Net architecture. In the second stage, the CT images were classified into two groups using a 3D DenseNet to discriminate between the causes of splenomegaly, first using the whole abdominal CT, and second using only the spleen segmentation mask. The classification performances were evaluated using the area under the receiver operating characteristic curve (AUC), accuracy (ACC), sensitivity (SEN), and specificity (SPE). Occlusion sensitivity maps were applied to the whole abdominal CT images, to illustrate which regions were important for the prediction. When trained on the whole abdominal CT volume, the DenseNet was able to differentiate between the lymphoma and liver cirrhosis in the test cohort with an AUC of 0.88 and an ACC of 0.88. When the model was trained on the spleen segmentation mask, the performance decreased (AUC = 0.81, ACC = 0.76). Our model was able to accurately segment splenomegaly and recognize the underlying cause. Training on whole abdomen scans outperformed training using the segmentation mask. Nonetheless, considering the performance, a broader and more general application to differentiate other causes for splenomegaly is also conceivable.
Sharing labeled data is crucial to acquire large datasets for various Deep Learning applications. In medical imaging, this is often not feasible due to privacy regulations. Whereas anonymization would be a solution, standard techniques have been shown to be partially reversible. Here, synthetic data using a Generative Adversarial Network (GAN) with differential privacy guarantees could be a solution to ensure the patient's privacy while maintaining the predictive properties of the data. In this study, we implemented a Wasserstein GAN (WGAN) with and without differential privacy guarantees to generate privacy-preserving labeled Time-of-Flight Magnetic Resonance Angiography (TOF-MRA) image patches for brain vessel segmentation. The synthesized image-label pairs were used to train a U-net which was evaluated in terms of the segmentation performance on real patient images from two different datasets. Additionally, the Fréchet Inception Distance (FID) was calculated between the generated images and the real images to assess their similarity. During the evaluation using the U-Net and the FID, we explored the effect of different levels of privacy which was represented by the parameter ϵ. With stricter privacy guarantees, the segmentation performance and the similarity to the real patient images in terms of FID decreased. Our best segmentation model, trained on synthetic and private data, achieved a Dice Similarity Coefficient (DSC) of 0.75 for ϵ = 7.4 compared to 0.84 for ϵ = ∞ in a brain vessel segmentation paradigm (DSC of 0.69 and 0.88 on the second test set, respectively). We identified a threshold of ϵ <5 for which the performance (DSC <0.61) became unstable and not usable. Our synthesized labeled TOF-MRA images with strict privacy guarantees retained predictive properties necessary for segmenting the brain vessels. Although further research is warranted regarding generalizability to other imaging modalities and performance improvement, our results mark an encouraging first step for privacy-preserving data sharing in medical imaging.
Rupture of an intracranial aneurysm often results in subarachnoid hemorrhage, a life-threatening condition with high mortality and morbidity. The Cerebral Aneurysm Detection and Analysis (CADA) competition was organized to support the development and benchmarking of algorithms for the detection, analysis, and risk assessment of cerebral aneurysms in X-ray rotational angiography (3DRA) images. 109 anonymized 3DRA datasets were provided for training, and 22 additional datasets were used to test the algorithmic solutions. Cerebral aneurysm detection was assessed using the F2 score based on recall and precision, and the fit of the delivered bounding box was assessed using the distance to the aneurysm. Segmentation quality was measured using Jaccard and a combination of different surface distance measurements. Systematic errors were analyzed using volume correlation and bias. Rupture risk assessment was evaluated using the F2 score. 158 participants from 22 countries registered for the CADAchallenge. The detection solutions presented by the community are mostly accurate (F2 score 0.92) with a small number of missed aneurysms with diameters of 3.5 mm. In addition, the delineation of these structures is very good with a Jaccard score of 0.915. The rupture risk estimation methods achieved an F2 score of 0.7. The performance of the detection and segmentation solutions is equivalent to that of human experts. In rupture risk estimation, the best results are obtained by combining different image-based, morphological and computational fluid dynamic parameters using machine learning methods.
The Cerebral Aneurysm Detection and Analysis (CADA) challenge was organized to support the development and benchmarking of algorithms for detecting, analyzing, and risk assessment of cerebral aneurysms in X-ray rotational angiography (3DRA) images. 109 anonymized 3DRA datasets were provided for training, and 22 additional datasets were used to test the algorithmic solutions. Cerebral aneurysm detection was assessed using the F2 score based on recall and precision, and the fit of the delivered bounding box was assessed using the distance to the aneurysm. The segmentation quality was measured using the Jaccard index and a combination of different surface distance measures. Systematic errors were analyzed using volume correlation and bias. Rupture risk assessment was evaluated using the F2 score. 158 participants from 22 countries registered for the CADA challenge. The U-Net-based detection solutions presented by the community show similar accuracy compared to experts (F2 score 0.92), with a small number of missed aneurysms with diameters smaller than 3.5 mm. In addition, the delineation of these structures, based on U-Net variations, is excellent, with a Jaccard score of 0.92. The rupture risk estimation methods achieved an F2 score of 0.71. The performance of the detection and segmentation solutions is equivalent to that of human experts. The best results are obtained in rupture risk estimation by combining different image-based, morphological, and computational fluid dynamic parameters using machine learning methods. Furthermore, we evaluated the best methods pipeline, from detecting and delineating the vessel dilations to estimating the risk of rupture. The chain of these methods achieves an F2-score of 0.70, which is comparable to applying the risk prediction to the ground-truth delineation (0.71).
The aim of this study was to develop a deep learning-based algorithm for fully automated spleen segmentation using CT images and to evaluate the performance in conditions directly or indirectly affecting the spleen (e.g., splenomegaly, ascites). For this, a 3D U-Net was trained on an in-house dataset (n = 61) including diseases with and without splenic involvement (in-house U-Net), and an open-source dataset from the Medical Segmentation Decathlon (open dataset, n = 61) without splenic abnormalities (open U-Net). Both datasets were split into a training (n = 32.52%), a validation (n = 9.15%) and a testing dataset (n = 20.33%). The segmentation performances of the two models were measured using four established metrics, including the Dice Similarity Coefficient (DSC). On the open test dataset, the in-house and open U-Net achieved a mean DSC of 0.906 and 0.897 respectively (p = 0.526). On the in-house test dataset, the in-house U-Net achieved a mean DSC of 0.941, whereas the open U-Net obtained a mean DSC of 0.648 (p < 0.001), showing very poor segmentation results in patients with abnormalities in or surrounding the spleen. Thus, for reliable, fully automated spleen segmentation in clinical routine, the training dataset of a deep learning-based algorithm should include conditions that directly or indirectly affect the spleen.
Anonymization and data sharing are crucial for privacy protection and acquisition of large datasets for medical image analysis. This is a big challenge, especially for neuroimaging. Here, the brain?s unique structure allows for re-identification and thus requires non-conventional anonymization. Generative adversarial networks (GANs) have the potential to provide anonymous images while preserving predictive properties. Analyzing brain vessel segmentation, we trained 3 GANs on time-of-flight (TOF) magnetic resonance angiography (MRA) patches for image-label generation: 1) Deep convolutional GAN, 2) Wasserstein-GAN with gradient penalty (WGAN-GP) and 3) WGAN-GP with spectral normalization (WGAN-GP-SN). The generated image-labels from each GAN were used to train a U-net for segmentation and tested on real data. Moreover, we applied our synthetic patches using transfer learning on a second dataset. For an increasing number of up to 15 patients we evaluated the model performance on real data with and without pre-training. The performance for all models was assessed by the Dice Similarity Coefficient (DSC) and the 95th percentile of the Hausdorff Distance (95HD). Comparing the 3 GANs, the U-net trained on synthetic data generated by the WGAN-GP-SN showed the highest performance to predict vessels (DSC/95HD 0.85/30.00) benchmarked by the U-net trained on real data (0.89/ 26.57). The transfer learning approach showed superior performance for the same GAN compared to no pre training, especially for one patient only (0.91/24.66 vs. 0.84/27.36). In this work, synthetic image-label pairs retained generalizable information and showed good performance for vessel segmentation. Besides, we showed that synthetic patches can be used in a transfer learning approach with independent data. This paves the way to overcome the challenges of scarce data and anonymization in medical imaging.
Anonymization and data sharing are crucial for privacy protection and acquisition of large datasets for medical image analysis. This is a big challenge, especially for neuroimaging. Here, the brain's unique structure allows for re-identification and thus requires non-conventional anonymization. Generative adversarial networks (GANs) have the potential to provide anonymous images while preserving predictive properties. Analyzing brain vessel segmentation, we trained 3 GANs on time-of-flight (TOF) magnetic resonance angiography (MRA) patches for image-label generation: 1) Deep convolutional GAN, 2) Wasserstein-GAN with gradient penalty (WGAN-GP) and 3) WGAN-GP with spectral normalization (WGAN-GP-SN). The generated image-labels from each GAN were used to train a U-net for segmentation and tested on real data. Moreover, we applied our synthetic patches using transfer learning on a second dataset. For an increasing number of up to 15 patients we evaluated the model performance on real data with and without pre-training. The performance for all models was assessed by the Dice Similarity Coefficient (DSC) and the 95th percentile of the Hausdorff Distance (95HD). Comparing the 3 GANs, the U-net trained on synthetic data generated by the WGAN-GP-SN showed the highest performance to predict vessels (DSC/95HD 0.82/28.97) benchmarked by the U-net trained on real data (0.89/26.61). The transfer learning approach showed superior performance for the same GAN compared to no pre-training, especially for one patient only (0.91/25.68 vs. 0.85/27.36). In this work, synthetic image-label pairs retained generalizable information and showed good performance for vessel segmentation. Besides, we showed that synthetic patches can be used in a transfer learning approach with independent data. This paves the way to overcome the challenges of scarce data and anonymization in medical imaging.
Background and purpose Handling missing values is a prevalent challenge in the analysis of clinical data. The rise of data-driven models demands an efficient use of the available data. Methods to impute missing values are thus crucial. Here, we developed a publicly available framework to test different imputation methods and compared their impact in a typical stroke clinical dataset as a use case. Methods A clinical dataset based on the 1000Plus stroke study with 380 completed-entries patients was used. 13 common clinical parameters including numerical and categorical values were selected. Missing values in a missing-at-random (MAR) and missing-completely-at-random (MCAR) fashion from 0% to 60% were simulated and consequently imputed using the mean, hot-deck, multiple imputation by chained equations, expectation maximization method and listwise deletion. The performance was assessed by the root mean squared error, the absolute bias and the performance of a linear model for discharge mRS prediction. Results Listwise deletion was the worst performing method and started to be significantly worse than any imputation method from 2% (MAR) and 3% (MCAR) missing values on. The underlying missing value mechanism seemed to have a crucial influence on the identified best performing imputation method. Consequently no single imputation method outperformed all others. A significant performance drop of the linear model started from 11% (MAR+MCAR) and 18% (MCAR) missing values. Conclusions In the presented case study of a typical clinical stroke dataset we confirmed that listwise deletion should be avoided for dealing with missing values. Our findings indicate that the underlying missing value mechanism and other dataset characteristics strongly influence the best choice of imputation method. For future studies with similar data structure, we thus suggest to use the developed framework in this study to select the most suitable imputation method for a given dataset prior to analysis.
Brain vessel status is a promising biomarker for better prevention and treatment in cerebrovascular disease. However, classic rule-based vessel segmentation algorithms need to be hand-crafted and are insufficiently validated. A specialized deep learning method-the U-net-is a promising alternative. Using labeled data from 66 patients with cerebrovascular disease, the U-net framework was optimized and evaluated with three metrics: Dice coefficient, 95% Hausdorff distance (95HD) and average Hausdorff distance (AVD). The model performance was compared with the traditional segmentation method of graph-cuts. Training and reconstruction was performed using 2D patches. A full and a reduced architecture with less parameters were trained. We performed both quantitative and qualitative analyses. The U-net models yielded high performance for both the full and the reduced architecture: A Dice value of ~0.88, a 95HD of ~47 voxels and an AVD of ~0.4 voxels. The visual analysis revealed excellent performance in large vessels and sufficient performance in small vessels. Pathologies like cortical laminar necrosis and a rete mirabile led to limited segmentation performance in few patients. The U-net outperfomed the traditional graph-cuts method (Dice ~0.76, 95HD ~59, AVD ~1.97). Our work highly encourages the development of clinically applicable segmentation tools based on deep learning. Future works should focus on improved segmentation of small vessels and methodologies to deal with specific pathologies.
Introduction: Perfusion imaging by DSC-MRI (dynamic susceptibility contrast MRI) is the clinical method of choice for identification of penumbral flow (PF) in acute stroke. To date, the tissue at risk is estimated by a single predefined perfusion map. However, integration of various perfusion parameters may amplify the pathophysiological information and yield better estimation of PF. We therefore combined the common perfusion maps in a generalized linear model (GLM) to predict PF as defined by positron emission tomography (PET). Methods: In 18 patients with (sub)acute stroke, consecutive DSC-MRI and O15-water PET was performed (median age/NIHSS: 58 y , 12). PF was defined as cerebral-blood-flow (CBF) < 20 mL/100g/min on PET. MRI perfusion maps included: CBF, CBV, MTT, Tmax, TTP (cerebral-blood-volume, mean-transit-time, time-to-maximum and time-to-peak respectively). Probability maps for PF prediction were generated by a) single maps and b) multi-parametric maps (GLM) and underwent cross validation. ROC analysis assessed performance for PF prediction as area-under the curve (AUC). Results: Single maps showed AUC values between 0.57 and 0.72 (Tmax and CFB showing best performance). The GLM approach yielded an AUC of 0.75. Comparison by the Wilcoxon signed rank test showed that while the absolute difference was moderate, it was significant (p<0.04). Conclusions: Our results suggest that a multi-parameter perfusion model yields the highest accuracy for PF prediction. This finding, while preliminary, suggest a straight-forward model that can be easily integrated in clinical routine for improved stroke stratification based on the mismatch paradigm. Figure 1: Performance in penumbral flow prediction The graph shows performance in PF prediction for perfusion parameters and GLM. The error-bars represent standard-error. (*) marks significance for p value<0.05.
Background and Purpose— Identification of salvageable penumbra tissue by dynamic susceptibility contrast magnetic resonance imaging is a valuable tool for acute stroke patient stratification for treatment. However, prior studies have not attempted to combine the different perfusion maps into a predictive model. In this study, we established a multiparametric perfusion imaging model and cross-validated it using positron emission tomography perfusion for detection of penumbral flow. Methods— In a retrospective analysis of 17 subacute stroke patients with consecutive magnetic resonance imaging and H2O15 positron emission tomography scans, perfusion maps of cerebral blood flow, cerebral blood volume, mean transit time, time-to-maximum, and time-to-peak were constructed and combined using a generalized linear model (GLM). Both the GLM maps and the single perfusion maps alone were cross-validated with positron emission tomography-cerebral blood flow scans to predict penumbral flow on a voxel-wise level. Performance was tested by receiver-operating characteristics curve analysis, that is, the area under the curve, and the models’ fits were compared using the likelihood ratio test. Results— The GLM demonstrated significantly improved model fit compared with each of the single perfusion maps (P<1×e-5) and demonstrated higher performance, with an area under the curve of 0.91. However, the absolute difference between the performance of GLM and the best-performing single perfusion parameter (time-to-maximum) was relatively low (area under the curve difference =0.04). Conclusions— Our results support a dynamic susceptibility contrast magnetic resonance imaging–based GLM as an improved model for penumbral flow prediction in stroke patients. With given perfusion maps, this model is a straightforward and observer-independent alternative for therapy stratification.