BackgroundThe “translational roadblock” between successful animal stroke studies and neutral clinical trials is usually attributed to conceptual weaknesses. However, we hypothesized that rodent studies cannot inform the human disease due to intrinsic pathophysiological differences between rodents and humans., i.e., differences in infarct evolution.MethodsTo verify our hypothesis, we employed a mixed study design and compared findings from meta-analyses of animal studies and a retrospective clinical cohort study. For animal data, we systematically searched pubmed to identify all rodent studies, in which stroke was induced by MCAO and at least two sequential MRI scans were performed for infarct volume assessment within the first two days. For clinical data, we included 107 consecutive stroke patients with large artery occlusion, who received MRI scans upon admission and one or two days later.ResultsOur preclinical meta-analyses included 50 studies with 676 animals. Untreated animals had a median post-reperfusion infarct volume growth of 74%. Neuroprotective treatments reduced this infarct volume growth to 23%. A retrospective clinical cohort study showed that stroke patients had a median infarct volume growth of only 2% after successful recanalization. Stroke patients with unsuccessful recanalization, by contrast, experienced a meaningful median infarct growth of 148%.ConclusionOur study shows that rodents have a significant post-reperfusion infarct growth, and that this post-reperfusion infarct growth is the target of neuroprotective treatments. Stroke patients with successful recanalization do not have such infarct growth and thus have no target for neuroprotection.
Automated tumor segmentation tools for glioblastoma show promising performance. To apply these tools for automated response assessment, longitudinal segmentation, and tumor measurement, consistency is critical. This study aimed to determine whether BraTumIA and HD-GLIO are suited for this task. We evaluated two segmentation tools with respect to automated response assessment on the single-center retrospective LUMIERE dataset with 80 patients and a total of 502 post-operative time points. Volumetry and automated bi-dimensional measurements were compared with expert measurements following the Response Assessment in Neuro-Oncology (RANO) guidelines. The longitudinal trend agreement between the expert and methods was evaluated, and the RANO progression thresholds were tested against the expert-derived time-to-progression (TTP). The TTP and overall survival (OS) correlation was used to check the progression thresholds. We evaluated the automated detection and influence of non-measurable lesions. The tumor volume trend agreement calculated between segmentation volumes and the expert bi-dimensional measurements was high (HD-GLIO: 81.1%, BraTumIA: 79.7%). BraTumIA achieved the closest match to the expert TTP using the recommended RANO progression threshold. HD-GLIO-derived tumor volumes reached the highest correlation between TTP and OS (0.55). Both tools failed at an accurate lesion count across time. Manual false-positive removal and restricting to a maximum number of measurable lesions had no beneficial effect. Expert supervision and manual corrections are still necessary when applying the tested automated segmentation tools for automated response assessment. The longitudinal consistency of current segmentation tools needs further improvement. Validation of volumetric and bi-dimensional progression thresholds with multi-center studies is required to move toward volumetry-based response assessment.
Collateral circulation results from specialized anastomotic channels which are capable of providing oxygenated blood to regions with compromised blood flow caused by arterial obstruction. The quality of collateral circulation has been established as a key factor in determining the likelihood of a favorable clinical outcome and goes a long way to determining the choice of a stroke care model. Though many imaging and grading methods exist for quantifying collateral blood flow, the actual grading is mostly done through manual inspection. This approach is associated with a number of challenges. First, it is time-consuming. Second, there is a high tendency for bias and inconsistency in the final grade assigned to a patient depending on the experience level of the clinician. We present a multi-stage deep learning approach to predict collateral flow grading in stroke patients based on radiomic features extracted from MR perfusion data. First, we formulate a region of interest detection task as a reinforcement learning problem and train a deep learning network to automatically detect the occluded region within the 3D MR perfusion volumes. Second, we extract radiomic features from the obtained region of interest through local image descriptors and denoising auto-encoders. Finally, we apply a convolutional neural network and other machine learning classifiers to the extracted radiomic features to automatically predict the collateral flow grading of the given patient volume as one of three severity classes - no flow (0), moderate flow (1), and good flow (2). Results from our experiments show an overall accuracy of 72% in the three-class prediction task. With an inter-observer agreement of 16% and a maximum intra-observer agreement of 74% in a similar experiment, our automated deep learning approach demonstrates a performance comparable to expert grading, is faster than visual inspection, and eliminates the problem of grading bias.
Epileptic seizures require a rapid and safe diagnosis to minimize the time from onset to adequate treatment. Some epileptic seizures can be diagnosed clinically with the respective expertise. For more subtle seizures, imaging is mandatory to rule out treatable structural lesions and potentially life-threatening conditions. MRI perfusion abnormalities associated with epileptic seizures have been reported in CT and MRI studies. However, the interpretation of transient peri-ictal MRI abnormalities is routinely based on qualitative visual analysis and therefore reader dependent. In this retrospective study, we investigated the diagnostic yield of visual analysis of perfusion MRI during ictal and postictal states based on comparative expert ratings in 51 patients. We further propose an automated semi-quantitative method for perfusion analysis to determine perfusion abnormalities observed during ictal and postictal MRI using dynamic susceptibility contrast MRI, which we validated on a subcohort of 27 patients. The semi-quantitative method provides a parcellation of 3D T1-weighted images into 32 standardized cortical regions of interests and subcortical grey matter structures based on a recently proposed method, direct cortical thickness estimation using deep learning-based anatomy segmentation and cortex parcellation for brain anatomy segmentation. Standard perfusion maps from a Food and Drug Administration-approved image analysis tool (Olea Sphere 3.0) were co-registered and investigated for region-wise differences between ictal and postictal states. These results were compared against the visual analysis of two readers experienced in functional image analysis in epilepsy. In the ictal group, cortical hyperperfusion was present in 17/18 patients (94% sensitivity), whereas in the postictal cohort, cortical hypoperfusion was present only in 9/33 (27%) patients while 24/33 (73%) showed normal perfusion. The (semi-)quantitative dynamic susceptibility contrast MRI perfusion analysis indicated increased thalamic perfusion in the ictal cohort and hypoperfusion in the postictal cohort. Visual ratings between expert readers performed well on the patient level, but visual rating agreement was low for analysis of subregions of the brain. The asymmetry of the automated image analysis correlated significantly with the visual consensus ratings of both readers. We conclude that expert analysis of dynamic susceptibility contrast MRI effectively discriminates ictal versus postictal perfusion patterns. Automated perfusion evaluation revealed favourable interpretability and correlated well with the classification of the visual ratings. It may therefore be employed for high-throughput, large-scale perfusion analysis in extended cohorts, especially for research questions with limited expert rater capacity.
Although machine learning (ML) has shown promise in numerous domains, there are concerns about generalizability to out-of-sample data. This is currently addressed by centrally sharing ample, and importantly diverse, data from multiple sites. However, such centralization is challenging to scale (or even not feasible) due to various limitations. Federated ML (FL) provides an alternative to train accurate and generalizable ML models, by only sharing numerical model updates. Here we present findings from the largest FL study to-date, involving data from 71 healthcare institutions across 6 continents, to generate an automatic tumor boundary detector for the rare disease of glioblastoma, utilizing the largest dataset of such patients ever used in the literature (25,256 MRI scans from 6,314 patients). We demonstrate a 33% improvement over a publicly trained model to delineate the surgically targetable tumor, and 23% improvement over the tumor's entire extent. We anticipate our study to: 1) enable more studies in healthcare informed by large and diverse data, ensuring meaningful results for rare diseases and underrepresented populations, 2) facilitate further quantitative analyses for glioblastoma via performance optimization of our consensus model for eventual public release, and 3) demonstrate the effectiveness of FL at such scale and task complexity as a paradigm shift for multi-site collaborations, alleviating the need for data sharing.
Background and Objectives Very poor outcome despite IV thrombolysis (IVT) and mechanical thrombectomy (MT) occurs in approximately 1 of 4 patients with ischemic stroke and is associated with a high logistic and economic burden. We aimed to develop and validate a multivariable prognostic model to identify futile recanalization therapies (FRTs) in patients undergoing those therapies. Methods Patients from a prospectively collected observational registry of a single academic stroke center treated with MT and/or IVT were included. The data set was split into a training (N = 1,808, 80%) and internal validation (N = 453, 20%) cohort. We used gradient boosted decision tree machine learning models after k-nearest neighbor imputation of 32 variables available at admission to predict FRT defined as modified Rankin scale 5–6 at 3 months. We report feature importance, ability for discrimination, calibration, and decision curve analysis. Results A total of 2,261 patients with a median (interquartile range) age of 75 years (64–83 years), 46% female, median NIH Stroke Scale 9 (4–17), 34% IVT alone, 41% MT alone, and 25% bridging were included. Overall, 539 (24%) had FRT, more often in MT alone (34%) as compared with IVT alone (11%). Feature importance identified clinical variables (stroke severity, age, active cancer, prestroke disability), laboratory values (glucose, C-reactive protein, creatinine), imaging biomarkers (white matter hyperintensities), and onset-to-admission time as the most important predictors. The final model was discriminatory for predicting 3-month FRT (area under the curve 0.87, 95% CI 0.87–0.88) and had good calibration (Brier 0.12, 0.11–0.12). Overall performance was moderate (F1-score 0.63 ± 0.004), and decision curve analyses suggested higher mean net benefit at lower thresholds of treatment (up to 0.8). Conclusions This FRT prediction model can help inform shared decision making and identify the most relevant features in the emergency setting. Although it might be particularly useful in low resource healthcare settings, incorporation of further multifaceted variables is necessary to further increase the predictive performance.
Background The degree of reperfusion is the most important modifiable predictor of 3 month functional outcome and mortality in ischemic stroke patients treated with mechanical thrombectomy. Whether the beneficial effect of reperfusion also leads to a reduction in long term mortality is unknown. Methods Patients undergoing mechanical thrombectomy between January 2010 and December 2018 were included. The post-thrombectomy degree of reperfusion and emboli in new territories were core laboratory adjudicated. Reperfusion was evaluated according to the expanded Thrombolysis in Cerebral Infarction (eTICI) scale. Vital status was obtained from the Swiss population register. Adjusted hazard ratios (aHRs) using time split Cox regression models were calculated. Subgroup analyses were performed in patients with borderline indications. Results Our study included 1264 patients (median follow-up per patient 2.5 years). Patients with successful reperfusion had longer survival times, attributable to a lower hazard of death within 0-90 days and for >90 days to 2 years (aHR 0.34, 95% CI 0.26 to 0.46; aHR 0.37, 95% CI 0.22 to 0.62). This association was homogeneous across all predefined subgroups (p for interaction >0.05). Among patients with successful reperfusion, a significant difference in the hazard of death was observed between eTICI2b50 and eTICI3 (aHR 0.51, 95% CI 0.33 to 0.79). Emboli in new territories were present in 5% of patients, and were associated with increased mortality (aHR 2.3, 95% CI 1.11 to 4.86). Conclusion Successful, and ideally complete, reperfusion without emboli in new territories is associated with a reduction in long term mortality in patients treated with mechanical thrombectomy, and this was evident across several subgroups.
The volumes of equine teeth may change considerably over time for several reasons including domestication, routine dental floating, and the hypsodont and anelodont nature of the teeth. Cone beam computed tomography (CBCT) of the head is routinely performed in standing horses and, in this proof of concept study, the feasibility of measuring tooth volume from CBCT datasets was determined. The CBCT images of 5 equine cadaver cheek teeth were segmented with a software 3-dimensional (3D) Slicer using a predefined protocol, corrected manually, and re-assembled into a 3D model. Individual tooth volume (V-S) was calculated from the model. After extraction, the volumes were also measured using the "gold-standard" water displacement method (V-W) for comparison. The V-S of 77 teeth ranged from 7114 to 42,300 mm(3) which strongly correlated with V-W (r = 0.99)(,) and on average V-S was 6.1% less than V-W. There was no significant difference in V-S between the right and left arcades in individual animals. Maxillary cheek tooth volume was on average 40% larger than it was for mandibular counterparts. Semi-automatic image segmentation of equine cheek teeth from CBCT data is feasible and accurate but requires some manual intervention. This preliminary study provides initial data on the volume of equine cheek teeth and creates new possibilities for future in vivo studies.
(1) Background: To test the accuracy of a fully automated stroke tissue estimation algorithm (FASTER) to predict final lesion volumes in an independent dataset in patients with acute stroke; (2) Methods: Tissue-at-risk prediction was performed in 31 stroke patients presenting with a proximal middle cerebral artery occlusion. FDA-cleared perfusion software using the AHA recommendation for the Tmax threshold delay was tested against a prediction algorithm trained on an independent perfusion software using artificial intelligence (FASTER). Following our endovascular strategy to consequently achieve TICI 3 outcome, we compared patients with complete reperfusion (TICI 3) vs. no reperfusion (TICI 0) after mechanical thrombectomy. Final infarct volume was determined on a routine follow-up MRI or CT at 90 days after the stroke; (3) Results: Compared to the reference standard (infarct volume after 90 days), the decision forest algorithm overestimated the final infarct volume in patients without reperfusion. Underestimation was observed if patients were completely reperfused. In cases where the FDA-cleared segmentation was not interpretable due to improper definitions of the arterial input function, the decision forest provided reliable results; (4) Conclusions: The prediction accuracy of automated tissue estimation depends on (i) success of reperfusion, (ii) infarct size, and (iii) software-related factors introduced by the training sample. A principal advantage of machine learning algorithms is their improved robustness to artifacts in comparison to solely threshold-based model-dependent software. Validation on independent datasets remains a crucial condition for clinical implementations of decision support systems in stroke imaging.
The BraTS dataset contains a mixture of high-grade and low-grade gliomas, which have a rather different appearance: previous studies have shown that performance can be improved by separated training on low-grade gliomas (LGGs) and high-grade gliomas (HGGs), but in practice this information is not available at test time to decide which model to use. By contrast with HGGs, LGGs often present no sharp boundary between the tumor core and the surrounding edema, but rather a gradual reduction of tumor-cell density. Utilizing our 3D-to-2D fully convolutional architecture, DeepSCAN, which ranked highly in the 2019 BraTS challenge and was trained using an uncertainty-aware loss, we separate cases into those with a confidently segmented core, and those with a vaguely segmented or missing core. Since by assumption every tumor has a core, we reduce the threshold for classification of core tissue in those cases where the core, as segmented by the classifier, is vaguely defined or missing. We then predict survival of high-grade glioma patients using a fusion of linear regression and random forest classification, based on age, number of distinct tumor components, and number of distinct tumor cores. We present results on the validation dataset of the Multimodal Brain Tumor Segmentation Challenge 2020 (segmentation and uncertainty challenge), and on the testing set, where the method achieved 4th place in Segmentation, 1st place in uncertainty estimation, and 1st place in Survival prediction.
Background Automated brain tumor segmentation methods are computational algorithms that yield tumor delineation from, in this case, multimodal magnetic resonance imaging (MRI). We present an automated segmentation method and its results for resection cavity (RC) in glioblastoma multiforme (GBM) patients using deep learning (DL) technologies. Methods Post-operative, T1w with and without contrast, T2w and fluid attenuated inversion recovery MRI studies of 30 GBM patients were included. Three radiation oncologists manually delineated the RC to obtain a reference segmentation. We developed a DL cavity segmentation method, which utilizes all four MRI sequences and the reference segmentation to learn to perform RC delineations. We evaluated the segmentation method in terms of Dice coefficient (DC) and estimated volume measurements. Results Median DC of the three radiation oncologist were 0.85 (interquartile range [IQR]: 0.08), 0.84 (IQR: 0.07), and 0.86 (IQR: 0.07). The results of the automatic segmentation compared to the three different raters were 0.83 (IQR: 0.14), 0.81 (IQR: 0.12), and 0.81 (IQR: 0.13) which was significantly lower compared to the DC among raters (chi-square = 11.63, p = 0.04). We did not detect a statistically significant difference of the measured RC volumes for the different raters and the automated method (Kruskal-Wallis test: chi-square = 1.46, p = 0.69). The main sources of error were due to signal inhomogeneity and similar intensity patterns between cavity and brain tissues. Conclusions The proposed DL approach yields promising results for automated RC segmentation in this proof of concept study. Compared to human experts, the DC are still subpar.
Objectives: The objectives of this study included the volumetric analysis of persistent infarction and lesion reversal in Diffusion-Weighted Imaging (DWI), as well as the assessment of accuracy of ADC thresholds to identify regions of persistent infarction in patients with acute ischemic stroke after successful endovascular treatment (EVT). Materials and Methods: A retrospective analysis of patients with M1 or proximal M2 occlusions, treated between 01/2012 and 07/2017, who underwent successful EVT ([≥]TICI 2b) and both pre- and post-interventional Magnetic Resonance Imaging (MRI), led to the inclusion of N=90 patients. Administration of recombinant tissue plasminogen activator (rTPA) for intra-venous thrombolysis was performed ahead of intervention in 45 cases (N=45/90, 50%). The majority of patients (N=78/90, 86.7%) were treated with second-generation thrombectomy devices with or without intra-arterial urokinase. DWI at admission and 24-hour follow-up DWI data were co-registered. Acute ischemic changes at baseline DWI, 24-hour DWI lesion, and the affected gray/white matter regions were manually annotated. Persistent infarction was defined as acute ischemic changes on baseline DWI, which were sustained on 24-hour follow-up DWI. Based on the manual annotations, persistent infarction and DWI reversal were quantified in a voxel-wise analysis. Thresholds for the identification of persistent infarction using baseline ADC images were estimated by maximizing Youden's J statistic (ROC-analysis). Results: Median age of the patients was 71.9 years (IQR 60.4-79.7 years), 55.6% were female, and NIHSS at admission was 11 (IQR 6-14). The median DWI lesion volume at baseline was 9.9 mL (IQR 3-23.6 mL) and the median DWI lesion volume around 24 hours was 12.1 mL (IQR 3.6-23.7 mL). Reversal of acute ischemic changes occurred frequently (49.8%, IQR 31.7%-65.4%; percentage of initial DWI lesion volume per subject). Sizeable DWI reversal (i.e. >10 mL and >10%) was observed in 26.7% (N=24/90) of the cases. Relative DWI reversal was significantly higher in white matter (58.6%, IQR 35.3-81.5%) than in gray matter (39.2%, IQR 24.9-56.6%; p<0.001). The volume of persistent infarction and DWI reversal were both significantly correlated with the DWI lesion volume at baseline (R=0.873-0.945, p<0.001), however, no correlations with time to reperfusion were found (relative volumes: R=-/+0.058, p=0.607). ROC analyses of ADC thresholds yielded optimal values which differed significantly for gray and white matter (p=0.003), and were lower than previously reported thresholds while having significantly improved accuracy (p[≤]0.015). No correlations between the estimated ADC thresholds and different covariates were found (time from imaging to reperfusion, time from baseline to follow-up imaging, volume of acute ischemic changes). Conclusions: DWI reversal occurs frequently in successfully reperfused patients treated with modern EVT. Identification of persistent infarction using ADC thresholds in baseline DWI remains challenging with notable differences for gray and white matter.
As artificial intelligence (AI) systems begin to make their way into clinical radiology practice, it is crucial to assure that they function correctly and that they gain the trust of experts. Toward this goal, approaches to make AI "interpretable" have gained attention to enhance the understanding of a machine learning algorithm, despite its complexity. This article aims to provide insights into the current state of the art of interpretability methods for radiology AI. This review discusses radiologists' opinions on the topic and suggests trends and challenges that need to be addressed to effectively streamline interpretability methods in clinical practice. Supplemental material is available for this article. © RSNA, 2020 See also the commentary by Gastounioti and Kontos in this issue.
ObjectivesTo identify qualitative VASARI (Visually AcceSIble Rembrandt Images) Magnetic Resonance (MR) Imaging features for differentiation of glioblastoma (GBM) and brain metastasis (BM) of different primary tumors.Materials and MethodsT1-weighted pre- and post-contrast, T2-weighted, and T2-weighted, fluid attenuated inversion recovery (FLAIR) MR images of a total of 239 lesions from 109 patients with either GBM or BM (breast cancer, non-small cell (NSCLC) adenocarcinoma, NSCLC squamous cell carcinoma, small-cell lung cancer (SCLC)) were included. A set of adapted, qualitative VASARI MR features describing tumor appearance and location was scored (binary; 1 = presence of feature, 0 = absence of feature). Exploratory data analysis was performed on binary scores using a combination of descriptive statistics (proportions with 95% binomial confidence intervals), unsupervised methods and supervised methods including multivariate feature ranking using either repeated fitting or recursive feature elimination with Support Vector Machines (SVMs).ResultsGBMs were found to involve all lobes of the cerebrum with a fronto-occipital gradient, often affected the corpus callosum (32.4%, 95% CI 19.1–49.2), and showed a strong preference for the right hemisphere (79.4%, 95% CI 63.2–89.7). BMs occurred most frequently in the frontal lobe (35.1%, 95% CI 28.9–41.9) and cerebellum (28.3%, 95% CI 22.6–34.8). The appearance of GBMs was characterized by preference for well-defined non-enhancing tumor margin (100%, 89.8–100), ependymal extension (52.9%, 36.7–68.5) and substantially less enhancing foci than BMs (44.1%, 28.9–60.6 vs. 75.1%, 68.8–80.5). Unsupervised and supervised analyses showed that GBMs are distinctively different from BMs and that this difference is driven by definition of non-enhancing tumor margin, ependymal extension and features describing laterality. Differentiation of histological subtypes of BMs was driven by the presence of well-defined enhancing and non-enhancing tumor margins and localization in the vision center. SVM models with optimal hyperparameters led to weighted F1-score of 0.865 for differentiation of GBMs from BMs and weighted F1-score of 0.326 for differentiation of BM subtypes.ConclusionVASARI MR imaging features related to definition of non-enhancing margin, ependymal extension, and tumor localization may serve as potential imaging biomarkers to differentiate GBMs from BMs.
AbstractObjectivesMachine learning (ML) has been demonstrated to improve the prediction of functional outcome in patients with acute ischemic stroke. However, its value in a specific clinical use case has not been investigated. Aim of this study was to assess the clinical utility of ML models with respect to predicting functional impairment and severe disability or death considering its potential value as a decision-support tool in an acute stroke workflow.Materials and MethodsPatients (n=1317) from a retrospective, non-randomized observational registry treated with Mechanical Thrombectomy (MT) were included. The final dataset of patients who underwent successful recanalization (TICI ≥ 2b) (n=932) was split in order to develop ML-based prediction models using data of (n=745, 80%) patients. Subsequently, the models were tested on the remaining patient data (n=187, 20%). For comparison, baseline algorithms using majority class prediction, SPAN-100 score, PRE score, and Stroke-TPI score were implemented. The ML methods included eight different algorithms (e.g. Support Vector Machines and Random forests), stacked ensemble method and tabular neural networks. Prediction of modified Rankin Scale (mRS) 3–6 (primary analysis) and mRS 5–6 (secondary analysis) at 3 months was performed using 25 baseline variables available at patient admission. ML models were assessed with respect to their ability for discrimination, calibration and clinical utility (decision curve analysis).ResultsAnalyzed patients (n=932) showed a median age of 74.7 (IQR 62.7–82.4) years with (n=461, 49.5%) being female. ML methods performed better than clinical scores with stacked ensemble method providing the best overall performance including an F1-score of 0.75 ± 0.01, an ROC-AUC of 0.81 ± 0.00, AP score of 0.81 ± 0.01, MCC of 0.48 ± 0.02, and ECE of 0.06 ± 0.01 for prediction of mRS 3–6, and an F1-score of 0.57 ± 0.02, an ROC-AUC of 0.79 ± 0.01, AP score of 0.54 ± 0.02, MCC of 0.39 ± 0.03, and ECE of 0.19 ± 0.01 for prediction of mRS 5–6. Decision curve analyses suggested highest mean net benefit of 0.09 ± 0.02 at a-priori defined threshold (0.8) for the stacked ensemble method in primary analysis (mRS 3–6). Across all methods, higher mean net benefits were achieved for optimized probability thresholds but with considerably reduced certainty (threshold probabilities 0.24–0.47). For the secondary analysis (mRS 5–6), none of the ML models achieved a positive net benefit for the a-priori threshold probability 0.8.ConclusionsThe clinical utility of ML prediction models in a decision-support scenario aimed at yielding a high certainty for prediction of functional dependency (mRS 3–6) is marginal and not evident for the prediction of severe disability or death (mRS 5–6). Hence, using those models for patient exclusion cannot be recommended and future research should evaluate utility gains after incorporating more advanced imaging parameters.
Predicting the final ischaemic stroke lesion provides crucial information regarding the volume of salvageable hypoperfused tissue, which helps physicians in the difficult decision-making process of treatment planning and intervention. Treatment selection is influenced by clinical diagnosis, which requires delineating the stroke lesion, as well as characterising cerebral blood flow dynamics using neuroimaging acquisitions. Nonetheless, predicting the final stroke lesion is an intricate task, due to the variability in lesion size, shape, location and the underlying cerebral haemodynamic processes that occur after the ischaemic stroke takes place. Moreover, since elapsed time between stroke and treatment is related to the loss of brain tissue, assessing and predicting the final stroke lesion needs to be performed in a short period of time, which makes the task even more complex. Therefore, there is a need for automatic methods that predict the final stroke lesion and support physicians in the treatment decision process. We propose a fully automatic deep learning method based on unsupervised and supervised learning to predict the final stroke lesion after 90 days. Our aim is to predict the final stroke lesion location and extent, taking into account the underlying cerebral blood flow dynamics that can influence the prediction. To achieve this, we propose a two-branch Restricted Boltzmann Machine, which provides specialized data-driven features from different sets of standard parametric Magnetic Resonance Imaging maps. These data-driven feature maps are then combined with the parametric Magnetic Resonance Imaging maps, and fed to a Convolutional and Recurrent Neural Network architecture. We evaluated our proposal on the publicly available ISLES 2017 testing dataset, reaching a Dice score of 0.38, Hausdorff Distance of 29.21 mm, and Average Symmetric Surface Distance of 5.52 mm.
We introduce a modification of our previous 3D-to-2D fully convolutional architecture, DeepSCAN, replacing batch normalization with instance normalization, and adding a lightweight local attention mechanism. These networks are trained using a previously described loss function which mo els label noise and uncertainty. We present results on the validation dataset of the Multimodal Brain Tumor Segmentation Challenge 2019.
Data on infarcts in new territory (INT) in patients undergoing endovascular stroke treatment for acute large-vessel occlusions are sparse. Aim of this study was to assess the prevalence, risk factors, and clinical relevance of INT. For this purpose, all patients in a single-center prospective registry who underwent endovascular stroke treatment and received pre- and post-interventional diffusion-weighted imaging were included (N = 259). Using an established scoring system, INT were classified according to size (I-III, ≤2 mm, >2 mm ≤20 mm, >20 mm) and likelihood of being related to the intervention (A, high likelihood; B, low likelihood). Additionally, a new type of infarct, that occurred in a territory distal to the occlusion, but was initially not hypoperfused, was defined as an infarct in initially not hypoperfused territory (IINHT). A total of 180 INT and 38 IINHT were observed in 32.8% (N = 85/259) of patients. In most patients, INT were angiographically occult (90.2%), and 13 patients had INT/IINHT larger than 2 cm (type III). Absence of protection during stent-retrieval and a cardio-embolic stroke origin were associated with higher incidence of INT/IINHT, whereas pretreatment with IV tPA showed no association, even when different bolus timing was considered. INT/IINHT were associated with lower rates of functional independence with increasing size type after adjusting for confounders ( adjusted Odds Ratio per size group increase 0.63, 95% confidence interval 0.46–0.86). In conclusion, INT and IINHT are not rare, are associated with poor outcome with increasing size, and they may serve as a surrogate endpoint for safety evaluation of new devices and endovascular techniques. Further research on associated factors is warranted.
It is a general assumption in deep learning that more training data leads to better performance, and that models will learn to generalize well across heterogeneous input data as long as that variety is represented in the training set. Segmentation of brain tumors is a well-investigated topic in medical image computing, owing primarily to the availability of a large publicly-available dataset arising from the long-running yearly Multimodal Brain Tumor Segmentation (BraTS) challenge. Research efforts and publications addressing this dataset focus predominantly on technical improvements of model architectures and less on properties of the underlying data. Using the dataset and the method ranked third in the BraTS 2018 challenge, we performed experiments to examine the impact of tumor type on segmentation performance. We propose to stratify the training dataset into high-grade glioma (HGG) and low-grade glioma (LGG) subjects and train two separate models. Although we observed only minor gains in overall mean dice scores by this stratification, examining case-wise rankings of individual subjects revealed statistically significant improvements. Compared to a baseline model trained on both HGG and LGG cases, two separately trained models led to better performance in 64.9% of cases (p < 0.0001) for the tumor core. An analysis of subjects which did not profit from stratified training revealed that cases were missegmented which had poor image quality, or which presented clinically particularly challenging cases (e.g., underrepresented subtypes such as IDH1-mutant tumors), underlining the importance of such latent variables in the context of tumor segmentation. In summary, we found that segmentation models trained on the BraTS 2018 dataset, stratified according to tumor type, lead to a significant increase in segmentation performance. Furthermore, we demonstrated that this gain in segmentation performance is evident in the case-wise ranking of individual subjects but not in summary statistics. We conclude that it may be useful to consider the segmentation of brain tumors of different types or grades as separate tasks, rather than developing one tool to segment them all. Consequently, making this information available for the test data should be considered, potentially leading to a more clinically relevant BraTS competition.
In applications of supervised learning applied to medical image segmentation, the need for large amounts of labeled data typically goes unquestioned. In particular, in the case of brain anatomy segmentation, hundreds or thousands of weakly-labeled volumes are often used as training data. In this paper, we first observe that for many brain structures, a small number of training examples, (n=9), weakly labeled using Freesurfer 6.0, plus simple data augmentation, suffice as training data to achieve high performance, achieving an overall mean Dice coefficient of $0.84 \pm 0.12$ compared to Freesurfer over 28 brain structures in T1-weighted images of $\approx 4000$ 9-10 year-olds from the Adolescent Brain Cognitive Development study. We then examine two varieties of heteroscedastic network as a method for improving classification results. An existing proposal by Kendall and Gal, which uses Monte-Carlo inference to learn to predict the variance of each prediction, yields an overall mean Dice of $0.85 \pm 0.14$ and showed statistically significant improvements over 25 brain structures. Meanwhile a novel heteroscedastic network which directly learns the probability that an example has been mislabeled yielded an overall mean Dice of $0.87 \pm 0.11$ and showed statistically significant improvements over all but one of the brain structures considered. The loss function associated to this network can be interpreted as performing a form of learned label smoothing, where labels are only smoothed where they are judged to be uncertain.