Glioblastoma is a highly infiltrative brain tumor with fast progression and poor prognosis for patients. Due to the rapid growth, close treatment response monitoring is key. In this study, we benchmark different machine learning approaches for automated progression classification with various radiomic feature sets extracted from longitudinal magnetic resonance imaging and classifiers. Our experiments show differences in robustness and performance and offer insights into common failure modes. The best ROC-AUC was achieved with a random forest classifier without feature selection (0.748), and the best F1 score was at 0.792 for an XGBoost classifier where features of the current time point and the change from the reference time point were provided. Analyzing misclassifications shows different behavior for statistical machine learning classifiers and Residual Neural Networks.
Automated tumor segmentation tools for glioblastoma show promising performance. To apply these tools for automated response assessment, longitudinal segmentation, and tumor measurement, consistency is critical. This study aimed to determine whether BraTumIA and HD-GLIO are suited for this task. We evaluated two segmentation tools with respect to automated response assessment on the single-center retrospective LUMIERE dataset with 80 patients and a total of 502 post-operative time points. Volumetry and automated bi-dimensional measurements were compared with expert measurements following the Response Assessment in Neuro-Oncology (RANO) guidelines. The longitudinal trend agreement between the expert and methods was evaluated, and the RANO progression thresholds were tested against the expert-derived time-to-progression (TTP). The TTP and overall survival (OS) correlation was used to check the progression thresholds. We evaluated the automated detection and influence of non-measurable lesions. The tumor volume trend agreement calculated between segmentation volumes and the expert bi-dimensional measurements was high (HD-GLIO: 81.1%, BraTumIA: 79.7%). BraTumIA achieved the closest match to the expert TTP using the recommended RANO progression threshold. HD-GLIO-derived tumor volumes reached the highest correlation between TTP and OS (0.55). Both tools failed at an accurate lesion count across time. Manual false-positive removal and restricting to a maximum number of measurable lesions had no beneficial effect. Expert supervision and manual corrections are still necessary when applying the tested automated segmentation tools for automated response assessment. The longitudinal consistency of current segmentation tools needs further improvement. Validation of volumetric and bi-dimensional progression thresholds with multi-center studies is required to move toward volumetry-based response assessment.
Publicly available Glioblastoma (GBM) datasets predominantly include pre-operative Magnetic Resonance Imaging (MRI) or contain few follow-up images for each patient. Access to fully longitudinal datasets is critical to advance the refinement of treatment response assessment. We release a single-center longitudinal GBM MRI dataset with expert ratings of selected follow-up studies according to the response assessment in neuro-oncology criteria (RANO). The expert rating includes details about the rationale of the ratings. For a subset of patients, we provide pathology information regarding methylation of the O6-methylguanine-DNA methyltransferase (MGMT) promoter status and isocitrate dehydrogenase 1 (IDH1), as well as the overall survival time. The data includes T1-weighted pre- and post-contrast, T2-weighted, and fluid-attenuated inversion recovery (FLAIR) MRI. Segmentations from state-of-the-art automated segmentation tools, as well as radiomic features, complement the data. Possible applications of this dataset are radiomics research, the development and validation of automated segmentation methods, and studies on response assessment. This collection includes MRI data of 91 GBM patients with a total of 638 study dates and 2487 images.
Patients with Glioblastoma multiforme (GBM) have a very low overall survival (OS) time, due to the rapid growth an invasiveness of this brain tumor. As a contribution to the overall survival (OS) prediction task within the Brain Tumor Segmentation Challenge (BraTS), we classify the OS of GBM patients into overall survival classes based on information derived from pre-treatment Magnetic Resonance Imaging (MRI). The top-ranked methods from the past years almost exclusively used shape and position features. This is a remarkable contrast to the current advances in GBM radiomics showing a benefit of intensity-based features. This discrepancy may be caused by the inconsistent acquisition parameters in a multi-center setting. In this contribution, we test if normalizing the images based on the healthy tissue intensities enables the robust use of intensity features in this challenge. Based on these normalized images, we test the performance of 176 combinations of feature selection techniques and classifiers. Additionally, we test the incorporation of a sequence and robustness prior to limit the performance drop when models are applied to unseen data. The most robust performance on the training data (accuracy: 0.52± 0.09 ) was achieved with random forest regression, but this accuracy could not be maintained on the test set.
Background Automated brain tumor segmentation methods are computational algorithms that yield tumor delineation from, in this case, multimodal magnetic resonance imaging (MRI). We present an automated segmentation method and its results for resection cavity (RC) in glioblastoma multiforme (GBM) patients using deep learning (DL) technologies. Methods Post-operative, T1w with and without contrast, T2w and fluid attenuated inversion recovery MRI studies of 30 GBM patients were included. Three radiation oncologists manually delineated the RC to obtain a reference segmentation. We developed a DL cavity segmentation method, which utilizes all four MRI sequences and the reference segmentation to learn to perform RC delineations. We evaluated the segmentation method in terms of Dice coefficient (DC) and estimated volume measurements. Results Median DC of the three radiation oncologist were 0.85 (interquartile range [IQR]: 0.08), 0.84 (IQR: 0.07), and 0.86 (IQR: 0.07). The results of the automatic segmentation compared to the three different raters were 0.83 (IQR: 0.14), 0.81 (IQR: 0.12), and 0.81 (IQR: 0.13) which was significantly lower compared to the DC among raters (chi-square = 11.63, p = 0.04). We did not detect a statistically significant difference of the measured RC volumes for the different raters and the automated method (Kruskal-Wallis test: chi-square = 1.46, p = 0.69). The main sources of error were due to signal inhomogeneity and similar intensity patterns between cavity and brain tissues. Conclusions The proposed DL approach yields promising results for automated RC segmentation in this proof of concept study. Compared to human experts, the DC are still subpar.
Background This study aims to identify robust radiomic features for Magnetic Resonance Imaging (MRI), assess feature selection and machine learning methods for overall survival classification of Glioblastoma multiforme patients, and to robustify models trained on single-center data when applied to multi-center data. Methods Tumor regions were automatically segmented on MRI data, and 8327 radiomic features extracted from these regions. Single-center data was perturbed to assess radiomic feature robustness, with over 16 million tests of typical perturbations. Robust features were selected based on the Intraclass Correlation Coefficient to measure agreement across perturbations. Feature selectors and machine learning methods were compared to classify overall survival. Models trained on single-center data (63 patients) were tested on multi-center data (76 patients). Priors using feature robustness and clinical knowledge were evaluated. Results We observed a very large performance drop when applying models trained on single-center on unseen multi-center data, e.g. a decrease of the area under the receiver operating curve (AUC) of 0.56 for the overall survival classification boundary at 1 year. By using robust features alongside priors for two overall survival classes, the AUC drop could be reduced by 21.2%. In contrast, sensitivity was 12.19% lower when applying a prior. Conclusions Our experiments show that it is possible to attain improved levels of robustness and accuracy when models need to be applied to unseen multi-center data. The performance on multi-center data of models trained on single-center data can be increased by using robust features and introducing prior knowledge. For successful model robustification, tailoring perturbations for robustness testing to the target dataset is key.
The authors declare no potential conflict of interest.
ObjectivesTo identify qualitative VASARI (Visually AcceSIble Rembrandt Images) Magnetic Resonance (MR) Imaging features for differentiation of glioblastoma (GBM) and brain metastasis (BM) of different primary tumors.Materials and MethodsT1-weighted pre- and post-contrast, T2-weighted, and T2-weighted, fluid attenuated inversion recovery (FLAIR) MR images of a total of 239 lesions from 109 patients with either GBM or BM (breast cancer, non-small cell (NSCLC) adenocarcinoma, NSCLC squamous cell carcinoma, small-cell lung cancer (SCLC)) were included. A set of adapted, qualitative VASARI MR features describing tumor appearance and location was scored (binary; 1 = presence of feature, 0 = absence of feature). Exploratory data analysis was performed on binary scores using a combination of descriptive statistics (proportions with 95% binomial confidence intervals), unsupervised methods and supervised methods including multivariate feature ranking using either repeated fitting or recursive feature elimination with Support Vector Machines (SVMs).ResultsGBMs were found to involve all lobes of the cerebrum with a fronto-occipital gradient, often affected the corpus callosum (32.4%, 95% CI 19.1–49.2), and showed a strong preference for the right hemisphere (79.4%, 95% CI 63.2–89.7). BMs occurred most frequently in the frontal lobe (35.1%, 95% CI 28.9–41.9) and cerebellum (28.3%, 95% CI 22.6–34.8). The appearance of GBMs was characterized by preference for well-defined non-enhancing tumor margin (100%, 89.8–100), ependymal extension (52.9%, 36.7–68.5) and substantially less enhancing foci than BMs (44.1%, 28.9–60.6 vs. 75.1%, 68.8–80.5). Unsupervised and supervised analyses showed that GBMs are distinctively different from BMs and that this difference is driven by definition of non-enhancing tumor margin, ependymal extension and features describing laterality. Differentiation of histological subtypes of BMs was driven by the presence of well-defined enhancing and non-enhancing tumor margins and localization in the vision center. SVM models with optimal hyperparameters led to weighted F1-score of 0.865 for differentiation of GBMs from BMs and weighted F1-score of 0.326 for differentiation of BM subtypes.ConclusionVASARI MR imaging features related to definition of non-enhancing margin, ependymal extension, and tumor localization may serve as potential imaging biomarkers to differentiate GBMs from BMs.
Deep learning for regression tasks on medical imaging data has shown promising results. However, compared to other approaches, their power is strongly linked to the dataset size. In this study, we evaluate 3D-convolutional neural networks (CNNs) and classical regression methods with hand-crafted features for survival time regression of patients with high-grade brain tumors. The tested CNNs for regression showed promising but unstable results. The best performing deep learning approach reached an accuracy of 51.5% on held-out samples of the training set. All tested deep learning experiments were outperformed by a Support Vector Classifier (SVC) using 30 radiomic features. The investigated features included intensity, shape, location and deep features. The submitted method to the BraTS 2018 survival prediction challenge is an ensemble of SVCs, which reached a cross-validated accuracy of 72.2% on the BraTS 2018 training set, 57.1% on the validation set, and 42.9% on the testing set. The results suggest that more training data is necessary for a stable performance of a CNN model for direct regression from magnetic resonance images, and that non-imaging clinical patient information is crucial along with imaging information.
It is a general assumption in deep learning that more training data leads to better performance, and that models will learn to generalize well across heterogeneous input data as long as that variety is represented in the training set. Segmentation of brain tumors is a well-investigated topic in medical image computing, owing primarily to the availability of a large publicly-available dataset arising from the long-running yearly Multimodal Brain Tumor Segmentation (BraTS) challenge. Research efforts and publications addressing this dataset focus predominantly on technical improvements of model architectures and less on properties of the underlying data. Using the dataset and the method ranked third in the BraTS 2018 challenge, we performed experiments to examine the impact of tumor type on segmentation performance. We propose to stratify the training dataset into high-grade glioma (HGG) and low-grade glioma (LGG) subjects and train two separate models. Although we observed only minor gains in overall mean dice scores by this stratification, examining case-wise rankings of individual subjects revealed statistically significant improvements. Compared to a baseline model trained on both HGG and LGG cases, two separately trained models led to better performance in 64.9% of cases (p < 0.0001) for the tumor core. An analysis of subjects which did not profit from stratified training revealed that cases were missegmented which had poor image quality, or which presented clinically particularly challenging cases (e.g., underrepresented subtypes such as IDH1-mutant tumors), underlining the importance of such latent variables in the context of tumor segmentation. In summary, we found that segmentation models trained on the BraTS 2018 dataset, stratified according to tumor type, lead to a significant increase in segmentation performance. Furthermore, we demonstrated that this gain in segmentation performance is evident in the case-wise ranking of individual subjects but not in summary statistics. We conclude that it may be useful to consider the segmentation of brain tumors of different types or grades as separate tasks, rather than developing one tool to segment them all. Consequently, making this information available for the test data should be considered, potentially leading to a more clinically relevant BraTS competition.
Clinical use of MRSI is limited by the level of experience required to properly translate MRSI examinations into relevant clinical information. To solve this, several methods have been proposed to automatically recognize a predefined set of reference metabolic patterns. Given the variety of metabolic patterns seen in glioma patients, the decision on the optimal number of patterns that need to be used to describe the data is not trivial. In this paper, we propose a novel framework to (1) separate healthy from abnormal metabolic patterns and (2) retrieve an optimal number of reference patterns describing the most important types of abnormality. Using 41 MRSI examinations (1.5 T, PRESS, T E 135 ms) from 22 glioma patients, four different patterns describing different types of abnormality were detected: edema , healthy without Glx , active tumor and necrosis . The identified patterns were then evaluated on 17 MRSI examinations from nine different glioma patients. The results were compared against BraTumIA, an automatic segmentation method trained to identify different tumor compartments on structural MRI data. Finally, the ability to predict future contrast enhancement using the proposed approach was also evaluated.
Deep learning methods for brain tumor segmentation are typically trained in an ad hoc fashion on all available data. Brain tumors are tremendously heterogeneous in image appearance and labeled training data is limited. We argue that incorporation of additional prior information, specifically tumor grade, associated with tumor imaging phenotypes during model training can significantly improve segmentation performance. Two strategies for incorporation of tumor grade during model training are proposed and their impact on segmentation performance is demonstrated on the BRATS 2018 dataset.
Introduction:Imaging-based diagnosis of intra-axial contrast-enhancing brain tumors is frequently challenging. We show that the diagnosis of medulloblastoma (MDB) versus pilocytic astrocytoma (PA) and ependymoma (EPM) profit from computational analyses, based on quantitative image properties (i.e. textural features from apparent diffusion coefficient (ADC)-maps) and an automated machine learning classification (random forests (RF)). Methods:Forty patients who were diagnosed with three types of brain tumors were included in this study: 16 with MDB, 4 with PA, and 10 EPM. Based on the analysis of multi parametric preoperative magnetic resonance images, neuroradiologists gave a clear-cut diagnosis if they were sure of the diagnosis; however, most diagnoses comprise several possible tumor types. To distinguish between the named tumor types, a computer-based differential diagnosis (DD) tool was developed. Tumor lesion volumes were manually defined using ADC-maps only. From the demarked ADC-map, texture-parameters were extracted to train RF classifiers for pairwise DD. Performance of the RF models and reproducibility of the manual segmentation were evaluated. Results:Neuroradiologists gave correct and clear-cut diagnoses for 31% of MDB, 14.3% of PA, and 10% of EPM. Most diagnoses comprised several tumor types and altogether diagnoses containing the right tumor were given in 69% of true MDB, 64% of true PA, and 30% of true EPM. Ambiguous diagnoses could be improved by RF classifiers showing the following PA versus MDB performance: sensitivity 0.888 ± 0.031, specificity 0.886 ± 0.036; EPM versus MDB: sensitivity: 0.938 (95% CI = (0.677, 0.997)) and specificity: 0.7 (95% CI = (0.354, 0.919)); EPM versus PA: sensitivity: 0.786 (95% CI = (0.488, 0.942) and specificity: 0.100 (95% CI = (0.005, 0.458). An inter- and intra-rater analysis (three human raters) was performed and the Fleiss’ kappa test revealed high inter-rater agreement of κ = 0.821 (p value << 0.001) and an intra-rater agreement of κ = 0.822 (p value << 0.001). Conclusion:In the frequent case of ambiguous neuroradiologist diagnoses, a subsequent differential RF classification improves the diagnoses in all cases. The largest benefit is gained for the discrimination PA versus MDB with an accuracy of 88.0 ± 3.0% followed by EPM versus MDB with an accuracy of 84.6%.
Purpose: To improve the detection of peritumoral changes in GBM patients by exploring the relation between MRSI information and the distance to the solid tumor volume (STV) defined using structural MRI (sMRI).Methods: Twenty-three MRSI studies (PRESS, TE 135 ms) acquired from different patients with untreated GBM were used in this study. For each MRSI examination, the STV was identified by segmenting the corresponding sMRI images using BraTumIA, an automatic segmentation method. The relation between different metabolite ratios and the distance to STV was analyzed. A regression forest was trained to predict the distance from each voxel to STV based on 14 metabolite ratios. Then, the trained model was used to determine the expected distance to tumor (EDT) for each voxel of the MRSI test data. EDT maps were compared against sMRI segmentation.Results: The features showing abnormal values at the longest distances to the tumor were: %NAA, Glx/NAA, Cho/NAA, and Cho/Cr. These four features were also the most important for the prediction of the distances to STV. Each EDT value was associated with a specific metabolic pattern, ranging from normal brain tissue to actively proliferating tumor and necrosis. Low EDT values were highly associated with malignant features such as elevated Cho/NAA and Cho/Cr.Conclusion: The proposed method enables the automatic detection of metabolic patterns associated with different distances to the STV border and may assist tumor delineation of infiltrative brain tumors such as GBM.
Gliomas are the most common primary brain malignancies, with different degrees of aggressiveness, variable prognosis and various heterogeneous histologic sub-regions, i.e., peritumoral edematous/invaded tissue, necrotic core, active and non-enhancing core. This intrinsic heterogeneity is also portrayed in their radio-phenotype, as their sub-regions are depicted by varying intensity profiles disseminated across multi-parametric magnetic resonance imaging (mpMRI) scans, reflecting varying biological properties. Their heterogeneous shape, extent, and location are some of the factors that make these tumors difficult to resect, and in some cases inoperable. The amount of resected tumor is a factor also considered in longitudinal scans, when evaluating the apparent tumor for potential diagnosis of progression. Furthermore, there is mounting evidence that accurate segmentation of the various tumor sub-regions can offer the basis for quantitative image analysis towards prediction of patient overall survival. This study assesses the state-of-the-art machine learning (ML) methods used for brain tumor image analysis in mpMRI scans, during the last seven instances of the International Brain Tumor Segmentation (BraTS) challenge, i.e., 2012-2018. Specifically, we focus on i) evaluating segmentations of the various glioma sub-regions in pre-operative mpMRI scans, ii) assessing potential tumor progression by virtue of longitudinal growth of tumor sub-regions, beyond use of the RECIST/RANO criteria, and iii) predicting the overall survival from pre-operative mpMRI scans of patients that underwent gross total resection. Finally, we investigate the challenge of identifying the best ML algorithms for each of these tasks, considering that apart from being diverse on each instance of the challenge, the multi-institutional mpMRI BraTS dataset has also been a continuously evolving/growing dataset.
Uncertainty measures of medical image analysis technologies, such as deep learning, are expected to facilitate their clinical acceptance and synergies with human expertise. Therefore, we propose a full-resolution residual convolutional neural network (FRRN) for brain tumor segmentation and examine the principle of Monte Carlo (MC) Dropout for uncertainty quantification by focusing on the Dropout position and rate. We further feed the resulting brain tumor segmentation into a survival prediction model, which is built on age and a subset of 26 image-derived geometrical features such as volume, volume ratios, surface, surface irregularity and statistics of the enhancing tumor rim width. The results show comparable segmentation performance between MC Dropout models and a standard weight scaling Dropout model. A qualitative evaluation further suggests that informative uncertainty can be obtained by applying MC Dropout after each convolution layer. For survival prediction, results suggest only using few features besides age. In the BraTS17 challenge, our method achieved the 2nd place in the survival task and completed the segmentation task in the 3rd best-performing cluster of statistically different approaches.
OBJECTIVE In the treatment of glioblastoma, residual tumor burden is the only prognostic factor that can be actively influenced by therapy. Therefore, an accurate, reproducible, and objective measurement of residual tumor burden is necessary. This study aimed to evaluate the use of a fully automatic segmentation method-brain tumor image analysis (BraTumIA)-for estimating the extent of resection (EOR) and residual tumor volume (RTV) of contrast-enhancing tumor after surgery. METHODS The imaging data of 19 patients who underwent primary resection of histologically confirmed supratentorial glioblastoma were retrospectively reviewed. Contrast-enhancing tumors apparent on structural preoperative and immediate postoperative MR imaging in this patient cohort were segmented by 4 different raters and the automatic segmentation BraTumIA software. The manual and automatic results were quantitatively compared. RESULTS First, the interrater variabilities in the estimates of EOR and RTV were assessed for all human raters. Interrater agreement in terms of the coefficient of concordance (W) was higher for RTV (W = 0.812; p < 0.001) than for EOR (W = 0.775; p < 0.001). Second, the volumetric estimates of BraTumIA for all 19 patients were compared with the estimates of the human raters, which showed that for both EOR (W = 0.713; p < 0.001) and RTV (W = 0.693; p < 0.001) the estimates of BraTumIA were generally located close to or between the estimates of the human raters. No statistically significant differences were detected between the manual and automatic estimates. BraTumIA showed a tendency to overestimate contrast-enhancing tumors, leading to moderate agreement with expert raters with respect to the literature-based, survival-relevant threshold values for EOR. CONCLUSIONS BraTumIA can generate volumetric estimates of EOR and RTV, in a fully automatic fashion, which are comparable to the estimates of human experts. However, automated analysis showed a tendency to overestimate the volume of a contrast-enhancing tumor, whereas manual analysis is prone to subjectivity, thereby causing considerable interrater variability.
This paper proposes a novel approach for uncertainty quantification in dense Conditional Random Fields (CRFs). The presented approach, called Perturb-and-MPM, enables efficient, approximate sampling from dense multi-label CRFs via random perturbations. An analytic error analysis was performed which identified the main cause of approximation error as well as showed that the error is bounded. Spatial uncertainty maps can be derived from the Perturb-and-MPM model, which can be used to visualize uncertainty in image segmentation results. The method is validated on synthetic and clinical Magnetic Resonance Imaging data. The effectiveness of the approach is demonstrated on the challenging problem of segmenting the tumor core in glioblastoma. We found that areas of high uncertainty correspond well to wrongly segmented image regions. Furthermore, we demonstrate the potential use of uncertainty maps to refine imaging biomarkers in the case of extent of resection and residual tumor volume in brain tumor patients.
This paper extends a previously published brain tumor segmentation method with a dense Conditional Random Field (CRF). Dense CRFs can overcome the shrinking bias inherent to many grid-structured CRFs. We focus on illustrating the impact of alleviating the shrinking bias on the performance of CRF-based brain tumor segmentation. The proposed segmentation method is evaluated using data from the MICCAI BRATS 2013 & 2015 data sets (up to 110 patient cases for testing) and compared to a baseline method using a grid-structured CRF. Improved segmentation performance for the complete and enhancing tumor was observed with respect to grid-structured CRFs.
Objective Comparison of a fully-automated segmentation method that uses compartmental volume information to a semi-automatic user-guided and FDA-approved segmentation technique. Methods Nineteen patients with a recently diagnosed and histologically confirmed glioblastoma (GBM) were included and MR images were acquired with a 1.5 T MR scanner. Manual segmentation for volumetric analyses was performed using the open source software 3D Slicer version 4.2.2.3 (www.slicer.org). Semi-automatic segmentation was done by four independent neurosurgeons and neuroradiologists using the computer-assisted segmentation tool SmartBrush® (referred to as SB), a semi-automatic user-guided and FDA-approved tumor-outlining program that uses contour expansion. Fully automatic segmentations were performed with the Brain Tumor Image Analysis (BraTumIA, referred to as BT) software. We compared manual (ground truth, referred to as GT), computer-assisted (SB) and fully-automated (BT) segmentations with regard to: (1) products of two maximum diameters for 2D measurements, (2) the Dice coefficient, (3) the positive predictive value, (4) the sensitivity and (5) the volume error. Results Segmentations by the four expert raters resulted in a mean Dice coefficient between 0.72 and 0.77 using SB. BT achieved a mean Dice coefficient of 0.68. Significant differences were found for intermodal (BT vs. SB) and for intramodal (four SB expert raters) performances. The BT and SB segmentations of the contrast-enhancing volumes achieved a high correlation with the GT. Pearson correlation was 0.8 for BT; however, there were a few discrepancies between raters (BT and SB 1 only). Additional non-enhancing tumor tissue extending the SB volumes was found with BT in 16/19 cases. The clinically motivated sum of products of diameters measure (SPD) revealed neither significant intermodal nor intramodal variations. The analysis time for the four expert raters was faster (1 minute and 47 seconds to 3 minutes and 39 seconds) than with BT (5 minutes). Conclusion BT and SB provide comparable segmentation results in a clinical setting. SB provided similar SPD measures to BT and GT, but differed in the volume analysis in one of the four clinical raters. A major strength of BT may its independence from human interactions, it can thus be employed to handle large datasets and to associate tumor volumes with clinical and/or molecular datasets ("-omics") as well as for clinical analyses of brain tumor compartment volumes as baseline outcome parameters. Due to its multi-compartment segmentation it may provide information about GBM subcompartment compositions that may be subjected to clinical studies to investigate the delineation of the target volumes for adjuvant therapies in the future.