Objective. The Dutch proton robustness evaluation protocol prescribes the dose of the clinical target volume (CTV) to the voxel-wise minimum (VWmin) dose of 28 scenarios. This results in a consistent but conservative near-minimum CTV dose (D98%,CTV). In this study, we analyzed (i) the correlation between VWmin/voxel-wise maximum (VWmax) metrics and actually delivered dose to the CTV and organs at risk (OARs) under the impact of treatment errors, and (ii) the performance of the protocol before and after its calibration with adequate prescription-dose levels. Approach. Twenty-one neuro-oncological patients were included. Polynomial chaos expansion was applied to perform a probabilistic robustness evaluation using 100,000 complete fractionated treatments per patient. Patient-specific scenario distributions of clinically relevant dosimetric parameters for the CTV and OARs were determined and compared to clinical VWmin and VWmax dose metrics for different scenario subsets used in the robustness evaluation protocol. Main results. The inclusion of more geometrical scenarios leads to a significant increase of the conservativism of the protocol in terms of clinical VWmin and VWmax values for the CTV and OARs. The protocol could be calibrated using VWmin dose evaluation levels of 93.0%-92.3%, depending on the scenario subset selected. Despite this calibration of the protocol, robustness recipes for proton therapy showed remaining differences and an increased sensitivity to geometrical random errors compared to photon-based margin recipes. Significance. The Dutch proton robustness evaluation protocol, combined with the photon-based margin recipe, could be calibrated with a VWmin evaluation dose level of 92.5%. However, it shows limitations in predicting robustness in dose, especially for the near-maximum dose metrics to OARs. Consistent robustness recipes could improve proton treatment planning to calibrate residual differences from photon-based assumptions.
PURPOSE:To demonstrate the feasibility of predicting the patient-specific treatment planning Pareto front (PF) for prostate cancer patients based only on delineations of PTV, rectum and body.MATERIAL/METHODS:Our methodology consists of four steps. First, using Erasmus-iCycle, the Pareto fronts of 112 prostate cancer patients were constructed by generating per patient 42 Pareto optimal treatment plans with different priorities. Dose parameters associated to homogeneity, conformity and dose to rectum were extracted. Second, a 3D convex function representing the PF spanned by the 42 plans was fitted for each patient using three patient-specific parameters. Third, ten features were extracted from the, aforementioned, structures to train a linear-regressor prediction algorithm to predict these three patient-specific parameters. Fourth, the quality of the predictions was assessed by calculating the average and maximum distances of the predicted PF to the 42 plans for patients in the validation cohort.RESULTS:The prediction model was able to predict the clinically relevant PF within 2 Gy for 90% of the patients with a median average distance of 0.6 Gy.CONCLUSIONS:We demonstrate the feasibility of fast, accurate predictions of the patient-specific PF for prostate cancer patients based only on delineations of PTV, rectum and body.
PURPOSEThe goal of this study was to generate a large treatment plan database for head and neck (H&N) cancer patients that can be considered as the gold standard to train and validate models for knowledge-based (KB) treatment planning and QA. With this dataset, the intrinsic prediction performance, the effect of interorgan dependency, and the impact of dataset inconsistency was investigated for an existing treatment planning QA model.METHODSThe CT scans of 108 previously treated oropharyngeal patients were used to establish the plan database. For each patient, 15 Pareto optimal treatment plans with different planning priorities for the parotid glands were generated with fully automatic multicriterial treatment planning (1620 plans in total). For each of the 15 sets of plans in the database, a KB model was trained with 54 patients and validated on the other 54 by comparing the predictions with the achieved doses. The dose prediction accuracy (predicted-achieved) of the KB models was assessed and compared among the different models to characterize the intrinsic performance and effect of interorgan dependency. In addition, the effect of dataset inconsistency with respect to planning prioritizations was investigated by mixing plans with different prioritizations, for the training, the validation dataset, and for both combined.RESULTSIn the case of a high planning priority, the mean ± SD of the prediction error for the mean dose of the parotid glands was only 0.2 ± 2.2 Gy, but this increased to 1.0 ± 5.0 Gy in the case that the parotid glands had a low planning priority. Dataset inconsistency (in planning priority) led to a large increase in prediction error for the parotid glands (mean ± SD) from 0.2 ± 2.2 Gy to 2.8 ± 3.3 Gy, -3.2 ± 5.0 Gy or -0.6 ± 5.4 Gy, depending on the way the datasets were mixed.CONCLUSIONSThe generated plan database can be used to validate and characterize KB prediction models for H&N cancer and will be made available upon request. The investigated KB model performed well in case the parotid glands had a high planning priority (little dependence on lower priority OARs), but poorly for organs for which the dose strongly depends on other higher priority OARs. To improve the performance of KB prediction models for H&N cancer, interorgan dependency should be modeled and accounted for. Dataset inconsistency has a large negative impact on the prediction errors of KB models and should be avoided as much as possible.
Background and purpose: With the advent of automatic treatment planning options like Pinnacle's Autoplanning (PAP), the challenge arises how to assess the quality of a plan that no dosimetrist did work on. The aim of this study was to assess plan quality consistency of PAP prostate cancer patients in clinical practice. Materials and methods: 100 prostate cancer patients were included from NKI and 129 from RadboudUMC (RUMC). Per institute a previously developed [1] treatment planning QA model, based on overlap volume histograms, was trained on PAP plans to predict achievable dose metrics which were then compared to the clinical PAP plans. A threshold of 3 Gy (DVH dose parameters)/3% (DVH volume parameters) was used to detect outliers. For the outlier plans, the PAP technique was adjusted with the aim of meeting the threshold. Results: The average difference between the prediction and the clinically achieved value was < 0.5 Gy (mean dose parameters) and < 1.2% (volume parameters), with standard deviation of 1.9 Gy/1.5% respectively. We found 8% (NKI)/25% (RUMC) of patients to exceed the 3 Gy/3% threshold, with deviations up to 6.7 Gy (mean dose rectum) and 6% (rectal wall V64Gy). In all cases the plans could be improved to fall within the thresholds, without compromising the other dose metrics. Conclusion: Independent treatment planning QA was used successfully to assess the quality of clinical PAP in a multi-institutional setting. Respectively 8% and 25% suboptimal clinical PAP plans were detected that all could be improved with replanning. Therefore we recommend the use of independent treatment plan QA in combination with PAP for prostate cancer patients. (C) 2018 Elsevier B.V. All rights reserved.
In the abdomen, it is challenging to assess the accuracy of deformable image registration (DIR) for individual patients, due to the lack of clear anatomical landmarks, which can hamper clinical applications that require high accuracy DIR, such as adaptive radiotherapy. In this study, we propose and evaluate a methodology for estimating the impact of uncertainties in DIR on calculated accumulated dose in the upper abdomen, in order to aid decision making in adaptive treatment approaches. Sixteen liver metastasis patients treated with SBRT were evaluated. Each patient had one planning and three daily treatment CT-scans. Each daily CT scan was deformably registered 132 times to the planning CT-scan, using a wide range of parameter settings for the registration algorithm. A subset of ‘realistic’ registrations was then objectively selected based on distances between mapped and target contours. The underlying 3D transformations of these registrations were used to assess the corresponding uncertainties in voxel positions, and delivered dose, with a focus on accumulated maximum doses in the hollow OARs, i.e. esophagus, stomach, and duodenum. The number of realistic registrations varied from 5 to 109, depending on the patient, emphasizing the need for individualized registration parameters. Considering for all patients the realistic registrations, the 99th percentile of the voxel position uncertainties was 5.6 ± 3.3 mm. This translated into a variation (difference between 1st and 99th percentile) in accumulated Dmax in hollow OARs of up to 3.3 Gy. For one patient a violation of the accumulated stomach dose outside the uncertainty band was detected. The observed variation in accumulated doses in the OARs related to registration uncertainty, emphasizes the need to investigate the impact of this uncertainty for any DIR algorithm prior to clinical use for dose accumulation. The proposed method for assessing on an individual patient basis the impact of uncertainties in DIR on accumulated dose is in principle applicable for all DIR algorithms allowing variation in registration parameters.
PURPOSE:To validate a novel deformable image registration (DIR) method for online adaptation of planning organ-at-risk (OAR) delineations to match daily anatomy during hypo-fractionated RT of abdominal tumors.MATERIALS AND METHODS:For 20 liver cancer patients, planning OAR delineations were adapted to daily anatomy using the DIR on corresponding repeat CTs. The DIR's accuracy was evaluated for the entire cohort by comparing adapted and expert-drawn OAR delineations using geometric (Dice Similarity Coefficient (DSC), Modified Hausdorff Distance (MHD) and Mean Surface Error (MSE)) and dosimetric (Dmax and Dmean) measures.RESULTS:For all OARs, DIR achieved average DSC, MHD and MSE of 86%, 2.1 mm, and 1.7 mm, respectively, within 20 s for each repeat CT. Compared to the baseline (translations), the average improvements ranged from 2% (in heart) to 24% (in spinal cord) in DSC, and 25% (in heart) to 44% (in right kidney) in MHD and MSE. Furthermore, differences in dose statistics (Dmax, Dmean and D2%) using delineations from an expert and the proposed DIR were found to be statistically insignificant (p > 0.01).CONCLUSION:The validated DIR showed potential for online-adaptive radiotherapy of abdominal tumors as it achieved considerably high geometric and dosimetric correspondences with the expert-drawn OAR delineations, albeit in a fraction of time required by experts.
PURPOSE:Treatment plans manually generated in clinical routine may suffer from variations and inconsistencies in quality. Using such plans for validating a DVH prediction algorithm might obscure its intrinsic prediction accuracy. In this study we used a recently published large database of Pareto-optimal prostate cancer plans to assess the prediction accuracy of a commercial knowledge-based DVH prediction algorithm, RapidPlan. The database plans were consistently generated with automated planning using an independent optimizer, and can be considered as aground truth of plan quality.METHODS:Prediction models were generated using training sets with 20, 30, 45, 55 and 114 Pareto-optimal plans. Model-20 and Model-30 were built using 5 groups of randomly selected training patients. For 60 independent Pareto-optimal validation plans, predicted and database DVHs were compared.RESULTS:For model-114, differences between predicted and database mean doses of more than ± 10% in rectum, anus and bladder, occurred for 23.3%, 55.0%, and 6.7% of the validation plans, respectively. For rectum V65Gy and V75Gy, differences outside the ±10% range were observed in 21.7% and 70.0% of validation plans, respectively. For 61.7% of validation plans, inaccuracies in predicted rectum DVHs resulted in a deviation in predicted NTCP for rectal bleeding outside ±10%. With smaller training sets the DVH prediction performance deteriorated, showing dependence on the selected training patients.CONCLUSION:Even when analysed with Pareto-optimal plans with highly consistent quality, clinically relevant deviations in DVH predictions were observed. Such deviations could potentially result in suboptimal plans for new patients. Further research on DVH prediction models is warranted.
PURPOSE:To prospectively investigate the use of an independent DVH prediction tool to detect outliers in the quality of fully automatically generated treatment plans for prostate cancer patients.MATERIALS/METHODS:A plan QA tool was developed to predict rectum, anus and bladder DVHs, based on overlap volume histograms and principal component analysis (PCA). The tool was trained with 22 automatically generated, clinical plans, and independently validated with 21 plans. Its use was prospectively investigated for 50 new plans by replanning in case of detected outliers.RESULTS:For rectum Dmean, V65Gy, V75Gy, anus Dmean, and bladder Dmean, the difference between predicted and achieved was within 0.4 Gy or 0.3% (SD within 1.8 Gy or 1.3%). Thirteen detected outliers were re-planned, leading to moderate but statistically significant improvements (mean, max): rectum Dmean (1.3 Gy, 3.4 Gy), V65Gy (2.7%, 4.2%), anus Dmean (1.6 Gy, 6.9 Gy), and bladder Dmean (1.5 Gy, 5.1 Gy). The rectum V75Gy of the new plans slightly increased (0.2%, p = 0.087).CONCLUSION:A high accuracy DVH prediction tool was developed and used for independent QA of automatically generated plans. In 28% of plans, minor dosimetric deviations were observed that could be improved by plan adjustments. Larger gains are expected for manually generated plans.
IMRT planning with commercial Treatment Planning Systems (TPSs) is a trial-and-error process. Consequently, the quality of treatment plans may not be consistent among patients, planners and institutions. Recently, different plan quality assurance (QA) models have been proposed, that could flag and guide improvement of suboptimal treatment plans. However, the performance of these models was validated using plans that were created using the conventional trail-and-error treatment planning process. Consequently, it is challenging to assess and compare quantitatively the accuracy of different treatment planning QA models. Therefore, we created a golden standard dataset of consistently planned Pareto-optimal IMRT plans for 115 prostate patients. Next, the dataset was used to assess the performance of a treatment planning QA model that uses the overlap volume histogram (OVH). 115 prostate IMRT plans were fully automatically planned using our in-house developed TPS Erasmus-iCycle. An existing OVH model was trained on the plans of 58 of the patients. Next it was applied to predict DVHs of the rectum, bladder and anus of the remaining 57 patients. The predictions were compared with the achieved values of the golden standard plans for the rectum D mean, V 65, and V 75, and D mean of the anus and the bladder. For the rectum, the prediction errors (predicted-achieved) were only -0.2 ± 0.9 Gy (mean ± 1 SD) for D mean,-1.0 ± 1.6% for V 65, and -0.4 ± 1.1% for V 75. For D mean of the anus and the bladder, the prediction error was 0.1 ± 1.6 Gy and 4.8 ± 4.1 Gy, respectively. Increasing the training cohort to 114 patients only led to minor improvements. A dataset of consistently planned Pareto-optimal prostate IMRT plans was generated. This dataset can be used to train new, and validate and compare existing treatment planning QA models, and has been made publicly available. The OVH model was highly accurate in predicting rectum and anus DVHs. For the bladder, larger prediction errors were observed.
______________________________________________________________________________________________________
IMRT treatment planning with commercial treatment planning systems is a trial-and-error process, based on a series of subjective human decisions. Therefore the quality of the IMRT treatment plans may not be consistent among patients, planners or institutions with different experience. Different plan quality assurance (QA) models have been proposed recently, that could flag suboptimal plans that may benefit from an additional treatment planning effort [1-7].
Purpose:Various studies have demonstrated that online adaptive radiotherapy by real‐time re‐optimization of the treatment plan can improve organs‐at‐risk (OARs) sparing in the abdominal region. Its clinical implementation, however, requires fast and accurate auto‐segmentation of OARs in CT scans acquired just before each treatment fraction. Autosegmentation is particularly challenging in the abdominal region due to the frequently observed large deformations. We present a clinical validation of a new auto‐segmentation method that uses fully automated non‐rigid registration for propagating abdominal OAR contours from planning to daily treatment CT scans.Methods:OARs were manually contoured by an expert panel to obtain ground truth contours for repeat CT scans (3 per patient) of 10 patients. For the non‐rigid alignment, we used a new non‐rigid registration method that estimates the deformation field by optimizing local normalized correlation coefficient with smoothness regularization. This field was used to propagate planning contours to repeat CTs. To quantify the performance of the auto‐segmentation, we compared the propagated and ground truth contours using two widely used metrics‐ Dice coefficient (Dc) and Hausdorff distance (Hd). The proposed method was benchmarked against translation and rigid alignment based auto‐segmentation.Results:For all organs, the auto‐segmentation performed better than the baseline (translation) with an average processing time of 15 s per fraction CT. The overall improvements ranged from 2% (heart) to 32% (pancreas) in Dc, and 27% (heart) to 62% (spinal cord) in Hd. For liver, kidneys, gall bladder, stomach, spinal cord and heart, Dc above 0.85 was achieved. Duodenum and pancreas were the most challenging organs with both showing relatively larger spreads and medians of 0.79 and 2.1 mm for Dc and Hd, respectively.Conclusion:Based on the achieved accuracy and computational time we conclude that the investigated auto‐segmentation method overcomes an important hurdle to the clinical implementation of online adaptive radiotherapy.Partial funding for this work was provided by Accuray Incorporated as part of a research collaboration with Erasmus MC Cancer Institute.
Background and purposeTo predict the lowest achievable rectum D35 for quality assurance of IMRT plans of prostate cancer patients.Materials and methodsFor each of 24 patients from a database of 47 previously treated patients, the anatomy was compared to the anatomies of the other 46 to predict the minimal achievable rectum D35. The 24 patients were then replanned to obtain maximally reduced rectum D35. Next, the newly derived plans were added to the database to replace the original clinical plans, and new predictions of the lowest achievable rectum D35 were made.ResultsAfter replanning, the rectum D35 reduced by 9.3Gy±6.1 (average±1 SD; p<0.001) compared to the original plan. The first predictions of the rectum D35 were 4.8Gy±4.2 (average±1 SD; p<0.001) too high when evaluated with the new plans. After updating the database, the replanned and newly predicted rectum D35 agreed within 0.1Gy±2.8 (average±1 SD; p=0.89). The doses to the bladder, anus and femoral heads did not increase compared to the original plans.ConclusionsFor individual prostate patients, the lowest achievable rectum D35 in IMRT planning can be accurately predicted from dose distributions of previously treated patients by quantitative comparison of patient anatomies. These predictions can be used to quantitatively assess the quality of IMRT plans.