BACKGROUND:Deep-learning neural network algorithms for detecting prostate cancer in MRI have proliferated in the literature. However, out of 30+ studies published since the PROSTATEx challenge, no studies tested the performance of their algorithm against using true external image data sets (studies came from an outside institution that did not supply any training data to the algorithm) while validating against MR-US fusion biopsy or whole-mount prostatectomy. Using true external data sets paints a much clearer picture of real-world clinical performance of an algorithm. PURPOSE:This work will assess the performance of a published deep learning (DL) neural network algorithm to detect prostate cancer using external studies. The main difference from other studies is the combination of using only MR-US fusion biopsy results as a gold standard; using test data from an institution that did not supply any training data for this version of the algorithm (including studies acquired with an endorectal coil, which were not in the original training set); and comparing the performance of algorithm-generated regions-of-interest (ROIs) versus algorithm heat maps. METHODS:Patients were included in the study if they had a prostate MRI with at least one radiologist-drawn target on MRI and underwent MR-US fusion biopsy where the target was sampled for pathological analysis. Patients were excluded if they had any history of prostate cancer treatment, had previously undergone MR-US fusion biopsy at our institution, were missing MRI acquisitions, had artifacts in image sets, or if the study had been shared for future algorithm development. MR image data was assessed using a DL research prototype (XProstate) from Siemens Healthineers that produced (a) ROIs in suspected cancer areas with a level of suspicion (LoS) score and (b) heat maps with LoS scores across the entire gland. The XProstate prototype had been trained with 2170 studies from eight different academic institutions. Clinical radiologist, XProstate ROI, and XProstate Heat Map scores were assessed with ROC analysis using pathology results from biopsy as a gold standard. RESULTS:202 unique patients were included for assessment of the XProstate research prototype. The ROC curve for the XProstate Heat Map LoS score generated the highest AUC (0.76, 95% CI: 0.70, 0.82) followed by clinical radiologist PI-RADS score (0.73, 95% CI: 0.68, 0.79) and by XProstate ROI LoS score (0.71, 95% CI: 0.65, 0.77). Neither the XProstate Heat Map (AUC difference = 0.03, 95% CI: -0.04, 0.10, p = 0.38) nor the XProstate ROI (AUC difference = -0.02, 95% CI: -0.09, 0.04, p = 0.43) was significantly different from the radiologist PI-RADS score. CONCLUSIONS:The XProstate prototype demonstrated equivalent performance as clinical radiologists when presented with de novo cases that would mirror a real-world clinical deployment. The automatic ROI delineation more closely matched clinical radiologist performance when using a cutoff of PI-RADS 5 for annotating suspicious regions. Overall, the XProstate prototype provided reasonable clinical performance and this study demonstrated the need to assess Deep Learning prototypes with external institutional test data.
A commercial MRI-based deep learning algorithm for prostate cancer detection showed greater positive predictive value, despite its lower sensitivity, potentially allowing it to assist radiologists in biopsy planning.
Background:A new diffusion-weighted imaging (DWI) technique, known as zoomed-field-of-view echo-planar DWI (z-DWI), has been developed to reduce geometric distortions and susceptibility artifacts and to achieve higher spatial resolution. However, it remains unclear whether z-DWI, compared with the traditional DWI technique, can enhance the diagnostic performance of deep-learning-based computer-aided diagnosis (DL-CAD) and radiologists using DL-CAD in detecting prostate cancer (PCa). This study aims to evaluate and compare the diagnostic performance and PI-RADS scores of DL-CAD in detecting PCa using conventional full-field-of-view single-shot echo-planar DWI (f-DWI) and advanced z-DWI and to extend this comparison to clinical practice, in which radiologists use DL-CAD. Methods:This study retrospectively included magnetic resonance imaging from 359 patients for suspected PCa. There were 496 prostate lesions included, with 253 (51%) being malignant. Using a DL-CAD system, images of f-DWI and z-DWI were uploaded separately to obtain the localizations and the prostate imaging reporting and data system (PI-RADS) scores of suspected malignant lesions. The results were compared to histopathologic results. The diagnostic performance of f-DWI and z-DWI were evaluated using the free-response receiver operating characteristics and the alternative free-response receiver operating characteristics curves. Discrepancies in PI-RADS scores were analyzed. Additionally, two radiologists participated in consensus reading images by using DL-CAD with different DWI techniques, and their performance and PI-RADS scores were compared. Lastly, the relationship between PI-RADS discrepancies and clinically significant prostate cancer (csPCa) risk was analyzed. Results:z-DWI enabled DL-CAD to exhibit better diagnostic performance [area under the curve (AUC), 0.857 vs. 0.841; P=0.02], with a higher mean PI-RADS score for PCa lesions (4.26 vs. 3.92; P<0.001), and improved scores for 66 PCa lesions compared to f-DWI. When radiologists used DL-CAD, z-DWI also enabled radiologists to exhibit a higher mean PI-RADS score for PCa lesions (4.31 vs. 4.02; P<0.001) and improved scores for 56 PCa lesions compared to f-DWI, however, no statistically significant difference was found in diagnostic performance (AUC, 0.887 vs. 0.881; P=0.16). In multivariable logistic regression analyses, upgraded PI-RADS scores by z-DWI were significantly associated with csPCa risk. Conclusions:z-DWI, in comparison to f-DWI, enhances the diagnostic performance of DL-CAD for PCa, assigning higher PI-RADS scores to malignant lesions. Despite offering limited improvement for radiologists using DL-CAD, z-DWI shows promise in enhancing the detection of csPCa.
Our hypothesis is that UDA using diffusion-weighted images, generated with a unified model, offers a promising and reliable strategy for enhancing the performance of supervised learning models in multi-site prostate lesion detection, especially when various b-values are present. This retrospective study included data from 5,150 patients (14,191 samples) collected across nine different imaging centers. A novel UDA method using a unified generative model was developed for multi-site PCa detection. This method translates diffusion-weighted imaging (DWI) acquisitions, including apparent diffusion coefficient (ADC) and individual DW images acquired using various b-values, to align with the style of images acquired using b-values recommended by Prostate Imaging Reporting and Data System (PI-RADS) guidelines. The generated ADC and DW images replace the original images for PCa detection. An independent set of 1,692 test cases (2,393 samples) was used for evaluation. The area under the receiver operating characteristic curve (AUC) was used as the primary metric, and statistical analysis was performed via bootstrapping. For all test cases, the AUC values for baseline SL and UDA methods were 0.73 and 0.79 (p<.001), respectively, for PI-RADS>=3, and 0.77 and 0.80 (p<.001) for PI-RADS>=4 PCa lesions. In the 361 test cases under the most unfavorable image acquisition setting, the AUC values for baseline SL and UDA were 0.49 and 0.76 (p<.001) for PI-RADS>=3, and 0.50 and 0.77 (p<.001) for PI-RADS>=4 PCa lesions. The results indicate the proposed UDA with generated images improved the performance of SL methods in multi-site PCa lesion detection across datasets with various b values, especially for images acquired with significant deviations from the PI-RADS recommended DWI protocol (e.g. with an extremely high b-value).
Deep learning has been utilized in knowledge-based radiotherapy planning in which a system trained with a set of clinically approved plans is employed to infer a three-dimensional dose map for a given new patient. However, previous deep methods are primarily limited to simple scenarios, e.g., a fixed planning type or a consistent beam angle configuration. This in fact limits the usability of such approaches and makes them not generalizable over a larger set of clinical scenarios. Herein, we propose a novel conditional generative model, Flexible-C^m GAN, utilizing additional information regarding planning types and various beam geometries. A miss-consistency loss is proposed to deal with the challenge of having a limited set of conditions on the input data, e.g., incomplete training samples. To address the challenges of including clinical preferences, we derive a differentiable shift-dose-volume loss to incorporate the well-known dose-volume histogram constraints. During inference, users can flexibly choose a specific planning type and a set of beam angles to meet the clinical requirements. We conduct experiments on an illustrative face dataset to show the motivation of Flexible-C^m GAN and further validate our model's potential clinical values with two radiotherapy datasets. The results demonstrate the superior performance of the proposed method in a practical heterogeneous radiotherapy planning application compared to existing deep learning-based approaches.
Deep-learning-based computer-aided diagnosis (DL-CAD) systems using MRI for prostate cancer (PCa) detection have demonstrated good performance. Nevertheless, DL-CAD systems are vulnerable to high heterogeneities in DWI, which can interfere with DL-CAD assessments and impair performance. This study aims to compare PCa detection of DL-CAD between zoomed-field-of-view echo-planar DWI (z-DWI) and full-field-of-view DWI (f-DWI) and find the risk factors affecting DL-CAD diagnostic efficiency. This retrospective study enrolled 354 consecutive participants who underwent MRI including T2WI, f-DWI, and z-DWI because of clinically suspected PCa. A DL-CAD was used to compare the performance of f-DWI and z-DWI both on a patient level and lesion level. We used the area under the curve (AUC) of receiver operating characteristics analysis and alternative free-response receiver operating characteristics analysis to compare the performances of DL-CAD using f- DWI and z-DWI. The risk factors affecting the DL-CAD were analyzed using logistic regression analyses. P values less than 0.05 were considered statistically significant. DL-CAD with z-DWI had a significantly better overall accuracy than that with f-DWI both on patient level and lesion level (AUCpatient: 0.89 vs. 0.86; AUClesion: 0.86 vs. 0.76; P < .001). The contrast-to-noise ratio (CNR) of lesions in DWI was an independent risk factor of false positives (odds ratio [OR] = 1.12; P < .001). Rectal susceptibility artifacts, lesion diameter, and apparent diffusion coefficients (ADC) were independent risk factors of both false positives (ORrectal susceptibility artifact = 5.46; ORdiameter, = 1.12; ORADC = 0.998; all P < .001) and false negatives (ORrectal susceptibility artifact = 3.31; ORdiameter = 0.82; ORADC = 1.007; all P ≤ .03) of DL-CAD. Z-DWI has potential to improve the detection performance of a prostate MRI based DL-CAD. ChiCTR, NO. ChiCTR2100041834 . Registered 7 January 2021.
PURPOSE We developed a deep neural network that queries the lung computed tomography–derived feature space to identify radiation sensitivity parameters that can predict treatment failures and hence guide the individualization of radiotherapy dose. In this article, we examine the transportability of this model across health systems. METHODS This multicenter cohort-based registry included 1,120 patients with cancer in the lung treated with stereotactic body radiotherapy. Pretherapy lung computed tomography images from the internal study cohort (n = 849) were input into a multitask deep neural network to generate an image fingerprint score that predicts time to local failure. Deep learning (DL) scores were input into a regression model to derive iGray, an individualized radiation dose estimate that projects a treatment failure probability of < 5% at 24 months. We validated our findings in an external, holdout cohort (n = 271). RESULTS There were substantive differences in the baseline patient characteristics of the two study populations, permitting an assessment of model transportability. In the external cohort, radiation treatments in patients with high DL scores failed at a significantly higher rate with 3-year cumulative incidences of local failure of 28.5% (95% CI, 19.8 to 37.8) versus 10.2% (95% CI, 5.9 to 16.2; hazard ratio, 3.3 [95% CI, 1.74 to 6.49]; P < .001). A model that included DL score alone predicted treatment failures with a concordance index of 0.68 (95% CI, 0.59 to 0.77), which had a similar performance to a nested model derived from within the internal cohort (0.70 [0.64 to 0.75]). External cohort patients with iGray values that exceeded the delivered doses had proportionately higher rates of local failure ( P < .001). CONCLUSION Our results support the development and implementation of new DL-guided treatment guidance tools in the image-replete and highly standardized discipline of radiation oncology.
Purpose The aim of this study was to evaluate the accuracy of prostate volume estimates calculated from the ellipsoid formula using the anteroposterior (AP) diameter measured on axial and sagittal images obtained through ultrasonography (US) and magnetic resonance imaging (MRI). Methods This retrospective study included 456 patients with transrectal US and MRI from two university hospitals. Two radiologists independently measured the prostate gland diameters on US and MRI: AP diameters on axial and sagittal images, transverse, and longitudinal diameters on midsagittal images. The volume estimates, volumeax and volumesag, were calculated from the ellipsoid formula by using the AP diameter on axial and sagittal images, respectively. The prostate volume extracted from MRI-based whole-gland segmentation was considered the gold standard. The intraclass correlation coefficient (ICC) was used to evaluate the inter-method agreement between volumeax and volumesag, and agreement with the gold standard. The Wilcoxon signedrank test was used to analyze the differences between the volume estimates and the gold standard. Results The prostate gland volume estimates showed excellent inter-method agreement, and excellent agreement with the gold standard (ICCs >0.9). Compared with the gold standard, the volume estimates were significantly larger on MRI and significantly smaller on US (P<0.001). The volume difference (segmented volume–volume estimate) was greater in patients with larger prostate glands, especially on US. Conclusion Volumeax and volumesag showed excellent inter-method agreement and excellent agreement with the gold standard on both US and MRI. However, prostate volume was overestimated on MRI and underestimated on US.
Purpose: The Prostate Imaging Reporting and Data System (PI-RADS) was introduced to standardize prostate cancer diagnosis by MRI. However, the inter-reader agreement by PI-RADS scoring is not always high. The purpose of this study was to validate a deep-learning-based diagnostic algorithm of PI-RADS. Methods: We applied a Siemens Healthineers Prostate Artificial Intelligence (AI) prototype (work in progress) for fully automated prostate lesion detection, classification and reporting. More than 2000 bi-parametric MRI studies along with the PI-RADS reports were included as training, validation, and test data. This prospective validation study includes 101 consecutive patients suspected of prostate cancer, and 100 patients were included in the analysis. All subjects underwent a noncontrast-enhanced bi-parametric MRI including T2-weighted and diffusion-weighted imaging. Two board-certified radiologists independently scored the PI-RADS, and if there were disagreements; another radiologist confirmed the diagnosis. We compared the AI results with the interpretation results by the radiologists. Results: The sensitivity of our AI model for PI-RADS ≥ 4 was 0.76, and the specificity was 0.76. For the cases with PI-RADS ≥ 3, the sensitivity was 0.69, and the specificity was 0.76. In the lesion-based analysis, AI detection rates of PI-RADS 3, 4, 5 lesions in the peripheral zone were 43%, 63%, and 100%, respectively. In the transition zone, AI detection rates of PI-RADS 3, 4, 5 were 30%, 54%, and 100%, respectively. Conclusion: Our deep-learning-based algorithm has been validated and shown to help score PI-RADS.
Abstract Introduction Deep learning (DL) models that use medical images to predict clinical outcomes are poised for clinical translation. For tumors that reside in organs that move, however, the impact of motion (i.e., degenerated object appearance or blur) on DL model accuracy remains unclear. We examine the impact of tumor motion on an image‐based DL framework that predicts local failure risk after lung stereotactic body radiotherapy (SBRT). Methods We input pre‐therapy free breathing (FB) computed tomography (CT) images from 849 patients treated with lung SBRT into a multitask deep neural network to generate an image fingerprint signature (or DL score) that predicts time‐to‐event local failure outcomes. The network includes a convolutional neural network encoder for extracting imaging features and building a task‐specific fingerprint, a decoder for estimating handcrafted radiomic features, and a task‐specific network for generating image signature for radiotherapy outcome prediction. The impact of tumor motion on the DL scores was then examined for a holdout set of 468 images from 39 patients comprising: (1) FB CT, (2) four‐dimensional (4D) CT, and (3) maximum‐intensity projection (MIP) images. Tumor motion was estimated using a 3D vector of the maximum distance traveled, and its association with DL score variance was assessed by linear regression. Findings The variance and amplitude in 4D CT image‐derived DL scores were associated with tumor motion (R 2 = 0.48 and 0.46, respectively). Specifically, DL score variance was deterministic and represented by sinusoidal undulations in phase with the respiratory cycle. DL scores, but not tumor volumes, peaked near end‐exhalation. The mean of the scores derived from 4D CT images and the score obtained from FB CT images were highly associated (Pearson r = 0.99). MIP‐derived DL scores were significantly higher than 4D‐ or FB‐derived risk scores (p < 0.0001). Interpretation An image‐based DL risk score derived from a series of 4D CT images varies in a deterministic, sinusoidal trajectory in a phase with the respiratory cycle. These results indicate that DL models of tumors in motion can be robust to fluctuations in object appearance due to movement and can guide standardization processes in the clinical translation of DL models for patients with lung cancer.
Artificial intelligence-based prostate cancer (PCa) detection models have been widely explored to assist clinical diagnosis. However, these trained models may generate erroneous results specifically on datasets that are not within training distribution. In this paper, we propose an approach to tackle this so-called out-of-distribution (OOD) data problem. Specifically, we devise an end-to-end unsupervised framework to estimate uncertainty values for cases analyzed by a previously trained PCa detection model. Our PCa detection model takes the inputs of bpMRI scans and through our proposed approach we identify OOD cases that are likely to generate degraded performance due to the data distribution shifts. The proposed OOD framework consists of two parts. First, an autoencoder-based reconstruction network is proposed, which learns discrete latent representations of in-distribution data. Second, the uncertainty is computed using perceptual loss that measures the distance between original and reconstructed images in the feature space of a pre-trained PCa detection network. The effectiveness of the proposed framework is evaluated on seven independent data collections with a total of 1,432 cases. The performance of pre-trained PCa detection model is significantly improved by excluding cases with high uncertainty.
Advances in tissue analysis methods, image analysis, high-throughput molecular profiling, and computational tools increasingly allow us to capture and quantify patient-to-patient variations that impact cancer risk, prognosis, and treatment response. Statistical models that integrate patient-specific information from multiple sources (e.g., family history, demographics, germline variants, imaging features) can provide individualized cancer risk predictions that can guide screening and prevention strategies. The precision, quality, and standardization of diagnostic imaging are improving through computer-aided solutions, and multigene prognostic and predictive tests improved predictions of prognosis and treatment response in various cancer types. A common theme across many of these advances is that individually moderately informative variables are combined into more accurate multivariable prediction models. Advances in machine learning and the availability of large data sets fuel rapid progress in this field. Molecular dissection of the cancer genome has become a reality in the clinic, and molecular target profiling is now routinely used to select patients for various targeted therapies. These technology-driven increasingly more precise and quantitative estimates of benefit versus risk from a given intervention empower patients and physicians to tailor treatment strategies that match patient values and expectations.
Better models to identify individuals at low risk of ventricular arrhythmia (VA) are needed for implantable cardioverter-defibrillator (ICD) candidates to mitigate the risk of ICD-related complications. We designed the CERTAINTY study (CinE caRdiac magneTic resonAnce to predIct veNTricular arrhYthmia) with deep learning for VA risk prediction from cine cardiac magnetic resonance (CMR). Using a training cohort of primary prevention ICD recipients (n = 350, 97 women, median age 59 years, 178 ischemic cardiomyopathy) who underwent CMR immediately prior to ICD implantation, we developed two neural networks: Cine Fingerprint Extractor and Risk Predictor. The former extracts cardiac structure and function features from cine CMR in a form of cine fingerprint in a fully unsupervised fashion, and the latter takes in the cine fingerprint and outputs disease outcomes as a cine risk score. Patients with VA (n = 96) had a significantly higher cine risk score than those without VA. Multivariate analysis showed that the cine risk score was significantly associated with VA after adjusting for clinical characteristics, cardiac structure and function including CMR-derived scar extent. These findings indicate that non-contrast, cine CMR inherently contains features to improve VA risk prediction in primary prevention ICD candidates. We solicit participation from multiple centers for external validation.
Purpose: To compare the performance of lesion detection and Prostate Imaging-Reporting and Data System (PIRADS) classification between a deep learning-based algorithm (DLA), clinical reports and radiologists with different levels of experience in prostate MRI. Methods: This retrospective study included 121 patients who underwent prebiopsy MRI and prostate biopsy. More than five radiologists (Reader groups 1, 2: residents; Readers 3, 4: less-experienced radiologists; Reader 5: expert) independently reviewed biparametric MRI (bpMRI). The DLA results were obtained using bpMRI. The reference standard was based on pathologic reports. The diagnostic performance of the PI-RADS classification of DLA, clinical reports, and radiologists was analyzed using AUROC. Dichotomous analysis (PI-RADS cutoff value > 3 or 4) was performed, and the sensitivities and specificities were compared using McNemar's test. Results: Clinically significant cancer [CSC, Gleason score > 7] was confirmed in 43 patients (35.5%). The AUROC of the DLA (0.828) for diagnosing CSC was significantly higher than that of Reader 1 (AUROC, 0.706; p = 0.011), significantly lower than that of Reader 5 (AUROC, 0.914; p = 0.013), and similar to clinical reports and other readers (p = 0.060-0.661). The sensitivity of DLA (76.7%) was comparable to those of all readers and the clinical reports at a PI-RADS cutoff value > 4. The specificity of the DLA (85.9%) was significantly higher than those of clinical reports and Readers 2-3 and comparable to all others at a PI-RADS cutoff value > 4. Conclusions: The DLA showed moderate diagnostic performance at a level between those of residents and an expert in detecting and classifying according to PI-RADS. The performance of DLA was similar to that of clinical reports from various radiologists in clinical practice.
Objective The aim of this study was to evaluate the effect of a deep learning based computer-aided diagnosis (DL-CAD) system on radiologists' interpretation accuracy and efficiency in reading biparametric prostate magnetic resonance imaging scans. Materials and Methods We selected 100 consecutive prostate magnetic resonance imaging cases from a publicly available data set (PROSTATEx Challenge) with and without histopathologically confirmed prostate cancer. Seven board-certified radiologists were tasked to read each case twice in 2 reading blocks (with and without the assistance of a DL-CAD), with a separation between the 2 reading sessions of at least 2 weeks. Reading tasks were to localize and classify lesions according to Prostate Imaging Reporting and Data System (PI-RADS) v2.0 and to assign a radiologist's level of suspicion score (scale from 1–5 in 0.5 increments; 1, benign; 5, malignant). Ground truth was established by consensus readings of 3 experienced radiologists. The detection performance (receiver operating characteristic curves), variability (Fleiss κ), and average reading time without DL-CAD assistance were evaluated. Results The average accuracy of radiologists in terms of area under the curve in detecting clinically significant cases (PI-RADS ≥4) was 0.84 (95% confidence interval [CI], 0.79–0.89), whereas the same using DL-CAD was 0.88 (95% CI, 0.83–0.94) with an improvement of 4.4% (95% CI, 1.1%–7.7%; P = 0.010). Interreader concordance (in terms of Fleiss κ) increased from 0.22 to 0.36 (P = 0.003). Accuracy of radiologists in detecting cases with PI-RADS ≥3 was improved by 2.9% (P = 0.10). The median reading time in the unaided/aided scenario was reduced by 21% from 103 to 81 seconds (P < 0.001). Conclusions Using a DL-CAD system increased the diagnostic accuracy in detecting highly suspicious prostate lesions and reduced both the interreader variability and the reading time.
Multi-parametric MRI (mp-MRI) has recently been established in major guidelines as a first-line diagnostic test for men suspected of having prostate cancer (PCa) primarily to detect and classify clinically significant lesions. However, widespread utilization is still challenged by 1) the difficulty of interpretation specifically for radiologists less experienced in reading mp-MRI scans, and 2) decreased productivity associated with increased time spent per case for reading these complex scans. Deep learning based lesion detection and segmentation methods have been proposed for radiologists to perform their tasks more accurately and efficiently. In this work, we present a novel panoptic lesion detection and segmentation method with both semantic and instance branches as well as an attention module to optimally incorporate both local and global image features. In a free-response receiver operating characteristics (FROC) analysis for lesion sensitivity on an independent dataset with 243 patients, our method has achieved 89% sensitivity and 85% with 0.94 and 0.62 false positives per patient, respectively. Using the proposed method, we have achieved an unprecedented area under ROC curve (AUC) of 0.897 in identifying clinically significant cases.
Background: Opportunistic prostate cancer (PCa) screening is a controversial topic. Magnetic resonance imaging (MRI) has proven to detect prostate cancer with a high sensitivity and specificity, leading to the idea to perform an image-guided prostate cancer (PCa) screening; Methods: We evaluated a prospectively enrolled cohort of 49 healthy men participating in a dedicated image-guided PCa screening trial employing a biparametric MRI (bpMRI) protocol consisting of T2-weighted (T2w) and diffusion weighted imaging (DWI) sequences. Datasets were analyzed both by human readers and by a fully automated artificial intelligence (AI) software using deep learning (DL). Agreement between the algorithm and the reports—serving as the ground truth—was compared on a per-case and per-lesion level using metrics of diagnostic accuracy and k statistics; Results: The DL method yielded an 87% sensitivity (33/38) and 50% specificity (5/10) with a k of 0.42. 12/28 (43%) Prostate Imaging Reporting and Data System (PI-RADS) 3, 16/22 (73%) PI-RADS 4, and 5/5 (100%) PI-RADS 5 lesions were detected compared to the ground truth. Targeted biopsy revealed PCa in six participants, all correctly diagnosed by both the human readers and AI. Conclusions: The results of our study show that in our AI-assisted, image-guided prostate cancer screening the software solution was able to identify highly suspicious lesions and has the potential to effectively guide the targeted-biopsy workflow.
Prostate cancer (PCa) is the most prevalent and one of the leading causes of cancer death among men. Multi-parametric MRI (mp-MRI) is a prominent diagnostic scan, which could help in avoiding unnecessary biopsies for men screened for PCa. Artificial intelligence (AI) systems could help radiologists to be more accurate and consistent in diagnosing clinically significant cancer from mp-MRI scans. Lack of specificity has been identified recently as one of weak points of such assistance systems. In this paper, we propose a novel false positive reduction network to be added to the overall detection system to further analyze lesion candidates. The new network utilizes multiscale 2D image stacks of these candidates to discriminate between true and false positive detections. We trained and validated our network on a dataset with 2170 cases from seven different institutions and tested it on a separate independent dataset with 243 cases. With the proposed model, we achieved area under curve (AUC) of 0.876 on discriminating between true and false positive detected lesions and improved the AUC from 0.825 to 0.867 on overall identification of clinically significant cases.
Our results indicate that there are image-distinct subpopulations that have differential sensitivity to radiotherapy. The image-based deep learning framework proposed herein is the first opportunity to use medical images to individualize radiotherapy dose.