PURPOSE To assess diagnostic performance of digital breast tomosynthesis (DBT) alone or combined with technologist-performed handheld screening ultrasound (US) in women with dense breasts. METHODS In an institutional review board–approved, Health Insurance Portability and Accountability Act–compliant multicenter protocol in western Pennsylvania, 6,179 women consented to three rounds of annual screening, interpreted by two radiologist observers, and had appropriate follow-up. Primary analysis was based on first observer results. RESULTS Mean participant age was 54.8 years (range, 40-75 years). Across 17,552 screens, there were 126 cancer events in 125 women (7.2/1,000; 95% CI, 5.9 to 8.4). In year 1, DBT-alone cancer yield was 5.0/1,000, and of DBT+US, 6.3/1,000, difference 1.3/1,000 (95% CI, 0.3 to 2.1; P = .005). In years 2 + 3, DBT cancer yield was 4.9/1,000, and of DBT+US, 5.9/1,000, difference 1.0/1,000 (95% CI, 0.4 to 1.5; P < .001). False-positive rate increased from 7.0% for DBT in year 1 to 11.5% for DBT+US and from 5.9% for DBT in year 2 + 3 to 9.7% for DBT+US ( P < .001 for both). Nine cancers were seen only by double reading DBT and one by double reading US. Ten interval cancers (0.6/1,000 [95% CI, 0.2 to 0.9]) were identified. Despite reduction in specificity, addition of US improved receiver operating characteristic curves, with area under receiver operating characteristic curve increasing from 0.83 for DBT alone to 0.92 for DBT+US in year 1 ( P = .01), with smaller improvements in subsequent years. Of 6,179 women, across all 3 years, 172/6,179 (2.8%) unique women had a false-positive biopsy because of DBT as did another 230/6,179 (3.7%) women because of US ( P < .001). CONCLUSION Overall added cancer detection rate of US screening after DBT was modest at 19/17,552 (1.1/1,000; CI, 0.5- to 1.6) screens but potentially overcomes substantial increases in false-positive recalls and benign biopsies.
Interpretations of breast ultrasound screening examinations result in high recall rates and large inter-radiologist variability, frequently leading to “conservative” recommendations. Double reading of all breast ultrasound screening examinations is cost prohibitive, but double reading of only “initially recalled” cases may prove efficacious. We assessed changes in recommendations, if any, by providing a consensus second opinion in a limited subset of examinations initially recommended for recall. We performed a retrospective reader study with 197 ultrasound examinations (97 not recalled and 100 recalled clinically). First, we generated a consensus “second opinion” consisting of the majority vote of three independent readings of each case by experienced ultrasound interpreters. During the reader study that followed, if the reader recommended a “recall” and the “consensus second opinion” did not, a message to that effect was displayed and the reader was asked to re-review the exam and re-assess if, knowing the second opinion, a re-rating of the case was warranted. We compared performance levels pre- and post- the second opinion. The second opinion resulted in “no recall” recommendations of 141 cases in the entire set, including four cancer cases missed by all three readers. On average, radiologists received “warning” messages in 30 cases (range 15-50), or in ~15% of cases. Rating changes (downgrades to no recall) occurred in 36 of these cases. These changes resulted in a possible recall rate reduction of 28% in prompted cases or 14% overall recall reduction, while increasing the false negative rate by only one case missed by 2 readers (~1%).
Aim: The aim of this study was to examine the factors influencing count density in low-dose molecular breast imaging and their impact on image quality. Methods: One hundred patients scheduled for a diagnostic MBI procedure were imaged using a commercially available MBI system following the SNM Practice Guideline for Breast Scintigraphy, with one modification. Each 20 mCi (740 MBq) diagnostic dose of Tc-99m sestamibi was separated into two syringes, with each containing 10 mCi (370 MBq) or with one containing 5 mCi (185 MBq) and the other containing 15 mCi (555 MBq). Patients were randomly injected with 5 mCi, 10 mCi, or 15 mCi and imaged bilaterally in the craniocaudal (CC) view. The remaining fraction of the 20 mCi was then injected, followed by a standard four-view study. Results: The two sets of CC view images were analyzed to determine the average count density for a given acquisition time and injected activity, after application of an effective half-life correction factor to account for the time-related combined effect of radioactive decay and cellular washout. The average count density of the MBI images was reduced by radiopharmaceutical decay and washout and imaging time should be adjusted accordingly in low dose imaging. Conclusions : It was observed that the patient weight was one physical characteristic that should be considered for optimal image quality, with the dose increasing in proportion to the patient weight.
OBJECTIVEVariations in the thickness of a compressed breast and the resulting variations in mammographic densities confound current automated procedures for estimating tissue composition of breasts from digitized mammograms. We sought to determine whether adjusting mammographic data for tissue thickness before estimating tissue composition could improve the accuracy of the tissue estimates.MATERIALS AND METHODSWe developed methods for locally estimating breast thickness from mammograms and then adjusting pixel values so that the values correlated with the tissue composition over the breast area. In our technique, the pixel values are corrected for the nonlinearity of the combined characteristic curve from the film and film digitizer; the approximate relative thickness as a function of distance from the skin line is measured; and the pixel values are adjusted to reflect their distance from the skin line. To estimate tissue composition, we created a backpropagation neural network classifier from features extracted from the histogram of pixel values, after the data had been adjusted for characteristic curve and tissue thickness. We used a 10-fold cross-validation method to evaluate the neural network. The averaged scores of three radiologists were our gold standard.RESULTSThe performance of the neural network was calculated as the percentage of correct classifications of images that were or were not corrected to reflect tissue thickness. With its parameters derived from the pixel-value histogram, the neural network based on corrected images performed better (71% accuracy) than that based on uncorrected images (67% accuracy) (p < 0.05).CONCLUSIONOur results show that adjusting tissue thickness before estimating tissue composition improved the performance of our estimation procedure in reproducing the tissue composition values determined by radiologists.
RATIONALE AND OBJECTIVES:The authors developed a computerized method for the quantitative assessment of breast tissue composition on digitized mammograms.MATERIALS AND METHODS:Three radiologists were asked to review 200 digitized mammograms and independently provide a Breast Imaging Reporting and Data System-like rating for breast tissue composition on a scale of 0 to 4. These values were incorporated into a "consensus" rating that was used as a reference point in the development and evaluation of a computerized method. After tissue segmentation that excluded nontissue areas, a set of quantitative features was computed. A computerized summary index that attempts to reproduce the radiologists' ratings was developed. Correlation coefficients (Pearson r) were used to compare the computerized index with the consensus ratings.RESULTS:Some individual features computed for the relatively dense breast areas showed good correlation (r > 0.8) with the radiologists' subjective ratings. The summary index of tissue composition demonstrated a significant correlation (r = 0.87), as well.CONCLUSION:Computerized methods that show good correlation with radiologists' ratings of breast tissue composition can be developed.
OBJECTIVE. We used receiver operating characteristic (ROC) analysis to compare two methods of evaluating observer performance in detecting an abnormality on chest radiographs. In the first method, the abnormality in question, rib fracture, was one of five investigated, and it was the only one of interest in the second.MATERIALS AND METHODS. Eight experienced observers viewed 117 posteroanterior chest radiographs in two interpretation modes. Fifty-four of these images depicted rib fractures that had been rated as subtle for detection. The likelihood of the presence of a rib fracture was rated as one of five abnormalities in question in one mode and the sole abnormality of interest in the other mode.RESULTS. Six of the observers performed better during the single-abnormality mode, one performed equally well in both modes, and one performed better during the multiple-abnormality mode. The average area under the ROC curves (A(Z)) was 0.73 +/- 0.07 for the multiple-abnormality mode and 0.80 +/- 0.04 for the single-abnormality mode. The results were significantly different (p < 0.05).CONCLUSION. Study methodology can significantly affect the results in ROC studies, particularly for abnormalities that may not be perceived as primary or important. The order in which abnormalities appear on a checklist report form may be important.
Small, efficient observer performance study methods, such as multi-point rank order and forced choice, can be used to either determine the necessity to perform a large scale, receiver operator characteristic (ROC) study or to set the boundary conditions at which it makes sense to perform an ROC study. Because these studies often require observers to discriminate between (among) modes having small visual differences, we decided to address the issue of observers' abilities to make distinctions between modes in these types of non-ROC studies. In this project we reviewed the data of six different non-ROC studies that included a total of 24 observers to determine whether some behave as better mode discriminators. Because of the actual small differences, most observers performed poorly in identifying differences between or among modes. At the same time, at least one and sometimes more observers could identify extremely small differences between modes. These differences were statistically significant. Our results indicate that a good mode discriminator in one study may not perform as well in another such study. Non-ROC studies can be highly sensitive to differences between modes. However, large differences in observer performance combined with observer inconsistency across studies necessitates that these studies include multiple observers.
PURPOSE:To assess the performance of radiologists in the detection of masses and microcalcification clusters on digitized mammograms by using different computer-assisted detection (CAD) cuing environments.MATERIALS AND METHODS:Two hundred nine digitized mammograms depicting 57 verified masses and 38 microcalcification clusters in 85 positive and 35 negative cases were interpreted independently by seven radiologists using five display modes. Except for the first mode, for which no CAD results were provided, suspicious regions identified with a CAD scheme were cued in all the other modes by using a combination of two cuing sensitivities (90% and 50%) and two false-positive rates (0.5 and 2.0 per image). A receiver operating characteristic study was performed by using soft-copy images.RESULTS:CAD cuing at 90% sensitivity and a rate of 0.5 false-positive region per image improved observer performance levels significantly (P < .01). As accuracy of CAD cuing decreased so did observer performances (P < .01). Cuing specificity affected mass detection more significantly, while cuing sensitivity affected detection of microcalcification clusters more significantly (P < .01). Reduction of cuing sensitivity and specificity significantly increased false-negative rates in noncued areas (P < .05). Trends were consistent for all observers.CONCLUSION:CAD systems have the potential to significantly improve diagnostic performance in mammography. However, poorly performing schemes could adversely affect observer performance in both cued and noncued areas.
The purpose of this work was to develop and evaluate a computer-aided detection (CAD) scheme for the improvement of mass identification on digitized mammograms using a knowledge-based approach. Three hundred pathologically verified masses and 300 negative, but suspicious, regions, as initially identified by a rule-based CAD scheme, were randomly selected from a large clinical database for development purposes. In addition, 500 different positive and 500 negative regions were used to test the scheme. This suspicious region pruning scheme includes a learning process to establish a knowledge base that is then used to determine whether a previously identified suspicious region is likely to depict a true mass. This is accomplished by quantitatively characterizing the set of known masses, measuring "similarity" between a suspicious region and a "known" mass, then deriving a composite "likelihood" measure based on all "known" masses to determine the state of the suspicious region. To assess the performance of this method, receiver-operating characteristic (ROC) analyses were employed. Using a leave-one-out validation method with the development set of 600 regions, the knowledge-based CAD scheme achieved an area under the ROC curve of 0.83. Fifty-one percent of the previously identified false-positive regions were eliminated, while maintaining 90% sensitivity. During testing of the 1,000 independent regions, an area under the ROC curve as high as 0.80 was achieved. Knowledge-based approaches can yield a significant reduction in false-positive detections while maintaining reasonable sensitivity. This approach has the potential of improving the performance of other rule-based CAD schemes.
PURPOSE: To compare the cost of magnetic resonance (MR) imaging and its ability to direct the use of lymph node dissection with the cost and ability of conventional surgery for the staging of endometrial carcinoma.MATERIALS AND METHODS: Preoperative MR images of 25 patients who underwent hysterectomy for endometrial carcinoma were retrospectively evaluated. MR imaging results were compared with those of intraoperative gross dissection of the uterus and final histopathologic examination. Medicare reimbursements for two scenarios were compared in each patient. In the MR imaging scenario, the necessity for lymph node dissection was based on MR imaging results and histologic findings at biopsy. In the actual scenario, lymph node dissection was performed at the surgeon's discretion on the basis of findings at gross dissection of the uterus and histologic examination at biopsy.RESULTS: The cost of the MR imaging scenario, as defined by Medicare reimbursements, was 1% ($1,265/$148,500) less than that of the actual scenario. In the MR imaging scenario, all patients who required lymph node dissection received it, and 86% of the lymph node dissections performed were necessary. In the actual scenario, one necessary lymph node dissection was not performed, and only 31% of the lymph node dissections performed were necessary.CONCLUSION: Staging with MR imaging has costs and accuracy similar to those of the current method of staging with intraoperative gross dissection of the uterus. In addition, MR imaging decreases the number of unnecessary lymph node dissections.
Rationale and Objectives. The authors compare a 43-mu m computed radiographic system with a mammographic screen-film system for detection of simulated microcalcifications in an observer-performance study.Materials and Methods. The task of detecting microcalcifications was simulated by imaging aluminum wire seg ments (200-500 mu m in length; 100, 125, or 150 mu m in diameter) that overlapped with tissue background structures produced by beef brisket. A total of 288 such simulations were generated and examined with both computed radiography and conventional screen-film mammography techniques. Computed radiography was performed with high-resolution plates, a 43-mu m image reader, and a 43-mu m laser film printer. Computed radiographic images were printed with simple contrast enhancement and compared with screen-film images in a receiver operating characteristic study in which experienced readers detected and scored the simulated microcalcifications. Observer performance was quantitated and compared by computing the area under the receiver operating characteristic curve.Results. Although the resolution of the computed radiography system was better than that of commercial systems, it fell short of that of screen-film systems. For the 100-mu m microcalcifications, the difference in the average area under the curve was not statistically significant, but it was significant for the larger simulated microcalcifications: the average area under the curve was 0.58 for computed radiography versus 0.76 for screen-film imaging for the 125-mu m microcalcifications and 0.83 versus 1.00, respectively, for the 150-mu m microcalcifications.Conclusion. Observer performance in the detection of small simulated microcalcifications (100-150 mu m in diameter) is better with screen-film images than with high-resolution computed radiographic images.
To evaluate the sensitivity of a non-receiver-operating characteristic (ROC) study in assessing small differences of perceived image quality of hand images acquired by computed radiography (CR) and conventional screen-film systems, hand images were acquired on 12 patients with both conventional screen-film and CR. Each CR image was then processed with three different edge-enhancement algorithms. One conventional film and four CR images were then viewed side by side by five radiologists. Observers rated perceived image quality of each radiograph using a 10-category discrete scale. The study was repeated after 6 weeks using a different block randomization scheme. Despite the small sample size, significant differences (P < .05) in assigned image quality were detected among CR images acquired at low, medium, and high resolutions. Image processing routines did not fully compensate for differences in quality between conventional film and CR-acquired images. The quality rating of the reference conventional image was found to be dependent on the quality of images with which it was compared. Small, highly sensitive study designs can be used to identify radiologists' perceived differences in image quality. "Reference" or "gold standard" quality are important in such studies. Edge-enhancement schemes cannot fully compensate for perceived image quality degradations because of reduced image resolution.
RATIONALE AND OBJECTIVES:We investigated non-receiver operating characteristic (non-ROC) methods for the selection of processing algorithms for digital image compression. METHODS:We performed a multipoint, rank-order study with 20 posteroanterior chest images, each processed using four different algorithms. Seven radiologists reviewed these alongside the digitized noncompressed image. Observers were forced to rank order the similarity and/or difference of the processed images to the nonprocessed image in each case. RESULTS:A two-way analysis of variance of the rankings was statistically significant (p = .025), indicating that one processing scheme yielded images that were clearly perceived as the most similar to the nonprocessed images. The selected processing scheme was not the one that yielded the lowest quantitative difference from the nonprocessed images as measured by root mean square error. CONCLUSION:Non-ROC study designs that are highly sensitive to small differences among similar images can be used to select processing algorithms.