Background The literature supports the use of short-interval follow-up as an alternative to biopsy for lesions assessed as probably benign, Breast Imaging Reporting and Data System (BI-RADS) category 3, with an expected malignancy rate of less than 2%. Purpose To assess outcomes from 6-, 12-, and 24-month follow-up of probably benign findings first identified at recall from screening mammography in the National Mammography Database (NMD). Materials and Methods This retrospective study included women recalled from screening mammography with BI-RADS category 3 assessment at additional evaluation from January 2009 through March 2018 from 471 NMD facilities. Only the first BI-RADS category 3 occurrence for women aged 25 years or older with no personal history of breast cancer was analyzed, with biopsy or 2-year imaging follow-up. Cancer yield and positive predictive value of biopsies performed (PPV3) were determined at each follow-up. Results Among 45 202 women (median age, 55 years; range, 25-90 years) with a BI-RADS category 3 lesion, 1574 (3.5%) underwent biopsy at the time of lesion detection, yielding 72 cancers (cancer yield, 4.6%; 72 of 1574 women). For the remaining 43 628 women who accepted surveillance, 922 were seen within 90 days (with 78 lesions biopsied and 12 [15%] classified as malignant). The women still in surveillance (31 465 of 43 381 women [72.5%]) underwent follow-up mammography at 6 months. Of 3001 (9.5%) lesions biopsied, 456 (15.2%) were malignant (cancer yield, 1.5%; 456 of 31 465 women; 95% confidence interval [CI]: 1.3%, 1.6%). Among 18 748 of 25 997 women (72.1%) in surveillance who underwent follow-up at 12 months, 1219 (6.5%) underwent biopsy with 230 (18.9%) malignant lesions found (cancer yield, 1.2%; 230 of 18 748 women; 95% CI: 1.1%, 1.4%). Through 2-year follow-up, the biopsy rate was 11.2% (4894 of 43 628 women) with a cancer yield of 1.86% (810 malignancies found among 43 628 women; 95% CI: 1.73%, 1.98%) and a PPV3 of 16.6% (810 malignancies found among 4894 women). Conclusion In the National Mammography Database, Breast Imaging Reporting and Data System (BI-RADS) category 3 use is appropriate, with 1.86% cumulative cancer yield through 2-year follow-up. Of 810 malignancies, 468 (57.8%) were diagnosed at or before 6 months, validating necessity of short-interval follow-up of mammographic BI-RADS category 3 findings. © RSNA, 2020 Online supplemental material is available for this article. See also the editorial by Moy in this issue.
HomeRadiologyVol. 292, No. 2 PreviousNext Reviews and CommentaryFree AccessEditorialOptimizing Breast Cancer Screening Programs: Experience and StructuresRobert D. Rosenberg , David SeidenwurmRobert D. Rosenberg , David SeidenwurmAuthor AffiliationsFrom the Radiology Associates of Albuquerque, 4411 The 25 Way NE, Suite 150, Albuquerque, NM 87109 (R.D.R.); and Department of Diagnostic Imaging, Sutter Health, Sacramento, Calif (D.S.).Address correspondence to R.D.R. (e-mail: [email protected]).Robert D. Rosenberg David SeidenwurmPublished Online:May 28 2019https://doi.org/10.1148/radiol.2019190924MoreSectionsPDF ToolsImage ViewerAdd to favoritesCiteTrack CitationsPermissionsReprints ShareShare onFacebookTwitterLinked In See also the article by Hoff et al in this issue.IntroductionThe standards for accuracy and efficiency of a screening test are higher than those of most medical tests because they are applied to asymptomatic healthy people, with the expectation of long-term benefit. Poor quality is harmful and costly, while overregulation or unreasonable requirements are burdensome, limit access, and, therefore, need rigorous justification. Demanding image quality requirements for mammography are similar across Europe and the United States. However, the annual volume requirements and recommendations vary widely, from 960 mammograms within the past 2 years under the Mammography Quality Standards Act (1) to 3000–5000 mammograms per year required or recommended in European screening programs (2). In addition, the required initial experience under the Mammography Quality Standards Act is just 240 mammograms within a 6-month period. Precisely how annual interpretation, initial and overall cumulative experience, and program structure affect breast cancer screening program quality remain open questions despite decades of effort.The recent study by Hoff and colleagues in this issue of Radiology (3) has three important findings concerning optimal interpretation volumes. First, radiologists’ false-positive rates, and therefore specificity, are greatly impacted by experience and annual volumes. False-positive rates are high for the first 20 000 career mammograms and with a low annual volume of less than 2000 mammograms. Second, a radiologist’s cancer detection rate is less associated with experience or annual volume and is optimal with 4000–10 000 annual mammograms. The few radiologists with annual volumes of more than 10 000 mammograms had the lowest cancer detection rates. Finally, double reading all mammograms with consensus review of discordant interpretations reduces false-positive rates and increases cancer detection.Thus, optimal interpretation accuracy is improved with adequate initial experience, ongoing practice, and internal review. The screening environment in Norway with biennial mammography and double reading with consensus is different from that in single-reader systems, so the incremental benefits of suggested program changes must be validated cautiously as they are applied to existing practices.Research in the United States by the Breast Cancer Surveillance Consortium (4–6) and in Canada (7) in single-reading screen-film environments showed results similar to those shown in Norway. There is improvement in performance associated with both higher initial experience and higher annual interpretation volumes. Specifically, performing work-up studies of recall cases (6), greater than 3 years of experience, and higher annual interpretation volumes are associated with improvement in the false-positive rate. Associations with sensitivity or cancer detection improvement were more difficult to identify (4,6,7).What remains to be determined is which specific aspect of experience leads to improvement and how this can be targeted to reduce the time needed to acquire the needed skills. There may be two processes involved—one for cancer detection sensitivity and one for specificity improvement.It seems that the consensus process, or review of the work-up process (7), provides useful feedback on false-positive cases because the overwhelming majority of diagnostic mammograms are negative. Cancer detection, however, requires adequate initial training and adequate interpretation time (3). Any research on sensitivity is limited by the relative scarcity of three to five cancers per 1000 in the screening environment and resultant poor statistical power of most research studies. This particularly limits the assessment of low-volume radiologists, as they will never achieve adequate cancer case volumes.One important issue is the diversity in false-positive rates in different systems. Typical false-positive rates or recall rates in the United States and Canada are 9%–10% (4,6,7), compared with 4% in the study by Hoff et al. This is undoubtedly multifactorial. A mix of the medical-legal environments, habits, peer modeling, training customs, and societal differences contribute to this variation. There are no concrete incentives in the United States to reduce recall. Breast Cancer Surveillance Consortium work on optimizing recall in the screen-film environment suggests that a recall rate of 6%–8% is achievable at similar sensitivity for many radiologists. Recall may also decrease with use of multiple prior mammograms and avoidance of findings unlikely to be clinically important (8).Several specific approaches for the improvement of overall accuracy are implied by Hoff and colleagues and the body of literature they build upon. However, these solutions may add new barriers to access or have other unintended consequences.The first approach is to increase the minimal reader volume in screening while monitoring the screening detection rate in very high-volume readers.The second approach is to minimize the medical-legal fears associated with low false-positive rates. General double reading may be optimal but is impractical in many systems. A first step could be double reading and consensus for any initially recalled and discordant cases. This should produce a diminished liability risk because at least two radiologists agreed at the time of screening that further work-up was not indicated. Careful monitoring of recall and cancer detection rates would be required during a transition period. Although any double reading would increase the radiologist time in screening mammography, the burden would be at least partly recouped through savings in the time- and resource-intensive diagnostic imaging process.The third approach is to address the problem of the low rate of cancers in the screening environment and decrease the training time needed for new readers. There are methods to either add known cancers to the screening environment (9,10) or provide access to cancer-enriched screening mammography case sets. Any current requirement of ongoing or initial minimal screening volumes should also include minimal numbers of cancers detected, including those in enriched environments, since cancer detection is the immediate screening goal. Availability of enriched case sets or enriched screening also allows for ongoing assessment of cancer detection ability.Incentives in payment structures that reward high cancer detection and lower recall rates are needed. Replacement of the procedure-based, atomized fee-for-service payment structure with payment for the annual screening episode, appropriately priced to include all services related to screening mammography, would accomplish this goal and would be administratively feasible. A single payment for the annual screening episode would include relevant professional and technical services through the initial biopsy. This type of plan would incentivize cancer diagnosis and disincentivize false-positive recall while permitting local decision making by each practice. Certain very rare patient populations may not be appropriate for this approach, but the majority of screening systems service typical populations. Medicare would be an ideal testing ground for this model of care due to the national scope and consistent enrollment of beneficiaries.Finally, technology has changed our work environment, so we should use its benefits. Digital imaging permits easy implementation and testing of cancer detection improvement strategies as original images may be shared. Tomosynthesis has reduced recall compared with digital mammography, and we anticipate further technology improvements (eg, computer-aided detection and/or artificial intelligence and screening MRI and breast US). However, for the foreseeable future the greatest gains are likely to be achieved by optimizing the most precious resources in breast cancer screening: the radiologists’ time and our patients’ well-being.In conclusion, sufficient data exist to recommend structural changes in screening mammography programs. Treating the whole patient and the system of care will lead us toward optimally performing systems.Disclosures of Conflicts of Interest: R.D.R. disclosed no relevant relationships. D.S. Activities related to the present article: receives fees as a professional liability witness and personal injury expert. Activities not related to the present article: receives fees for expert testimony as a professional liability witness and personal injury expert; is employed by Sutter Medical Group; has stock/stock options in Sutter Medical Group and Radiological Associated Medical Group. Other relationships: disclosed no relevant relationships.References1. Mammography Quality Standards Act Regulations. U.S. Food & Drug Administration. https://www.fda.gov/Radiation-EmittingProducts/MammographyQualityStandardsActandProgram/Regulations/ucm110906.htm. Updated November 29, 2017. Accessed April 21, 2019. Google Scholar2. Perry N, Broeders M, de Wolf C, Törnberg S, Holland R, von Karsa L, eds. European guidelines for quality assurance in breast cancer screening and diagnosis. 4th ed. Luxembourg: Office for Official Publications of the European Communities, 2006. Google Scholar3. Hoff SR, Myklebust TA, Lee CI, Hofvind S. Influence of mammography volume on radiologists’ performance: results from BreastScreen Norway. Radiology 2019; https://doi.org/10.1148/radiol.2019182684. Published online May 28, 2019. Link, Google Scholar4. Elmore JG, Jackson SL, Abraham L, et al. Variability in interpretive performance at screening mammography and radiologists’ characteristics associated with accuracy. Radiology 2009;253(3):641–651. Link, Google Scholar5. Miglioretti DL, Gard CC, Carney PA, et al. When radiologists perform best: the learning curve in screening mammogram interpretation. Radiology 2009;253(3):632–640. Link, Google Scholar6. Buist DS, Anderson ML, Haneuse SJ, et al. Influence of annual interpretive volume on screening mammography performance in the United States. Radiology 2011;259(1):72–84. Link, Google Scholar7. Théberge I, Chang SL, Vandal N, et al. Radiologist interpretive volume and breast cancer screening accuracy in a Canadian organized screening program. J Natl Cancer Inst 2014;106(3):djt461. Crossref, Medline, Google Scholar8. Sickles EA. Successful methods to reduce false-positive mammography interpretations. Radiol Clin North Am 2000;38(4):693–700. Crossref, Medline, Google Scholar9. Gordon PB, Borugian MJ, Warren Burhenne LJ. A true screening environment for review of interval breast cancers: pilot study to reduce bias. Radiology 2007;245(2):411–415. Link, Google Scholar10. Evans KK, Birdwell RL, Wolfe JM. If you don’t find it often, you often don’t find it: why some cancers are missed in breast cancer screening. PLoS One 2013;8(5):e64366. Crossref, Medline, Google ScholarArticle HistoryReceived: Apr 22 2019Revision requested: Apr 29 2019Revision received: May 1 2019Accepted: May 1 2019Published online: May 28 2019Published in print: Aug 2019 FiguresReferencesRelatedDetailsCited ByCorrelation of Test Sets and Actual Clinical PerformanceDavid J. Seidenwurm, ,Robert D. Rosenberg, 29 January 2021 | Radiology: Imaging Cancer, Vol. 3, No. 1Understanding the Mammography AuditKimberlyFunaro, DanaAtaya, BethanyNiell2021 | Radiologic Clinics of North America, Vol. 59, No. 1Impact of radiomics on the breast ultrasound radiologist’s clinical practice: From lumpologist to data wranglerEduardo de Faria CastroFleury, KarenMarcomini2020 | European Journal of Radiology, Vol. 131Accompanying This ArticleInfluence of Mammography Volume on Radiologists’ Performance: Results from BreastScreen NorwayMay 28 2019RadiologyRecommended Articles Clinical Performance of Synthesized Two-dimensional Mammography Combined with Tomosynthesis in a Large Screening PopulationRadiology2017Volume: 283Issue: 1pp. 70-76Implementation of Synthesized Two-dimensional Mammography in a Population-based Digital Breast Tomosynthesis Screening ProgramRadiology2016Volume: 281Issue: 3pp. 730-736An Effective Method to Reduce the Interpretation Time in the Clinical Use of Digital Breast TomosynthesisRadiology2020Volume: 297Issue: 3pp. 543-544Does Reader Performance with Digital Breast Tomosynthesis Vary according to Experience with Two-dimensional Mammography?Radiology2017Volume: 283Issue: 2pp. 371-380Equalizing the Performance of Radiologists and Physician Extenders in Screening MammographyRadiology2022Volume: 306Issue: 1pp. 110-111See More RSNA Education Exhibits Letâs Talk about Next-Generation Breast Cancer Screening Programs: How Should We Do? What Should We Use?Digital Posters2020Non-Contrast-Enhanced Breast MR Screening for Women with Dense BreastsDigital Posters2019Breast Density Included in the Modern Rules of Mammographic ScreeningDigital Posters2019 RSNA Case Collection Invasive Lobular CarcinomaRSNA Case Collection2021Breast edemaRSNA Case Collection2021Asymmetric lactational change RSNA Case Collection2021 Vol. 292, No. 2 Metrics Altmetric Score PDF download
PURPOSE:The National Mammography Database (NMD) contains nearly 20 million examinations from 693 facilities; it is the largest information source for use and effectiveness of breast imaging in the United States. NMD collects demographic, imaging, interpretation, biopsy, and basic pathology results, enabling facility and physician comparison for quality improvement. However, NMD lacks treatment and clinical outcomes data. The network of state cancer registries (CRs) contains detailed pathologic, treatment, and clinical outcomes data. This pilot study assessed electronic linkage of NMD and CR data at a multicenter institution as proof of concept. MATERIALS AND METHODS:We obtained Quality Oversight Committee approval for this retrospective study. Data of patients diagnosed with breast cancer in 2014 and 2015 were retrieved from our NMD-approved radiology information system (RIS) and matched with reportable patients in our CR using social security number (SSN), first name (fname), last name (lname), and date of birth (DOB). Matching was repeated without SSN. Percentage and reasons for mismatch were evaluated. RESULTS:The RIS query identified 1,316 patients. CR linkage was 99.2% successful (n = 1,305 of 1,316) using SSN, fname, lname, and DOB. Eleven mismatches included four CR case-finding failures, one NMD fname error, five nonreportable in the CR, and one with correct identifiers in both databases. Without SSN, linkage was 97.3% successful (n = 1,281 of 1,316); name errors accounted for 19 and DOB accounted for 5 additional mismatches. CONCLUSION:Using common data elements, linkage between the NMD and state CRs may be feasible and could provide critical outcomes information to advance accurate assessment of breast imaging in the United States.
HomeRadiologyVol. 287, No. 3 PreviousNext Reviews and CommentaryEditorialBreast Cancer Screening: Two (or Three) Heads Are Better than One?Robert D. Rosenberg , David SeidenwurmRobert D. Rosenberg , David SeidenwurmAuthor AffiliationsFrom Radiology Associates of Albuquerque, 4411 25th Way, Suite 150, Albuquerque, NM 87109 (R.D.R.); Department of Radiology, University of New Mexico HSC, Albuquerque, NM (R.D.R.); and Department of Neuroradiology, Diagnostic Imaging, Sutter Health, Sacramento, Calif (D.S.).Address correspondence to R.D.R. (e-mail: [email protected]).Robert D. Rosenberg David SeidenwurmPublished Online:Apr 10 2018https://doi.org/10.1148/radiol.2018180207MoreSectionsFull textPDF ToolsImage ViewerAdd to favoritesCiteTrack CitationsPermissionsReprints ShareShare onFacebookTwitterLinked In References1. International Cancer Screening Network. National Cancer Institute. https://healthcaredelivery.cancer.gov/icsn/. Accessed December 26, 2017. Google Scholar2. Taylor-Phillips S, Jenkinson D, Stinton C, Wallis MG, Dunn J, Clarke A. Double reading in breast cancer screening: cohort evaluation in the CO-OPS trial. Radiology 2018;287:749–757. Link, Google Scholar3. Wilson R, Liston J. Quality Assurance Guidelines for Breast Cancer Screening Radiology. NHS Breast Screening Programme Publication Number 59. 2nd ed. Sheffield, England: NHS Cancer Screening Programmes, 2011. Google Scholar4. Dang PA, Freer PE, Humphrey KL, Halpern EF, Rafferty EA. Addition of tomosynthesis to conventional digital mammography: effect on image interpretation time of screening examinations. Radiology 2014;270(1):49–56. Link, Google Scholar5. Sprague BL, Arao RF, Miglioretti DL, et al; Breast Cancer Surveillance Consortium. National performance benchmarks for modern diagnostic digital mammography: update from the Breast Cancer Surveillance Consortium. Radiology 2017;283(1):59–69. Link, Google ScholarArticle HistoryReceived January 24, 2018; revision requested January 26; final revision received January 29; accepted January 30.Published online: Apr 10 2018Published in print: June 2018 FiguresReferencesRelatedDetailsRecommended Articles Performance of Radiologists and Radiographers in Double Reading Mammograms: The UK National Health Service Breast Screening ProgramRadiology2022Volume: 306Issue: 1pp. 102-109Clinical Performance of Synthesized Two-dimensional Mammography Combined with Tomosynthesis in a Large Screening PopulationRadiology2017Volume: 283Issue: 1pp. 70-76Comparison of Mammography AI Algorithms with a Clinical Risk Model for 5-year Breast Cancer Risk Prediction: An Observational StudyRadiology2023Volume: 307Issue: 5Consensus Reads: The More Sets of Eyes Interpreting a Mammogram, the Better for WomenRadiology2020Volume: 295Issue: 1pp. 42-43Standalone AI for Breast Cancer Detection at Screening Digital Mammography and Digital Breast Tomosynthesis: A Systematic Review and Meta-AnalysisRadiology2023Volume: 307Issue: 5See More RSNA Education Exhibits Non-Contrast-Enhanced Breast MR Screening for Women with Dense BreastsDigital Posters2019Letâs Talk about Next-Generation Breast Cancer Screening Programs: How Should We Do? What Should We Use?Digital Posters2020The New Era for Breast Cancer Screening: Abbreviated Breast Magnetic Resonance ImagingDigital Posters2019 RSNA Case Collection Ductal carcinoma in situRSNA Case Collection2020Mixed Mucinous Breast Carcinoma RSNA Case Collection2022DCIS within breast fibroadenomaRSNA Case Collection2020 Vol. 287, No. 3 Metrics Altmetric Score PDF download
Background.-Recall for assessment in mammographic screening entails an inevitable number of false-positive screening results. This study aimed to investigate the variation in the cumulative risk of a false positive screening result and the positive predictive value across the screening centres in the Norwegian Breast Cancer Screening Program.Methods.-We studied 618,636 women aged 50-69 years who underwent 2,090,575 screening exams (1996-2010). Recall rate, positive predictive value, rate of screen-detected cancer, and the cumulative risk of a false positive screening result, without and with invasive procedures across the screening centres were calculated. Generalized linear models were used to estimate the probability of a false positive screening result and to compute the cumulative false-positive risk for up to ten biennial screening examinations.Results.-The cumulative risk of a false-positive screening exam varied from 10.7% (95% CI: 9.4-12.0%) to 41.5% (95% CI: 34.1-48.9%) across screening centres, with a highest to lowest ratio of 3.9 (95% CI: 3.7-4.0). The highest to lowest ratio for the cumulative risk of undergoing an invasive procedure with a benign outcome was 4.3 (95% CI: 4.0-4.6). The positive predictive value of recall varied between 12.0% (95% CI: 11.0-12.9%) and 19.9% (95% CI: 18.3-21.5%), with a highest to lowest ratio of 1.7 (95% CI: 1.5-1.9).Conclusions.-A substantial variation in the performance measures across the screening centres in the Norwegian Breast Cancer Screening Program was identified, despite of similar administration, procedures, and quality assurance requirements. Differences in the readers' performance is probably of influence for the variability. This results underscore the importance of continuous surveillance of the screening centres and the radiologists in order to sustain and improve the performance and effectiveness of screening programs (Figs 1 and 2).
OBJECTIVE Using a combination of performance measures, we updated previously proposed criteria for identifying physicians whose performance interpreting screening mammography may indicate suboptimal interpretation skills. MATERIALS AND METHODS In this study, six expert breast imagers used a method based on the Angoff approach to update criteria for acceptable mammography performance on the basis of two sets of combined performance measures: set 1, sensitivity and specificity for facilities with complete capture of false-negative cancers; and set 2, cancer detection rate (CDR), recall rate, and positive predictive value of a recall (PPV1) for facilities that cannot capture false-negative cancers but have reliable cancer follow-up information for positive mammography results. Decisions were informed by normative data from the Breast Cancer Surveillance Consortium (BCSC). RESULTS Updated combined ranges for acceptable sensitivity and specificity of screening mammography are sensitivity≥80% and specificity≥85% or sensitivity 75-79% and specificity 88-97%. Updated ranges for CDR, recall rate, and PPV1 are: CDR≥6 per 1000, recall rate 3-20%, and any PPV1; CDR 4-6 per 1000, recall rate 3-15%, and PPV1≥3%; or CDR 2.5-4.0 per 1000, recall rate 5-12%, and PPV1 3-8%. Using the original criteria, 51% of BCSC radiologists had acceptable sensitivity and specificity; 40% had acceptable CDR, recall rate, and PPV1. Using the combined criteria, 69% had acceptable sensitivity and specificity and 62% had acceptable CDR, recall rate, and PPV1. CONCLUSION The combined criteria improve previous criteria by considering the interrelationships of multiple performance measures and broaden the acceptable performance ranges compared with previous criteria based on individual measures.
PURPOSETo develop criteria to identify thresholds for the minimally acceptable performance of physicians interpreting diagnostic mammography studies.MATERIALS AND METHODSIn an institutional review board-approved HIPAA-compliant study, an Angoff approach was used to set criteria for identifying minimally acceptable interpretive performance for both workup after abnormal screening examinations and workup of a breast lump. Normative data from the Breast Cancer Surveillance Consortium (BCSC) was used to help the expert radiologist identify the impact of cut points. Simulations, also using data from the BCSC, were used to estimate the expected clinical impact from the recommended performance thresholds.RESULTSFinal cut points for workup of abnormal screening examinations were as follows: sensitivity, less than 80%; specificity, less than 80% or greater than 95%; abnormal interpretation rate, less than 8% or greater than 25%; positive predictive value (PPV) of biopsy recommendation (PPV2), less than 15% or greater than 40%; PPV of biopsy performed (PPV3), less than 20% or greater than 45%; and cancer diagnosis rate, less than 20 per 1000 interpretations. Final cut points for workup of a breast lump were as follows: sensitivity, less than 85%; specificity, less than 83% or greater than 95%; abnormal interpretation rate, less than 10% or greater than 25%; PPV2, less than 25% or greater than 50%; PPV3, less than 30% or greater than 55%; and cancer diagnosis rate, less than 40 per 1000 interpretations. If underperforming physicians moved into the acceptable range after remedial training, the expected result would be (a) diagnosis of an additional 86 cancers per 100,000 women undergoing workup after screening examinations, with a reduction in the number of false-positive examinations by 1067 per 100,000 women undergoing this workup, and (b) diagnosis of an additional 335 cancers per 100,000 women undergoing workup of a breast lump, with a reduction in the number of false-positive examinations by 634 per 100,000 women undergoing this workup.CONCLUSIONInterpreting physicians who fall outside one or more of the identified cut points should be reviewed in the context of an overall assessment of all their performance measures and their specific practice setting to determine if remedial training is indicated.
OBJECTIVE:The purposes of this study were to determine whether U.S. radiologists accurately estimate their own interpretive performance of screening mammography and to assess how they compare their performance with that of their peers.SUBJECTS AND METHODS:Between 2005 and 2006, 174 radiologists from six Breast Cancer Surveillance Consortium registries completed a mailed survey. The radiologists' estimated and actual recall, false-positive, and cancer detection rates and positive predictive value of biopsy recommendation (PPV(2)) for screening mammography were compared. Radiologists' ratings of their performance as lower than, similar to, or higher than that of their peers were compared with their actual performance. Associations with radiologist characteristics were estimated with weighted generalized linear models.RESULTS:Although most radiologists accurately estimated their cancer detection and recall rates (74% and 78% of radiologists), fewer accurately estimated their false-positive rate (19%) and PPV(2) (26%). Radiologists reported having recall rates similar to (43%) or lower than (31%) and false-positive rates similar to (52%) or lower than (33%) those of their peers and similar (72%) or higher (23%) cancer detection rates and similar (72%) or higher (38%) PPV(2). Estimation accuracy did not differ by radiologist characteristics except that radiologists who interpreted 1000 or fewer mammograms annually were less accurate at estimating their recall rates.CONCLUSION:Radiologists perceive their performance to be better than it actually is and at least as good as that of their peers. Radiologists have particular difficulty estimating their false-positive rates and PPV(2).
Purpose The aim of this study was to assess agreement of mammographic interpretations by community radiologists with consensus interpretations of an expert radiology panel to inform approaches that improve mammographic performance. Methods From 6 mammographic registries, 119 community-based radiologists were recruited to assess 1 of 4 randomly assigned test sets of 109 screening mammograms with comparison studies for no recall or recall, giving the most significant finding type (mass, calcifications, asymmetric density, or architectural distortion) and location. The mean proportion of agreement with an expert radiology panel was calculated by cancer status, finding type, and difficulty level of identifying the finding at the patient, breast, and lesion level. Concordance in finding type between study radiologists and the expert panel was also examined. For each finding type, the proportion of unnecessary recalls, defined as study radiologist recalls that were not expert panel recalls, was determined. Results Recall agreement was 100% for masses and for examinations with obvious findings in both cancer and noncancer cases. Among cancer cases, recall agreement was lower for lesions that were subtle (50%) or asymmetric (60%). Subtle noncancer findings and benign calcifications showed 33% agreement for recall. Agreement for finding responsible for recall was low, especially for architectural distortions (43%) and asymmetric densities (40%). Most unnecessary recalls (51%) were asymmetric densities. Conclusions Agreement in mammographic interpretation was low for asymmetric densities and architectural distortions. Training focused on these interpretations could improve the accuracy of mammography and reduce unnecessary recalls.
PURPOSE:To examine whether U.S. radiologists' interpretive volume affects their screening mammography performance.MATERIALS AND METHODS:Annual interpretive volume measures (total, screening, diagnostic, and screening focus [ratio of screening to diagnostic mammograms]) were collected for 120 radiologists in the Breast Cancer Surveillance Consortium (BCSC) who interpreted 783 965 screening mammograms from 2002 to 2006. Volume measures in 1 year were examined by using multivariate logistic regression relative to screening sensitivity, false-positive rates, and cancer detection rate the next year. BCSC registries and the Statistical Coordinating Center received institutional review board approval for active or passive consenting processes and a Federal Certificate of Confidentiality and other protections for participating women, physicians, and facilities. All procedures were compliant with the terms of the Health Insurance Portability and Accountability Act.RESULTS:Mean sensitivity was 85.2% (95% confidence interval [CI]: 83.7%, 86.6%) and was significantly lower for radiologists with a greater screening focus (P = .023) but did not significantly differ by total (P = .47), screening (P = .33), or diagnostic (P = .23) volume. The mean false-positive rate was 9.1% (95% CI: 8.1%, 10.1%), with rates significantly higher for radiologists who had the lowest total (P = .008) and screening (P = .015) volumes. Radiologists with low diagnostic volume (P = .004 and P = .008) and a greater screening focus (P = .003 and P = .002) had significantly lower false-positive and cancer detection rates, respectively. Median invasive tumor size and proportion of cancers detected at early stages did not vary by volume.CONCLUSION:Increasing minimum interpretive volume requirements in the United States while adding a minimal requirement for diagnostic interpretation could reduce the number of false-positive work-ups without hindering cancer detection. These results provide detailed associations between mammography volumes and performance for policymakers to consider along with workforce, practice organization, and access issues and radiologist experience when reevaluating requirements.
Rationale and Objectives: Mammography quality assurance programs have been in place for more than a decade. We studied radiologists' self-reported performance goals for accuracy in screening mammography and compared them to published recommendations.Materials and Methods: A mailed survey of radiologists at mammography registries in seven states within the Breast Cancer Surveillance Consortium (BCSC) assessed radiologists' performance goals for interpreting screening mammograms. Self-reported goals were compared to published American College of Radiology (ACR) recommended desirable ranges for recall rate, false-positive rate, positive predictive value of biopsy recommendation (PPV2), and cancer detection rate. Radiologists' goals for interpretive accuracy within desirable range were evaluated for associations with their demographic characteristics, clinical experience, and receipt of audit reports.Results: The survey response rate was 71% (257 of 364 radiologists). The percentage of radiologists reporting goals within desirable ranges was 79% for recall rate, 22% for false-positive rate, 39% for PPV2, and 61% for cancer detection rate. The range of reported goals was 0%-100% for false-positive rate and PPV2. Primary academic affiliation, receiving more hours of breast imaging continuing medical education, and receiving audit reports at least annually were associated with desirable PPV2 goals. Radiologists reporting desirable cancer detection rate goals were more likely to have interpreted mammograms for 10 or more years, and >1000 mammograms per year.Conclusion: Many radiologists report goals for their accuracy when interpreting screening mammograms that fall outside of published desirable benchmarks, particularly for false-positive rate and PPV2, indicating an opportunity for education.
PURPOSETo describe the timeliness of follow-up care in community-based settings among women who receive a recommendation for immediate follow-up during the screening mammography process and how follow-up timeliness varies according to facility and facility-level characteristics.MATERIALS AND METHODSThis was an institutional review board-approved and HIPAA-compliant study. Screening mammograms obtained from 1996 to 2007 in women 40-80 years old in the Breast Cancer Surveillance Consortium were examined. Inclusion criteria were a recommendation for immediate follow-up at screening, or subsequent imaging, and observed follow-up within 180 days of the recommendation. Recommendations for additional imaging (AI) and biopsy or surgical consultation (BSC) were analyzed separately. The distribution of time to follow-up care was estimated by using the Kaplan-Meier estimator.RESULTSData were available on 214,897 AI recommendations from 118 facilities and 35,622 BSC recommendations from 101 facilities. The median time to subsequent follow-up care after recommendation was 14 days for AI and 16 days for BSC. Approximately 90% of AI follow-up and 81% of BSC follow-up occurred within 30 days. Facilities with higher recall rates tended to have longer AI follow-up times (P < .001). Over the study period, BSC follow-up rates at 15 and 30 days improved (P < .001). Follow-up times varied substantially across facilities. Timely follow-up was associated with larger volumes of the recommended procedures but not notably associated with facility type nor observed facility-level characteristics.CONCLUSIONMost patients with follow-up returned within 3 weeks of the recommendation.
It is difficult for ultrasound to image small targets such as breast microcalcifications. Synthetic aperture ultrasound imaging has recently developed as a promising tool to improve the capabilities of medical ultrasound. We use two different tissue-equivalent phantoms to study the imaging capabilities of a real-time synthetic aperture ultrasound system for imaging small targets. The InnerVision ultrasound system DAS009 is an investigational system for real-time synthetic aperture ultrasound imaging. We use the system to image the two phantoms, and compare the images with those obtained from clinical scanners Acuson Sequoia 512 and Siemens S2000. Our results show that synthetic aperture ultrasound imaging produces images with higher resolution and less image artifacts than Acuson Sequoia 512 and Siemens S2000. In addition, we study the effects of sound speed on synthetic aperture ultrasound imaging and demonstrate that an accurate sound speed is very important for imaging small targets.
PURPOSE:To investigate the association between radiologist interpretive volume and diagnostic mammography performance in community-based settings. MATERIALS AND METHODS:This study received institutional review board approval and was HIPAA compliant. A total of 117,136 diagnostic mammograms that were interpreted by 107 radiologists between 2002 and 2006 in the Breast Cancer Surveillance Consortium were included. Logistic regression analysis was used to estimate the adjusted effect on sensitivity and the rates of false-positive findings and cancer detection of four volume measures: annual diagnostic volume, screening volume, total volume, and diagnostic focus (percentage of total volume that is diagnostic). Analyses were stratified by the indication for imaging: additional imaging after screening mammography or evaluation of a breast concern or problem. RESULTS:Diagnostic volume was associated with sensitivity; the odds of a true-positive finding rose until a diagnostic volume of 1000 mammograms was reached; thereafter, they either leveled off (P < .001 for additional imaging) or decreased (P = .049 for breast concerns or problems) with further volume increases. Diagnostic focus was associated with false-positive rate; the odds of a false-positive finding increased until a diagnostic focus of 20% was reached and decreased thereafter (P < .024 for additional imaging and P < .001 for breast concerns or problems with no self-reported lump). Neither total volume nor screening volume was consistently associated with diagnostic performance. CONCLUSION:Interpretive volume and diagnostic performance have complex multifaceted relationships. Our results suggest that diagnostic interpretive volume is a key determinant in the development of thresholds for considering a diagnostic mammogram to be abnormal. Current volume regulations do not distinguish between screening and diagnostic mammography, and doing so would likely be challenging.
Ultrasound image resolution and quality need to be significantly improved for breast microcalcification detection. Super-resolution imaging with the factorization method has recently been developed as a promising tool to break through the resolution limit of conventional imaging. In addition, wave-equation reflection imaging has become an effective method to reduce image speckles by properly handling ultrasound scattering/diffraction from breast heterogeneities during image reconstruction. We explore the capabilities of a novel super-resolution ultrasound imaging method and a wave-equation reflection imaging scheme for detecting breast microcalcifications. Super-resolution imaging uses the singular value decomposition and a factorization scheme to achieve an image resolution that is not possible for conventional ultrasound imaging. Wave-equation reflection imaging employs a solution to the acoustic-wave equation in heterogeneous media to backpropagate ultrasound scattering/diffraction waves to scatters and reconstruct images of heterogeneities. We construct numerical breast phantoms using in vivo breast images, and use a finite-difference wave-equation scheme to generate ultrasound data scattered from inclusions that mimic microcalcifications. We demonstrate that microcalcifications can be detected at full spatial resolution using the super-resolution ultrasound imaging and wave-equation reflection imaging methods.
Background: Hispanic women in New Mexico (NM) are more likely than non-Hispanic women to die of breast cancer–related causes. We determined whether survival differences between Hispanic and non-Hispanic women might be attributable to the method of detection, an independent breast cancer prognostic factor in previous studies.Methods: White women diagnosed with invasive breast cancer from 1995 through 2004 were identified from NM Surveillance Epidemiology End Results (SEER) files (n = 5,067) and matched to NM Mammography Project records. Method of cancer detection was categorized as "symptomatic" or "screen-detected." The proportion of Hispanic survival disparity accounted for by included variables was assessed using Cox models.Results: In the median follow-up of 87 months, 490 breast cancer deaths occurred. Symptomatic versus screen-detection was classifiable for 3,891 women (76.8%), and was independently related to breast cancer-specific survival [hazard ratio (HR), 1.6; 95% confidence interval (95% CI), 1.3-2.0]. Hispanic women had a 1.5-fold increased risk of breast cancer–related death, relative to non-Hispanic women (95% CI, 1.2-1.8). After adjustment for detection method, the Hispanic HR declined from 1.50 to 1.45 (10%), but after inclusion of other prognostic indicators the Hispanic HR equaled 1.23 (95% CI, 1.01-1.48).Conclusions: Although the Hispanic HR declined 50% after adjustment, the decrease was largely due to adverse tumor prognostic characteristics.Impact: Reduction of disparate survival in Hispanic women may rely not only on increased detection of tumors when asymptomatic but on the development of greater understanding of biological factors that predispose to poor prognosis tumors. Cancer Epidemiol Biomarkers Prev; 19(10); 2453–60. ©2010 AACR.
Diagnostic mammography is the primary imaging modality to diagnose breast cancer. However, few studies have evaluated variability in diagnostic mammography performance in communities, and none has done so between countries. We compared diagnostic mammography performance in community-based settings in the United States and Denmark. The performance of 93,585 diagnostic mammograms from 180 facilities contributing data to the US Breast Cancer Surveillance Consortium (BCSC) from 1999 to 2001 was compared to that of all 51,313 diagnostic mammograms performed at Danish clinics in 2000. We used the imaging workup's final assessment to determine sensitivity, specificity and an estimate of accuracy: area under the receiver-operating characteristics (ROCs) curve (AUC). Diagnostic mammography had slightly higher sensitivity in the United States (85%) than in Denmark (82%). In contrast, it had higher specificity in Denmark (99%) than in the United States (93%). The AUC was high in both countries: 0.91 in United States and 0.95 in Denmark. Denmark's higher accuracy may result from supplementary ultrasound examinations, which are provided to 74% of Danish women but only 37% to 52% of US women. In addition, Danish mammography facilities specialize in either diagnosis or screening, possibly leading to greater diagnostic mammography expertise in facilities dedicated to symptomatic patients. Performance of community-based diagnostic mammography settings varied markedly between the 2 countries, indicating that it can be further optimized.