
Objective To describe the technical workflow enabling scalable automated artificial intelligence (AI) monitoring in the first national imaging AI registry, Assess-AI. Materials and Methods Large language model (LLM) prompts are developed to extract clinically relevant findings from radiology reports through collaboration between data scientists and subspecialty radiologists. Prompts are optimized using tuning cohorts of use case-specific radiology reports and LLMs available through AWS Bedrock. Such cohorts are used to evaluate prompt accuracy and consistency across 10 repeated runs. Report-AI result pairs submitted to the Assess-AI registry for actively monitored use cases are additionally used to further optimize corresponding prompts. Results Prompts were developed for nine use cases. In Stage 2 development cohorts, final-prompt agreement with hybrid report-derived reference standard labels was 0.985 for ICH and 0.997 for PE. Because these cohorts informed prompt refinement and label construction, they were not independent validation sets. Discussion The workflow demonstrates feasible report-finding extraction at scale; independent accuracy and clinical utility remain unestablished. Conclusion LLM-based extraction within Assess-AI enables scalable, report-anchored AI performance monitoring in radiology.
Incidental findings on pediatric imaging studies can adversely impact patients and their families when inappropriately managed. However, there is limited literature on management of these findings in the pediatric population, and guidelines for similar findings in adult patients may not be suitable. Here, examples of incidental findings in children are illustrated to give readers a practical understanding of how pediatric incidental findings may differ from those in adults. The ACR's planned approach to developing pediatric-specific guidance for incidental findings is outlined, and a strategy for managing pediatric incidental findings in the absence of specific pediatric guidance is also provided.
OBJECTIVE:To assess the association between imaging ordering rates after office-based evaluation and management (E&M) visits and the percentage of a practice's imaging studies interpreted by a radiologist versus the ordering provider. METHODS:CMS 5% Research Identifiable Files (2021-2023) were used to identify office-based E&M claims, patient and clinical characteristics, and associated imaging (CT, MR, nuclear medicine, ultrasound, radiography/fluoroscopy) within 30 days. Provider and practice characteristics were derived from the CMS Medicare Data on Provider Practice and Specialty. Multivariable logistic regression assessed the association between imaging after an E&M visit and the share of imaging interpreted by radiologists or ordering providers, controlling for provider, practice, patient, and clinical characteristics. RESULTS:There were 24,243,642 E&M visits that met the selection criteria. Of these, 13.6% had imaging, and the ordering provider's self-interpretation rate was 33.2%. Practices with 100% versus 0% radiologist interpretation had lower adjusted odds ratios for 30-day imaging for all modalities except CT ranging from 0.59 (95% confidence interval [CI], 0.57-0.62) for ultrasound to 0.95 (95% CI, 0.90-0.99) for MR. Conversely, ordering providers with 100% versus 0% self-interpretation had higher odds ratios for visit-associated imaging, overall (2.88; 95% CI, 2.79-2.98) and for each modality ranging from 1.54 (95% CI, 1.42-1.66) for MR to 3.51 (95% CI, 3.37-3.65) for ultrasound. Odds varied by ordering provider subspecialty and patient diagnosis. DISCUSSION:Imaging orders were more likely after visits to providers that always versus never self-interpret imaging, while visits to practices that refer a higher percentage of imaging to radiologists were less likely to result in imaging.
Background Medicare adjusts radiologists' reimbursement under the Merit-based Incentive Payment System (MIPS); CMS plans a mandatory transition to MIPS Value Pathways (MVPs), which restrict multispecialty group-level reporting. Objective To analyze 2023 MIPS performance for diagnostic radiologists, identify structural factors affecting payment adjustments, and estimate the potential impact of the ExRad quality measure (Q494). Methods Cross-sectional analysis of publicly available 2023 MIPS data for 249,848 physicians, including 25,160 diagnostic radiologists. Radiologists were classified by reporting pattern (mixed, non-radiology only, radiology only, no measures) and by group versus individual reporting. Twenty radiology measures were assessed for topped-out status. A simulation replaced each radiology-only reporter's lowest scoring measure with ExRad at scores of 5, 7, and 9. Results Radiologists ranked 28th of 62 specialties in mean payment adjustment (+0.39% versus +0.20%). Mixed reporters (40.1%) had the highest quality score (85.1) and adjustment (+0.90%; 9.8% penalized); radiology-only reporters (15.8%) had the lowest among measure reporters (quality score 71.4; +0.14%; 27.3% penalized). Six of 20 radiology measures were topped-out in 2023, and the most frequently reported measures averaged 5.7-8.3 among radiology-only reporters. Among individually reporting radiologists, 87.3% reported no measures (mean adjustment -1.99%). In simulation, ExRad substitution increased estimated adjustments by 0.15-0.45 percentage points; at a score of 9, 40.0% of penalized radiology-only reporters avoided negative adjustments. Conclusion Radiologists' MIPS performance depended on multispecialty group reporting of non-radiology measures, a strategy MVPs will largely eliminate. Modeled estimates suggest ExRad, the sole outcome measure in the 2026 diagnostic radiology MVP, could partially offset this loss.
BACKGROUND:Radiologists frequently provide informal consultations regarding imaging interpretation, modality selection, and protocol planning. These interactions are valuable but often undocumented and may interrupt workflow. Electronic consultation (eConsult) systems enable asynchronous provider-to-provider communication within the electronic health record (EHR). OBJECTIVE:To describe the implementation and early utilization of a pediatric radiology eConsult platform integrated within the EHR at a large academic pediatric healthcare system. MATERIALS AND METHODS:This institutional review board-approved retrospective study evaluated an eConsult tool integrated into the Epic EHR between April 2021 and July 2025. The system enabled asynchronous communication between pediatric care providers and pediatric radiologists for non-urgent outpatient imaging-related questions. Consultation requests were extracted from the radiology database and categorized by provider type and consultation purpose. Provider and radiologist perceptions were assessed using structured REDCap surveys. RESULTS:A total of 257 eConsults were submitted (mean, 6.8/month). Most originated from pediatric gastroenterology (147/257, 57.2%) and primary care (64/257, 24.9%). The most common categories were diagnostic interpretation (124/257, 48.2%) and imaging planning (77/257, 29.9%). Utilization increased from 43 in 2022 to 73 in 2023 and 76 in 2024, indicating progressive adoption. Among providers (n = 26), all reported satisfaction (Likert ≥4), and 80.8% reported impact on clinical decision-making. Radiologists reported good usability with variable workload impact. CONCLUSION:A pediatric radiology eConsult platform supports structured, asynchronous communication within the EHR. Despite modest early utilization, it provides a documented pathway for non-urgent imaging questions and supports the consultative role of radiologists.
OBJECTIVE:SPECT-CT (Single Photon Emission Computed Tomography integrated with Computed Tomography) is an advanced multimodal imaging technique with established applications across musculoskeletal, oncological, and neurological indications. However, the national geographic and sociodemographic distribution of SPECT-CT access in the United States remains poorly characterized. This study aimed to build a comprehensive national database of SPECT-CT facilities, and quantify access disparities across geographic and sociodemographic dimensions. METHODS:Facility data were compiled from the American College of Radiology (ACR) database as of April 2024. Demographic and socioeconomic characteristics for each ZIP Code Tabulation Area (ZCTA) were obtained from the 2021 American Community Survey. Spatial access was measured using the Two-Step Floating Catchment Area (2SFCA) method. Inequality was assessed via Lorenz curves, Gini coefficient, and Theil index. Decomposition of the Theil index was used to evaluate between- and within-group disparities by income, race, sex, and population. Geographically Weighted Regression (GWR) was used to evaluate spatial variation in the association between sociodemographic predictors and distance to the nearest SPECT-CT facility, compared against a global Ordinary Least Squares (OLS) model. RESULTS:A total of 1,661 SPECT-CT facilities were identified, with coastal and urban states showing higher density. Spatial access was highly unequal (Gini = 0.327; Theil = 0.205), with many ZCTAs receiving no meaningful access. Decomposition analyses revealed significant between-group inequality by race and population (p = 0.001 and 0.004), but not by income or sex. Within-group disparities were also substantial, with lower-White and high-population areas exhibiting the highest internal inequality. The GWR model demonstrated markedly better fit than OLS (adjusted R2 = 0.984 vs. 0.115) and identified strong regional variation in predictor effects, including reversals in the direction of racial impact (local β for percent White: -7.82 to +9.37 miles/%). CONCLUSION:SPECT-CT access in the US is marked by stark geographic and demographic disparities, characterized by inequity across multiple axes. By improving access, we can also promote the adoption of preventative care over reactive interventions.
Recommendations for additional imaging (RAI) are common in radiology reports and are intended to facilitate timely diagnosis and patient management. However, adherence to these recommendations is often low and little is known about recommendations for dedicated orbital imaging following head and neck CT or MRI examinations. We therefore extracted radiology reports containing recommendations for additional orbital CT or MRI from an available data set of head and neck CT and MRI radiology reports from 6/1/2021 to 5/31/2022 from a single large healthcare system with multiple academic and community radiology practices. 133/60,543 (0.2%) of reports included recommendations for additional orbit CT or MRI. Follow-up CT or MRI of the orbits was performed within 1 year for 34% (45/133), of which 93% (42/45) showed pathological findings (62% neoplastic; 22% inflammatory; 16% infection, traumatic, or vascular). Patients with retrobulbar pathology were significantly more likely to undergo follow-up imaging (p<0.001). The high prevalence of pathological findings among completed follow-up examinations may reflect the clinical importantance of the recommendation, however many examinations may have been performed as part of planned surveillance rather and not as a direct result of the radiologist’s recommendation.
To evaluate the operational impact and accuracy of a vendor-agnostic, artificial intelligence (AI)-based optical character recognition (OCR) system for drafting dual-energy x-ray absorptiometry (DXA) reports across academic and community practice settings, we implemented a DXA reporting pipeline across four outpatient imaging sites within a single health system (two academic, two community). The system used AI OCR to extract measurements from DXA DICOM images and rule-based logic to generate complete draft reports within the radiology reporting system. Operational impact was assessed using a pre/post design, measuring report creation time (RCT; report start to first signature) and report turnaround time (TAT; study completion to final signature). Report accuracy was assessed by comparing AI drafts and original reports against source images (n = 400). Median RCT decreased from 3.48 to 0.87 min and median TAT from 2.40 to 0.96 hours at the academic sites. At the community sites, median RCT decreased from 1.42 to 0.63 min and median TAT from 133.82 to 42.10 hours. Mean differences for all four comparisons were significant by Welch’s t test (all P < .0001). AI drafts had comparable to slightly higher numerical accuracy than original reports (academic: 99.9% versus 99.4%, P = .022; community: 100% versus 99.6%, P = .031), increased completeness at community sites (100% versus 45%, P < .001), and preserved diagnostic accuracy (≥99.5% across cohorts). Implementation of a vendor-agnostic, AI OCR DXA reporting system was associated with substantial reductions in RCT and report TAT, including an approximately 4-day median TAT reduction at community sites, while maintaining numerical and diagnostic accuracy and improving report completeness.
Diffuse lung diseases (DLD) also referred to as diffuse parenchymal lung diseases or interstitial lung diseases encompass diverse disorders affecting the lung parenchyma with possible multicompartment involvement in the chest. DLD include several hundred established clinical syndromes and pathologies, with a variety of possible etiologies. Imaging plays a central role during multidisciplinary discussion, which constitutes the current standard for diagnosis and monitoring of DLD. This document aims to establish guidelines for evaluation of diffuse lung diseases for 1) initial imaging of suspected diffuse lung disease, 2) initial imaging of suspected acute exacerbation or acute deterioration in cases of confirmed diffuse lung disease, and 3) surveillance of confirmed diffuse lung disease without acute deterioration. The American College of Radiology Appropriateness Criteria are evidence-based guidelines for specific clinical conditions that are reviewed annually by a multidisciplinary expert panel. The guideline development and revision process support the systematic analysis of the medical literature from peer reviewed journals. Established methodology principles such as Grading of Recommendations Assessment, Development, and Evaluation or GRADE are adapted to evaluate the evidence. The RAND/UCLA Appropriateness Method User Manual provides the methodology to determine the appropriateness of imaging and treatment procedures for specific clinical scenarios. In those instances where peer reviewed literature is lacking or equivocal, experts may be the primary evidentiary source available to formulate a recommendation.
PURPOSE:To compare the diagnostic performance and interreader agreement of structured ACR Bone Reporting and Data System (Bone-RADS) versus unstructured assessment of bone tumors on radiographs and to assess their effects on management recommendations. MATERIALS AND METHODS:Multicenter, multireader study with a primary retrospective cohort (n = 1,423; 4 centers), external retrospective cohort (n = 354; 6 centers), and prospective cohort (n = 152; 2 of these 10 centers; written informed consent obtained from all prospective participants). After a standardized training session with practice cases, nine musculoskeletal radiologists (early career 3-5 years; midcareer 15-20 years; late career >20 years) independently performed unstructured classification and Bone-RADS 1 to 4 grading with a standardized 4-week washout. Primary end point was reader-averaged difference in area under the curve (ΔAUC) via an Obuchowski-Rockette-Hillis multireader, multicase receiver operating characteristic model; interreader agreement and management recommendation reclassification were also assessed. RESULTS:Overall reader-averaged AUC did not differ significantly between unstructured assessment and Bone-RADS across the primary retrospective, external retrospective, and prospective cohorts (ΔAUC: -0.0012, 0.0160, and 0.0053, respectively; all P > .05). However, early-career readers demonstrated significant AUC improvements with Bone-RADS in the primary and external cohorts ΔAUC: +0.0302 and +0.0615; both P < .001), with a similar but nonsignificant difference in the prospective cohort (ΔAUC = +0.0201; P = .572). In contrast, midcareer readers showed nonsignificant changes, and late-career readers exhibited slight performance declines. Interreader agreement increased across all cohorts (eg, in the prospective cohort, early-career Fleiss κ improved from 0.394 to 0.502). Feature-level agreement was highest for pathologic fracture but lower for margins and endosteal erosion (κ ∼0.518-0.660). Potential malignancy rates increased monotonically across Bone-RADS categories 1 to 4 (9.44%, 25.44%, 46.67%, and 79.67%, respectively). Among early-career readers, Bone-RADS shifted management recommendations by increasing referral or further workup recommendations for potentially malignant lesions in the primary (+8.61%), external (+15.30%), and prospective (+7.14%) cohorts, while concurrently reducing routine surveillance recommendation rates for benign lesions (-17.27%, -15.90%, and -19.51%, respectively). CONCLUSION:Bone-RADS provides stable risk stratification, improves diagnostic performance and interreader agreement in early-career readers, and shifts management recommendations toward greater referral or workup of potentially malignant lesions. Its sensitivity-oriented design may increase overmanagement of some benign lesions. To mitigate this, future refinements should consider incorporating age and anatomic-site information, adopting context-aware management thresholds, and providing structured training on lower-agreement features such as margin classification and endosteal erosion.