Artificial intelligence (AI) is rapidly changing the legal landscape of radiology. Here we examined whether the manner in which AI is integrated into the radiologist workflow impacts the perception of legal liability among mock jurors in the USA. Participants (n = 282) read about a hypothetical malpractice case where a patient suffered irreversible brain damage because a radiologist failed to detect a brain bleed from a computerized tomography (CT) scan, even though AI correctly identified the CT as abnormal. In the single-read condition, the radiologist interpreted the CT once after seeing AI feedback. In the double-read condition, the radiologist interpreted the CT twice, first without AI and then with AI feedback. Participants were asked whether the radiologist met their duty of care (yes or no). Participants were more likely to side with the plaintiff in the single-read versus the double-read condition (74.7% versus 52.9%), P = 0.0002 (odds ratio 2.6; 95% confidence interval 1.6, 4.4). This suggests that the penalty for disagreeing with correct AI can be mitigated when images are interpreted twice. In a randomized preregistered study, 282 participants were asked to judge a hypothetical malpractice case in different conditions of artificial intelligence (AI) assistance to a radiologist. An AI-assisted second-read framework was found to reduce the fraction of participants who would side with the plaintiff.
AI has been proposed as a triage or "rule-out" device to reduce radiologist workload, but it is presently unclear how an AI "rule-out" threshold should be determined. We present a framework for determining an optimal threshold. Using a retrospective study design, 114,229 bilateral 2D digital screening mammograms were analyzed from 2006-2023 at a single study site. All mammograms were given an AI score using Mirai, an open-source deep-learning model which provides a 1-year risk score. Several metrics were examined using two thresholds for determining ruled out versus retained cases: 1) Caseload Reduction Rate (CRR; percent of caseload reduced due to rule-out), 2) Gross AI False Omission Rate (G-FOR; probability of a patient having breast cancer if ruled out), 3) AI Net False Omission Rate (N-FOR; probability of a patient having breast cancer if ruled out and the radiologist would have caught in standard care [i.e., no triage]), 4) AI Adjusted Net False Omission Rate (30%) (AN-FOR[30%]; N-FOR adjusted for the hypothetical scenario where radiologists detect an extra 30% of breast cancers among AI retained cases). The two thresholds were risk scores of 0.2 (Youden's J) and 0.05 (AN-FOR[30%]=0). The former is mathematically optimal; the latter reflects a threshold where AI "rule-out" does not introduce any total increase in False Negatives. At the 0.20 threshold, G-FOR, N-FOR, and AN-FOR (30%) are 0.26%, 0.17%, and 0.14%, respectively (223, 141, and 121, respectively, missed cancer cases) and CRR = 75%. At the 0.05 threshold, the G-FOR, N-FOR, and AN-FOR (30%) are 0.12%, 0.07%, and 0.00% (49, 30, and 0, respectively, missed cancer cases) and CRR = 36%. We demonstrate how radiology practices can consider the trade-offs of using different AI scores as "rule-out" thresholds. At the AN-FOR rate of 30%, the Youden's J threshold results in 121 additional missed cancers for a 75% caseload reduction. We estimate no additional missed cancers at a 36% caseload reduction.
BACKGROUND:Artificial intelligence (AI) is rapidly being integrated into breast cancer screening, improving cancer detection and workflow efficiency. As patients become increasingly exposed to information about AI in mammography through news outlets, hospitals, and commercial entities, they are likely to seek information online. However, the readability and understandability of Online patient education materials (OPEM) for AI in mammography have not been examined. METHODS:We collected the top 20 nonsponsored results for each of 5 AI-related internet search terms in mammography. After removing duplicates, each webpage (n = 56) was categorized by the source type and evaluated for readability using 6 readability algorithms, as well as for understandability using the Patient Education Materials Assessment Tool for Printable Materials. RESULTS:The average grade-level readability across all webpages was 14.2, exceeding both the American Medical Association (sixth grade) and Centers for Disease Control and Prevention (eighth grade) recommendations. Reading ease scores placed most content in the "difficult" range, which is most suitable for college-level readers. Understandability averaged 72.4%, with variation by source type: commercial, government, and patient advocacy pages scored highest, while academic and medical media sources scored lowest. CONCLUSION:Online information about AI in mammography is generally written at a level too advanced for most patients and meets only the minimum standards of understandability. Because patients are likely to encounter both OPEM and non-OPEM sources when searching online, clinicians should be prepared to guide them toward accessible, reliable resources. Developing standardized, patient-focused education materials by professional medical societies could ensure access to comprehensible information.
BACKGROUND:Some radiology practices ask patients to pay out of pocket for supplemental artificial intelligence (AI) interpretations of screening mammograms. PURPOSE:To determine if different price points and information conditions influence patient willingness to pay for AI. MATERIALS AND METHODS:Women aged 40+ with a prior mammogram were recruited from an online research platform and asked if they would pay for supplemental AI interpretations after being randomized to different price points ($50, $200, $500) and information conditions (no AI information, two AI accuracy rates, two AI error rates) versus a no AI condition. If respondents declined, they were asked why (cannot afford it versus don't think it is worth it). Statistical analysis assessed permutations of the price points and information conditions. RESULTS:There were 2,534 respondents (median age 53 years [interquartile range: 46-61]). Among information conditions, respondents were most likely to pay for AI when shown an advertisement (26.5%) or good AI accuracy rates (25.3%) and least likely to pay when shown good (14.2%) or poor (7%) AI error rates (P < .0001). Among price conditions, respondents were more likely to pay for less expensive AI (P < .0001): 24.4% at $50, 17.1% at $200, and 13.4% at $500. Reasons for declining AI use varied by information condition and price point (P < .0001). CONCLUSION:A woman's willingness to pay for AI to interpret her mammogram varies according to price and the type of AI information presented. Radiology practices may wish to consider these findings when presenting patients with an option to pay for AI interpretation.
ABSTRACT Background With growing impetus to integrate artificial intelligence (AI) tools into radiology, clinical practices must navigate workflow redesign. This carries implications for medical malpractice liability. Methods We conducted an online vignette experiment with United States adults who acted as hypothetical jurors in a malpractice case involving a missed intracranial hemorrhage. Participants (n=2,347) were randomized to one of 22 conditions: a no-AI control and 21 conditions involving a hypothetical AI system. These twenty-one conditions varied by whether (1) a single-read or double-read workflow was used, (2) the radiologist’s initial interpretation was documented, (3) the radiologist changed their interpretation after viewing AI output, (4) the AI detected the abnormality, and (5) the AI error rate—False Discovery Rate (FDR) or False Omission Rate (FOR)—was provided to participants only, both participants and radiologist, or neither. The primary outcome was perceived liability, assessed by whether the radiologist met their duty of care. Findings Perceived liability differed across conditions (p<0.0001). Double-read workflows (p<0.0001), documenting initial interpretations (p=0.0125), and providing participants with AI error rates, including the FDR (p=0.0038) or FOR (p=0.0035), reduced perceived liability. Liability was also lower when AI was incorrect (p<0.0001). Radiologists’ awareness of AI error rates did not significantly impact liability. Notably, we observed an “erroneous change penalty”: the greatest liability occurred when radiologists initially identified an abnormality but later changed their interpretation to normal after seeing that AI identified the case as normal; conversely, perceived liability was lowest with documented, double-read workflows. Interpretation Double-read workflows with documented initial interpretations and disclosure of AI error rates reduce perceived liability, though changing a correct initial interpretation increases it. Strategic workflow design is critical for successful AI implementation that can mitigate malpractice risk. RESEARCH IN CONTEXT Evidence before this study Emerging research has identified several factors that shape how artificial intelligence (AI) systems are integrated into clinical workflows. Beyond technical performance, factors such as disease prevalence, documentation practices, and workflow design have all been shown to play a role in the implementation of AI tools and how they will inevitably affect physician liability. Vignette experiments have separately identified cognitive biases and mitigating measures that can be integrated into workflow design, such as disclosing AI error rates and documenting independent interpretations before reviewing AI output. Added value of this study This study builds on prior work by examining how multiple aspects of AI implementation influence perceptions of legal liability when hypothetical jurors are asked to adjudicate a medical malpractice lawsuit arising from a false negative interpretation. In particular, we show how combining double-read workflows, documentation practices, and the inclusion of AI error rates can reduce perceived liability. We also show that, when a double-read workflow is used, changing an initially correct interpretation to an incorrect one incurs greater liability for the radiologist than being incorrect in both the initial and final interpretations. These findings underscore the need to address cognitive biases that will undoubtedly arise at the human-AI interface. Implications of all the available evidence Optimizing the integration of AI tools into radiology requires strategic attention to workflow redesign, as combinations of features can collectively affect perceived liability, likely through well-known cognitive biases. Based on the results of our study, we propose one workflow that can mitigate a radiologist’s risk of legal liability. Moving forward, clinical practices and stakeholders should remain cognizant of these factors as they work toward building sustainable AI-physician systems.
The published performance of artificial intelligence (AI) models in radiology is typically based on the reporting of sensitivity, specificity, and receiver operating characteristic area under the curve, both in the peer-reviewed literature and for Food and Drug Administration 510(k) submissions. Interestingly, these metrics cannot inform radiologists, the users of these systems, of the rate or quantity of each type of error to anticipate if they implement the candidate AI product(s) into their own clinical practice. Only the positive predictive value (PPV) and negative predictive value (NPV), or rather, their complements, the false discovery rate (1-PPV, FDR), and the false omission rate (1-NPV, FOR) can provide these error rates. Although some published articles and 510(k) submissions include the PPV and NPV of the concerned AI models, many do not, and the ones that do sometimes test on enriched datasets with artificially high prevalence rates of the target condition, thus inflating PPV and deflating NPV that would be found in clinical practice. This manuscript demonstrates how clinical practices can estimate and evaluate an AI's FDR and FOR for their clinical population using Bayes' Theorem. We also propose a risk-based evaluation matrix (RADDE) which allows radiologists to consider the medical, legal, financial, workflow, psychological, and reputational impact of these estimated AI error rates.
Objective To explore the effect of rationales on placebos described honestly as inactive pills, (open-label placebos; OLPs) on chronic pain. Design Dismantling 4-arm randomized controlled trial. Setting Remote study with United-States residents. Subjects Chronic pain patients aged 18 to 89. Methods We plan to recruit 340 subjects, randomized before consent (Zelen randomization procedure) into one of 4 groups. Participants in a no-treatment group will not receive any OLPs. Participants in the three OLP groups will be told to take an OLP pill twice daily for 21 days. The information participants are given about placebos will vary. Those in the “Standard-OLP” group will be provided with a rationale similar to those used in prior OLP trials. Those in the “Mindfulness-OLP” group will be provided with a rationale taking a mindfulness approach. Those in the “Control-OLP” group (and no-treatment group) will be provided with length-matched information about pain demographics; no placebo information will be given. This dismantling design will allow us to compare rationales (Standard-OLP vs Mindfulness-OLP), and examine the rationale effect (Standard-OLP or Mindfulness-OLP vs. Control-OLP), the placebo effect (Standard-OLP or Mindfulness-OLP vs. No-treatment), and the pill effect (Control-OLP vs no-treatment). Pain intensity over 42 days is the primary outcome. Conclusions This trial will investigate how different components of OLPs impact pain among a chronic pain population. We also highlight novel ways to address limitations of prior OLP studies; namely, lack of blinding and improper controls.
Severity scores, which often refer to the likelihood or probability of a pathology, are commonly provided by artificial intelligence (AI) tools in radiology. However, little attention has been given to the use of these AI scores, and there is a lack of transparency into how they are generated. In this comment, we draw on key principles from psychological science and statistics to elucidate six human factors limitations of AI scores that undermine their utility: (1) variability across AI systems; (2) variability within AI systems; (3) variability between radiologists; (4) variability within radiologists; (5) unknown distribution of AI scores; and (6) perceptual challenges. We hypothesize that these limitations can be mitigated by providing the false discovery rate and false omission rate for each score as a threshold. We discuss how this hypothesis could be empirically tested. KEY POINTS: The radiologist-AI interaction has not been given sufficient attention. The utility of AI scores is limited by six key human factors limitations. We propose a hypothesis for how to mitigate these limitations by using false discovery rate and false omission rate.
Background For more than a decade, studies have supported the efficacy and safety of placebos without deception-so-called open-label placebos (OLPs)-to harness placebo effects in primary care while aligning with key ethical principles. Since treatment acceptance, feasibility, and successful implementation of novel interventions into clinical practice depend on patients' attitudes, patients' perspectives, perceived obstacles, and ideas on OLP use in clinical practice have yet to be elucidated. Therefore, patient and public involvement is increasingly demanded in research and its implementation into clinical practice. Qualitative research offers a unique opportunity to comprehensively understand attitudes, expectations, perceived benefits, and barriers from a patient's point of view. Thus, we studied patients' attitudes, concerns, and ideas toward OLP implementation into clinical practice with focus group discussions (FGDs).Methods In 2022, three exploratory online FGDs, each including two patients with the same condition, were conducted with adult patients affected by chronic back pain (n = 2), chronic migraine (n = 2), or chemotherapy-induced emesis/nausea (n = 2). Physicians recruited participants in three outpatient clinics at the University Hospital Basel in Switzerland. The FGDs were held online for 60 min. Qualitative data was analyzed using Reflexive Thematic Analysis, applying an inductive-deductive hybrid approach within a social constructivist framework.Results In total, five semantic-latent subthemes were identified, entailing: (i) Placebos: Promising but risky; (ii) Acceptance of OLPs depends on a myriad; (iii) Be trustworthy, but deception may be necessary; (iv) Harnessing placebo effects without placebos; (v) From bench to bedside: Clinical transference of OLPs. The themes reflect an in-depth discussion of the usage of OLPs in the clinical context, accompanied by different ambivalences regarding implementation, prerequisites, and the provider role.Conclusion The FGDs provided insights into distinct attitudes, concerns, varying acceptance, and patients' ideas regarding the clinical implementation of OLP interventions. While some patients displayed high acceptance, several concerns regarding ethical and practical issues have been expressed. OLP acceptance and attitudes toward practical issues of OLP intake differed between groups and within the same clinical condition.Trail registration ClinicalTrials.gov, identifier NCT05166213.
Although artificial intelligence (AI) tools are growing in breast imaging, little research has examined how to communicate AI findings in patient portals. English-speaking US women with ≥1 prior mammogram (n = 1623) were randomized to one of thirteen conditions. All participants viewed a radiologist report negative for breast cancer; twelve conditions included an AI report with one of four AI scores (0, 29 [no suspicion of cancer]; 31, 50 [suspicion of cancer]), presented alone, with an abnormality cutoff threshold, or with both the threshold and the AI’s False Discovery Rate (FDR) or False Omission Rate (FOR). Participants reported whether they would consult an attorney for litigation if advanced cancer were detected one year later (primary outcome). Secondary outcomes included follow-up decisions, concern for cancer, and trust. Litigation was higher when AI was included, especially when AI concluded suspicion of cancer. Providing the FDR/FOR reduced litigation; similar effects were observed for follow-up behaviors.
We respond to the Matters Arising article by Bernstein et al. 1 in response to our recent article “Assessing the impact of AI on physician decision-making for mental health treatment in primary care”. Their response highlights the significance of understanding the influence of AI—both unintended and underreported, as we found—on physician decision making, and a few notes regarding aspects of our study’s methodology. We welcome the opportunity to provide additional context to address their comments.
Artificial Intelligence (AI) will have unintended consequences for radiology. When a radiologist misses an abnormality on an image, their liability may differ according to whether or not AI also missed the abnormality. U.S. adults viewed a vignette describing a radiologist being sued for missing a brain bleed (N=652) or cancer (N=682). Participants were randomized to one of five conditions. In four conditions, they were told an AI system was used. Either AI agreed with the radiologist, also failing to find pathology (AI agree) or did find pathology (AI disagree). In the AI agree+FOR condition, AI agreed with the radiologist and an AI false omission rate (FOR) of 1% was presented. In the AI disagree+FDR condition, AI disagreed and an AI false discovery rate (FDR) of 50% was presented. There was also a no AI control condition. Otherwise, vignettes were identical. Participants indicated whether the radiologist met their duty of care as a proxy for whether they would side with defense (radiologist) or plaintiff in trial. Participants were more likely to side with the plaintiff in the AI disagree vs. AI agree condition (brain bleed: 72.9% vs. 50.0%, p=0.0054; cancer: 78.7% vs. 63.5%, p=0.00365) and in the AI disagree vs. no AI condition (brain bleed: 72.9% vs. 56.3%, p=0.0054; cancer: 78.7% vs. 65.2%, p=0.00895). Participants were less likely to side with the plaintiff when FDR or FOR were provided: AI disagree vs AI disagree+FDR (brain bleed: 72.9% vs. 48.8%, p=0.00005; cancer: 78.7% vs. 73.1%, p=0.1507), and AI agree vs. AI agree+FOR (brain bleed: 50.0% vs. 34.0%, p=0.0044; cancer: 63.5% vs. 56.4%, p=0.1085). Radiologists who failed to find an abnormality are viewed as more culpable when they used an AI system that detected the abnormality. Presenting participants with AI accuracy data decreased perceived liability. These findings have relevance for courtroom proceedings. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This study was funded by the Department of Diagnostic Imaging Research at Rhode Island Hospital ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Lifespan IRB #1 of Rhode Island Hospital gave ethical approval for this work I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors
We report our experience implementing an algorithm for the detection of large vessel occlusion (LVO) for suspected stroke in the emergency setting, including its performance, and offer an explanation as to why it was poorly received by radiologists. An algorithm was deployed in the emergency room at a single tertiary care hospital for the detection of LVO on CT angiography (CTA) between September 1st–27th, 2021. A retrospective analysis of the algorithm’s accuracy was performed. During the study period, 48 patients underwent CTA examination in the emergency department to evaluate for emergent LVO, with 2 positive cases (60.3 years ± 18.2; 32 women). The LVO algorithm demonstrated a sensitivity and specificity of 100
Some of the most controversial information in psychology involves genetic or evolutionary explanations for sex differences in educational-vocational outcomes (Clark et al., 2024a). We investigated whether men and women react differently to controversial information about sex differences and whether their reaction depends on who provides the information. In the experiment, college students (n=396) and U.S. middle-aged adults (n=154) reviewed a handout, purportedly provided by either a male or a female professor. The handout stated that (1) women in STEM are no longer discriminated against in hiring and publishing and (2) sex differences in educational-vocational outcomes are better explained by evolved differences between men and women in various personal attributes. We found that college women were less receptive to the information than college men were and wanted to censor it more than men did; also, in both the college student and community adult samples, women were less receptive and more censorious when the messenger was a male professor than when the messenger was a female professor. In both samples, participants who leaned to the left politically and who held stronger belief that words can cause harm reacted with more censoriousness. Our findings imply that the identity of a person presenting controversial scientific information and the receiver’s pre-existing identity and beliefs have the potential to influence how that information will be received.
Randomised placebo-controlled trials (RPCTs) are the gold standard for evaluating novel treatments. However, this design is rarely used in the context of orthopaedic interventions where participants are assigned to a real or placebo surgery. The present study examines attitudes towards RPCTs for orthopaedic surgery among 687 orthopaedic surgeons across the USA. When presented with a vignette describing an RPCT for orthopaedic surgery, 52.3% of participants viewed it as 'completely' or 'mostly' unethical. Participants were also asked to rank-order the value of five different types of evidence supporting the efficacy of a surgery, ranging from RPCT to an anecdotal report. Responses regarding RPCTs were polarised with 26.4% viewing it as the least valuable (even less valuable than an anecdote) and 35.7 .% viewing it as the most valuable. Where equipoise exists, if we want to subject orthopaedic surgeries to the highest standard of evidence (RPCTs) before they are implemented in clinical practice, it will be necessary to educate physicians on the value and ethics of placebo surgery control conditions. Otherwise, invasive procedures may be performed without any benefits beyond possible placebo effects.