This research applies the People, Process, Technology, and Operations (PPTO) framework to develop AI governance within a large hospital system in Canada that is early in AI adoption. Stakeholder interviews identified the organization’s strengths, gaps, and priorities for AI governance, providing foundational insights into the organization’s readiness and needs. Co-design workshops then adapted the PPTO framework to the organization’s specific context. Together, these efforts led to the creation of policies and the formation of an AI governance committee within the organization. This work demonstrates that the PPTO framework is a practical and adaptable tool for developing AI governance in real-world healthcare settings. It also addresses a critical gap in the field by generating empirical evidence of how a conceptual AI governance framework can be implemented within healthcare delivery organizations to drive organizational change.
Purpose: This shadow deployment evaluated an externally-developed AI tool to predict disposition using chest X-rays (CXR) in patients with community-acquired pneumonia (CAP) in the Emergency Department (ED). Retrospective and prospective external validations were conducted to assess differences between the 2 evaluations and across subgroups to inform deployment decisions. Methods: The CNN was retrospectively validated (n = 17 689) from November 1, 2020, to June 30, 2021, and prospectively validated on "suspected-CAP" patients (n = 3062) from Jan 1 to Jan 31, 2023. Calibration and standard metrics, including AUC, accuracy, sensitivity, specificity, PPV, and NPV, were calculated. Subgroup analyses were conducted for age, sex, modality, and CXR projection (PA vs AP). Results: The model's AUC was 67% in both validations. The prospective evaluation showed a non-significant increase in sensitivity (65% vs 59%) and PPV (64% vs 63%), while specificity (68% vs 73%) and NPV (69% vs 70%) slightly decreased. NPV was very high for younger patients in the prospective evaluation (95%); PPV was moderately high for older patients (81%). Sensitivity dropped significantly in females under 31 years (50%), and specificity was reduced in females over 86 years (38%). Conclusion: This study showed moderate, consistent performance in both retrospective and prospective validations. While this consistency is encouraging, further direct comparisons are needed to determine whether both validation approaches are necessary in different clinical settings. Subgroup analysis suggests the tool may be helpful to accelerate discharge in younger patients (high NPV) and possibly for admission in older patients (high PPV).
Objective: This study introduces an innovative end-to-end deep learning pipeline designed to automatically classify and order fetal ultrasound standard planes in alignment with the guidelines of the Canadian Association of Radiologists, while also assessing the diagnostic usability of each view. The primary objective is to address the manual and cumbersome challenges that interpreting radiologists encounter in the existing obstetric ultrasound workflow. Methods: We compiled a diverse dataset, comprising 33,561 de-identified two-dimensional obstetrical ultrasound images acquired from January 1, 2010, to June 1, 2020. This dataset was categorized into 19 distinct classes associated with standard planes and further partitioned into training, validation, and testing subsets via a 60:20:20 stratified split. The standard plane and diagnostic usability networks are founded on a convolutional neural network framework and employ the benefits of transfer learning. Results: The standard plane classification network demonstrated promising results by achieving 99.4 % and 98.7 % for accuracy and F1 score, respectively. Subsequently, the diagnostic usability network demonstrated strong performance, registering 80 % accuracy and an 82 % F1 score. Notably, this study is the first to investigate whether deep learning methods can surpass sonographers in the standard plane labeling task, with some instances revealing the algorithm's capacity to rectify sonographer mislabeled planes. Conclusion: The results highlight the algorithm's potential to be integrated into a clinical setting by serving as a reliable assistive tool, alleviating the cognitive workload faced by radiologists and enhancing efficiency and diagnostic outcomes in the current obstetric ultrasound process.
While it is common to monitor deployed clinical artificial intelligence (AI) models for performance degradation, it is less common for the input data to be monitored for data drift – systemic changes to input distributions. However, when real-time evaluation may not be practical (eg., labeling costs) or when gold-labels are automatically generated, we argue that tracking data drift becomes a vital addition for AI deployments. In this work, we perform empirical experiments on real-world medical imaging to evaluate three data drift detection methods’ ability to detect data drift caused (a) naturally (emergence of COVID-19 in X-rays) and (b) synthetically. We find that monitoring performance alone is not a good proxy for detecting data drift and that drift-detection heavily depends on sample size and patient features. Our work discusses the need and utility of data drift detection in various scenarios and highlights gaps in knowledge for the practical application of existing methods.
Despite frequent reports of imaging artificial intelligence (AI) that parallels human performance, clinicians often question the safety and robustness of AI products in practice. This work explores two underreported sources of noise that negatively affect imaging AI: (a) variation in labeling schema definitions and (b) noise in the labeling process. First, the overlap between the schemas of two publicly available datasets and a third-party vendor are compared, showing there is low agreement (<50%) between them. The authors also highlight the problem of label inconsistency, where different annotation schemas are selected for the same clinical prediction task; this results in inconsistent use of medical ontologies through intermingling or duplicate observations and diseases. Second, the individual radiologist annotations for the CheXpert test set are used to quantify noise in the labeling process. The analysis demonstrated that label noise varies by class, as agreement was high for pneumothorax and medical devices (percent agreement > 90%). Among low agreement classes (pneumonia, consolidation), the labels assigned as "ground truth" were unreliable, suggesting that the result of majority voting is highly dependent on which group of radiologists is assigned to annotation. Noise in labeling schemas and gold label annotations are pervasive in medical imaging classification and affect downstream clinical deployment. Possible solutions (eg, changes to task design, annotation methods, and model training) and their potential to improve trust in clinical AI are discussed. Keywords: Radiology AI, Dataset Creation, Noise in Datasets Supplemental material is available for this article. © RSNA, 2023 See also the commentary by Ursprung and Woitek in this issue.
Clinical AI model reporting cards should be expanded to incorporate a broad bias reporting of both social and non-social factors. Non-social factors consider the role of other factors, such as disease dependent, anatomic, or instrument factors on AI model bias, which are essential to ensure safe deployment.
Artificial intelligence (AI) software in radiology is becoming increasingly prevalent and performance is improving rapidly with new applications for given use cases being developed continuously, oftentimes with development and validation occurring in parallel. Several guidelines have provided reporting standards for publications of AI-based research in medicine and radiology. Yet, there is an unmet need for recommendations on the assessment of AI software before adoption and after commercialization. As the radiology AI ecosystem continues to grow and mature, a formalization of system assessment and evaluation is paramount to ensure patient safety, relevance and support to clinical workflows, and optimal allocation of limited AI development and validation resources before broader implementation into clinical practice. To fulfil these needs, we provide a glossary for AI software types, use cases and roles within the clinical workflow; list healthcare needs, key performance indicators and required information about software prior to assessment; and lay out examples of software performance metrics per software category. This conceptual framework is intended to streamline communication with the AI software industry and provide healthcare decision makers and radiologists with tools to assess the potential use of these software. The proposed software evaluation framework lays the foundation for a radiologist-led prospective validation network of radiology AI software. Learning Points: The rapid expansion of AI applications in radiology requires standardization of AI software specification, classification, and evaluation. The Canadian Association of Radiologists' AI Tech & Apps Working Group Proposes an AI Specification document format and supports the implementation of a clinical expert evaluation process for Radiology AI software.
Purpose: To observe interactions of practicing radiologists with a chest x-ray AI tool and evaluate its usability and impact on workflow efficiency. Methods: Using a simulated clinical workflow and remote multi-monitor screensharing, we prospectively assessed the interactions of 10 staff radiologists (5-33 years of experience) with a PACS-embedded, regulatory-approved chest x-ray AI tool. Qualitatively, we collected feedback using a think-aloud method and post-testing semi-structured interview; transcript themes were categorized by: (1) AI tool features, (2) deployment considerations, and (3) broad human-AI interactions. Quantitatively, we used time-stamped video recordings to compare reporting and decision-making efficiency with and without AI assistance. Results: For AI tool features, radiologists appreciated the simple binary classification (normal vs abnormal) and found the heatmap essential to understand what the AI considered abnormal; users were uncertain of how to interpret confidence values. Regarding deployment considerations, radiologists thought the tool would be especially helpful for identifying subtle diagnoses; opinions were mixed on whether the tool impacted perceived efficiency, accuracy, and confidence. Considering general human-AI interactions, radiologists shared concerns about automation bias especially when relying on an automated triage function. Regarding decision-making and workflow efficiency, participants began dictating 5 seconds later (42% increase, P = .02) and took 14 seconds longer to complete cases (33% increase, P = .09) with AI assistance. Conclusions: Radiologist usability testing provided insights into effective AI tool features, deployment considerations, and human-AI interactions that can guide successful AI deployment. Early AI adoption may increase radiologists' decision-making and total reporting time but improves with experience.
Purpose: To externally test four chest radiograph classifiers on a large, diverse, real-world dataset with robust subgroup analysis.Materials and Methods: In this retrospective study, adult posteroanterior chest radiographs (January 2016-December 2020) and associ-ated radiology reports from Trillium Health Partners in Ontario, Canada, were extracted and de-identified. An open-source natural language processing tool was locally validated and used to generate ground truth labels for the 197 540-image dataset based on the associated radiology report. Four classifiers generated predictions on each chest radiograph. Performance was evaluated using accuracy, positive predictive value, negative predictive value, sensitivity, specificity, F1 score, and Matthews correlation coefficient for the overall dataset and for patient, setting, and pathology subgroups.Results: Classifiers demonstrated 68%-77% accuracy, 64%-75% sensitivity, and 82%-94% specificity on the external testing dataset. Algorithms showed decreased sensitivity for solitary findings (43%-65%), patients younger than 40 years (27%-39%), and patients in the emergency department (38%-60%) and decreased specificity on normal chest radiographs with support devices (59%-85%). Differences in sex and ancestry represented movements along an algorithm's receiver operating characteristic curve.Conclusion: Performance of deep learning chest radiograph classifiers was subject to patient, setting, and pathology factors, demonstrat-ing that subgroup analysis is necessary to inform implementation and monitor ongoing performance to ensure optimal quality, safety, and equity.
INTRODUCTION:We aimed to develop an explainable machine learning (ML) model to predict side-specific extraprostatic extension (ssEPE) to identify patients who can safely undergo nerve-sparing radical prostatectomy using preoperative clinicopathological variables.METHODS:A retrospective sample of clinicopathological data from 900 prostatic lobes at our institution was used as the training cohort. Primary outcome was the presence of ssEPE. The baseline model for comparison had the highest performance out of current biopsy-derived predictive models for ssEPE. A separate logistic regression (LR) model was built using the same variables as the ML model. All models were externally validated using a testing cohort of 122 lobes from another institution. Models were assessed by area under receiver-operating-characteristic curve (AUROC), precision-recall curve (AUPRC), calibration, and decision curve analysis. Model predictions were explained using SHapley Additive exPlanations. This tool was deployed as a publicly available web application.RESULTS:Incidence of ssEPE in the training and testing cohorts were 30.7 and 41.8%, respectively. The ML model achieved AUROC 0.81 (LR 0.78, baseline 0.74) and AUPRC 0.69 (LR 0.64, baseline 0.59) on the training cohort. On the testing cohort, the ML model achieved AUROC 0.81 (LR 0.76, baseline 0.75) and AUPRC 0.78 (LR 0.75, baseline 0.70). The ML model was explainable, well-calibrated, and achieved the highest net benefit for clinically relevant cutoffs of 10-30%.CONCLUSIONS:We developed a user-friendly application that enables physicians without prior ML experience to assess ssEPE risk and understand factors driving these predictions to aid surgical planning and patient counselling (https://share.streamlit.io/jcckwong/ssepe/main/ssEPE_V2.py).
Introduction: We aimed to develop an explainable machine learning (ML) model to predict side-specific extraprostatic extension (ssEPE) to identify patients who can safely undergo nerve-sparing radical prostatectomy using preoperative clinicopathological variables. Methods: A retrospective sample
In this paper, we demonstrate the use of a “Challenge Dataset”: a small, site-specific, manually curated dataset – enriched with uncommon, risk-exposing, and clinically important edge cases – that can facilitate pre-deployment evaluation and identification of clinically relevant AI performance deficits. The five major steps of the Challenge Dataset process are described in detail, including defining use cases, edge case selection, dataset size determination, dataset compilation, and model evaluation. Evaluating performance of four chest X-ray classifiers (one third-party developer model and three models trained on open-source datasets) on a small, manually curated dataset (410 images), we observe a generalization gap of 20.7% (13.5% - 29.1%) for sensitivity and 10.5% (4.3% - 18.3%) for specificity compared to developer-reported values. Performance decreases further when evaluated against edge cases (critical findings: 43.4% [27.4% - 59.8%]; unusual findings: 45.9% [23.1% - 68.7%]; solitary findings 45.9% [23.1% - 68.7%]). Expert manual audit revealed examples of critical model failure (e.g., missed pneumomediastinum) with potential for patient harm. As a measure of effort, we find that the minimum required number of Challenge Dataset cases is about 1% of the annual total for our site (approximately 400 of 40,000). Overall, we find that the Challenge Dataset process provides a method for local pre-deployment evaluation of medical imaging AI models, allowing imaging providers to identify both deficits in model generalizability and specific points of failure prior to clinical deployment.### Competing Interest StatementBF is a shareholder of Pocket Health and Eva Center and has received consultant fees from Canon Medical. JS, MA, MA, LSK, and SM have no conflicts of interest to disclose.### Funding StatementThis work was funded by Canada's Digital Technology Supercluster. The funder and AI developer had no role in the study design or analysis.### Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesThe details of the IRB/oversight body that provided approval or exemption for the research described are given below:The research ethics board of Trillium Health Partners gave ethical approval for this workI confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable.YesThe dataset from this study is held securely at THP. Coded / aggregate data can be made accessible (contact senior author).
BACKGROUND Nonpharmaceutical interventions (NPIs) are the primary tools to mitigate early spread of the coronavirus disease 2019 (COVID-19) pandemic; however, such policies are implemented variably at the federal, provincial or territorial, and municipal levels without centralized documentation. We describe the development of the comprehensive open Canadian Non-Pharmaceutical Intervention (CAN-NPI) data set, which identifies and classifies all NPIs implemented in regions across Canada in response to COVID-19, and provides an accompanying description of geographic and temporal heterogeneity. METHODS We performed an environmental scan of government websites, news media and verified government social media accounts to identify NPIs implemented in Canada between Jan. 1 and Apr. 19, 2020. The CAN-NPI data set contains information about each intervention's timing, location, type, target population and alignment with a response stringency measure. We conducted descriptive analyses to characterize the temporal and geographic variation in early NPI implementation. RESULTS We recorded 2517 NPIs grouped in 63 distinct categories during this period. The median date of NPI implementation in Canada was Mar. 24, 2020. Most jurisdictions heightened the stringency of their response following the World Health Organization's global pandemic declaration on Mar. 11, 2020. However, there was variation among provinces or territories in the timing and stringency of NPI implementation, with 8 out of 13 provinces or territories declaring a state of emergency by Mar. 18, and all by Mar. 22, 2020. INTERPRETATION There was substantial geographic and temporal heterogeneity in NPI implementation across Canada, highlighting the importance of a subnational lens in evaluating the COVID-19 pandemic response. Our comprehensive open-access data set will enable researchers to conduct robust interjurisdictional analyses of NPI impact in curtailing COVID-19 transmission.
Seamless sharing between imaging facilities of medical images obtained on the same patient is crucial in providing accurate and efficient care to patients. However, the terminology used to describe semantically similar examinations can vary widely between facilities. Current practice is manual table-based mapping to a standard terminology, which has substantial potential for mislabelled and missing examinations. In this work, we establish several baseline methods for automating the mapping of radiology imaging procedure descriptions to a SNOMED CT based standard terminology. Our best performing baseline, consisting of a bag of words representation and shallow neural network, achieved 96.3% accuracy. In addition, we explore an unsupervised clustering method that explores relevancy matching without the need for an intervening standard. Lastly, we make the procedure name dataset used in this work available to encourage extension of this application.
PURPOSE:The aim of this study was to enhance multispecialty CT and MRI protocol assignment quality and efficiency through development, testing, and proposed workflow design of a natural language processing (NLP)-based machine learning classifier.METHODS:NLP-based machine learning classification models were developed using order entry input data and radiologist-assigned protocols from more than 18,000 unique CT and MRI examinations obtained during routine clinical use. k-Nearest neighbor, random forest, and deep neural network classification models were evaluated at baseline and after applying class frequency and confidence thresholding techniques. To simulate performance in real-world deployment, the model was evaluated in two operating modes in combination: automation (automated assignment of the top result) and clinical decision support (CDS; top-three protocol suggestion for clinical review). Finally, model-radiologist discordance was subjectively reviewed to guide explainability and safe use.RESULTS:Baseline protocol assignment performance achieved weighted precision of 0.757 to 0.824. Simulating real-world deployment using combined thresholding techniques, the optimized deep neural network model assigned 69% of protocols in automation mode with 95% accuracy. In the remaining 31% of cases, the model achieved 92% accuracy in CDS mode. Analysis of discordance with subspecialty radiologist labels revealed both more and less appropriate model predictions.CONCLUSIONS:A multiclass NLP-based classification algorithm was designed to drive local operational improvement in CT and MR radiology protocol assignment at subspecialist quality. The results demonstrate a simulated workflow deployment enabling automated assignment of protocols in nearly 7 of 10 cases with very few errors combined with top-three CDS for remaining cases supporting a high-quality, efficient radiology workflow.
Non-pharmaceutical interventions (NPIs) have been the primary tool used by governments and organizations to mitigate the spread of the ongoing pandemic of COVID-19. Natural experiments are currently being conducted on the impact of these interventions, but most of these occur at the subnational level - data not available in early global datasets. We describe the rapid development of the first comprehensive, labelled dataset of 1640 NPIs implemented at federal, provincial/territorial and municipal levels in Canada to guide COVID-19 research. For each intervention, we provide: a) information on timing to aid in longitudinal evaluation, b) location to allow for robust spatial analyses, and c) classification based on intervention type and target population, including classification aligned with a previously developed measure of government response stringency. This initial dataset release (v1.0) spans January 1st, and March 31st, 2020; bi-weekly data updates to continue for the duration of the pandemic. This novel dataset enables robust, inter-jurisdictional comparisons of pandemic response, can serve as a model for other jurisdictions and can be linked with other information about case counts, transmission dynamics, health care utilization, mobility data and economic indicators to derive important insights regarding NPI impact.
Choosing Wisely campaigns have spread to more than 20 countries worldwide. Choosing Wisely is a clinician-led effort to produce clinician recommendations and patient resources around overused tests, procedures and medications. The digitization of medical records and workflows and growth in computing power has enabled artificial intelligence (AI) to be applied to augment health care decision-making. Overuse is an important problem for AI-enabled technologies to address. This commentary offers a roadmap of opportunities for AI-augmented efforts to reduce overuse which are presented according to a patient’s journey of care, beginning from when patients experience symptoms through clinical decision making and subsequent evaluation efforts. This roadmap can be used to guide the cross-discipline development and implementation of AI-enabled tools to reduce overuse and drive value.
In concurrence with the introduction of the internet, widely networked computers, and the collection of large amounts of digital data, the medical profession as a whole has become more self-aware and self-critical. It is increasingly apparent that suboptimal decisions are made at times and, on other occasions, are fatally flawed. Most clinical decisions rest largely on what is referred to as the art of medicine: that is, decision-making that is based on inconsistent and incomplete provider knowledge; variable skills, training, and experience; and last but not the least, an array of biases. Unsurprisingly, the result is an unacceptable degree of care variation that is not explained by patient factors or the clinical context.1Wennberg J Unwarranted variation in healthcare delivery: implications for academic medical centres.BMJ. 2001; 325: 961-964Crossref Scopus (537) Google Scholar Every minute, a medical decision is being made somewhere that could be more informed, more objective, more precise, and more safe. How does medicine move on to adapt to an era of big data and a need to make consistent, data driven, evidence and value-based clinical decisions? Artificial intelligence (AI) refers to the ability of computers to learn the associations within troves of data to assist with classification (eg, diagnosis), prediction (eg, triaging and prognostication), and optimisation (eg, precision treatment). Studies have reviewed current applications of AI, as well as the opportunities and challenges it poses in the field of health care.2The Lancet Digital HealthWalking the tightrope of artificial intelligence guidelines in clinical practice.Lancet Digit Health. 2019; 1: e100Summary Full Text Full Text PDF Scopus (13) Google Scholar However, it is important to reiterate the vast potential for AI beyond the use in medical imaging, in a wide array of disparate clinical situations. For example, AI has long been used in the development of severity scoring systems, but could AI assist the psychiatrist in assessing suicide risk3Simon GE Johnson E Lawrence JM et al.Predicting suicide attempts and suicide deaths following outpatient visits using electronic health records.Am J Psychiatry. 2018; 175: 951-960Crossref PubMed Scopus (171) Google Scholar or assist the pediatrician in the recognition of rare genetic syndromes?4Hsieh T-C Mensah MA Pantel JT et al.PEDIA: prioritization of exome data by image analysis.Genet Med. 2019; (published online June 5.)DOI:10.1038/s41436-019-0566-2Summary Full Text Full Text PDF PubMed Scopus (30) Google Scholar AI is already being built into the development of physiological monitors and has been proposed as an aid to decisions such as the treatment of sepsis, assessment of readmission risk, and recognition of consciousness in unresponsive patients.5Rush B Celi LA Stone DJ Applying machine learning to continuously monitored physiological data.J Clin Monit Comput. 2018; (published online Nov 11.)DOI:10.1007/s10877-018-0219-zPubMed Google Scholar, 6Komorowski M Celi LA Badawi O Gordon AC Faisal AA The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care.Nat Med. 2018; 241716Crossref PubMed Scopus (393) Google Scholar, 7Badawi O Breslow MJ Readmissions and death after ICU discharge: development and validation of two predictive models.PloS one. 2012; 7e48758Crossref PubMed Scopus (67) Google Scholar, 8Claassen J Doyle K Matory A et al.Detection of brain activation in unresponsive patients with acute brain injury.New Engl J Med. 2019; 380: 2497-2505Crossref PubMed Scopus (173) Google Scholar One can envision the possibilities of AI guidance and support in the care of patients with complex conditions, such as those with multiple, chronic medical problems, or the decision to proceed to major surgery in fragile, complicated patients. The question is not whether computers can outperform humans in specific tasks, but how humanity will embrace and adopt these capabilities into the practice of medicine. For AI to achieve adoption in medicine, hospitals and health-care systems must be willing to buy it, and patients and providers must accept it. The AI task must be deemed useful to patients, providers, and payors. If the gains are trivial, the unsystematic and uncontrolled mushrooming of minimal value-added medical AI algorithms will only add to the over-referral, over-diagnosis, over-treatment and, ultimately, to overall health-care costs, as well as a distrust of the technology. With compelling use cases defined, one can get to the challenge of building trusted models. The development of AI algorithms requires data extraction and integration that is engineering intensive, as well as data standardisation and curation that is clinical domain expertise intensive. Like the scientist to clinician partnership in translational medicine, a new partnership of data scientist, engineers, and clinicians needs to be developed. However, this process is wrought with challenges, both technical and non-technical. The most fundamental challenge to the development and implementation of machine learning in health-care is access to reliable, well curated data. Clinical data usually resides across different systems, often locked in a proprietary format requiring additional costly software for extraction. And once data has been freed from proprietary servers, health data standards and common data models for research purposes are far from ideal, creating unnecessary work to unlock the value in the data. The task of establishing these standards requires not only expertise in medical informatics but time-consuming input from domain experts and researchers who might have little interest in providing this important housekeeping task. It would be worthwhile to learn from the experience of the Laboratory for Computational Physiology at the Massachusetts Institute of Technology who have created the publicly available and greatly used Medical Information Mart for Intensive Care (MIMIC) database and the eICU Collaborative Research Database, and have crowd-sourced the data curation across MIMIC's 12,000 users.9Johnson AEW Pollard TJ Shen L et al.MIMIC-III, a freely accessible critical care database.Sci data. 2016; 3160035Crossref PubMed Scopus (2986) Google Scholar With trusted models established, AI adoption will then require addressing the practical, but too often overlooked, matter of designing AI applications that are intelligently integrated with clinical workflows. Following the example of the current generation of electronic medical records, burnout has often followed as providers struggled with usability. Poorly designed AI could analogously worsen information overload and cognitive fatigue. The controversial IBM Watson-MD Anderson partnership is an example of the difficulties encountered in applying AI to complex clinical issues. Will it require decades of trial and (unfortunately, sometime disastrous) errors in human-machine interfaces? To achieve adoption, AI will need to be an invisible, seamless, and unbiased aid, helping patients and physicians make better decisions in an efficient, effective, and acceptable manner.9Johnson AEW Pollard TJ Shen L et al.MIMIC-III, a freely accessible critical care database.Sci data. 2016; 3160035Crossref PubMed Scopus (2986) Google Scholar It is clear that humans can trust machines for particular decisions, such as in airline travel today. But software has wrestled with pilots and won – with catastrophic results.10Hatton L Rutkowski A "Lessons must be learned"—but are they?.IEEE software. 2019; 36: 91-95Crossref Scopus (8) Google Scholar The Boeing 737 control issues showed that any changes made to a functional system must be well communicated, completely tested, and accompanied by a thorough education of all involved in the use of these systems. Clinical practice should evolve as a hybrid enterprise with clinicians who know what to expect from, and how to work with, what is fundamentally a very sophisticated clinical support tool. Working together, humans and machines can address many of the decisional fragilities intrinsic to current practice. The human-driven scientific method can be powerfully augmented by computational methods sifting through the necessarily large amounts of longitudinal patient- and provider-generated data. But the guidance of these methods also requires more precise definitions of some of the foundational principles in medicine: eg, what is normal, what abnormalities require clinical intervention, what outcomes are we trying to achieve, and what costs are acceptable to do so? For example, defining normal is fundamentally important when considering the use of AI for evaluating chest x-rays. Today, we use a radiologist's opinion as the gold standard. But it is also crucial to know which radiological abnormalities are significant and require intervention versus those which represent clear-cut overdiagnoses with the potential for overtreatment generating unnecessary costs, complications, and suffering. Similarly, for diagnosis of diabetic retinopathy from fundus photographs, the use of a consensus agreement standard with ophthalmologists might not be the gold standard. One opportunity is to mine huge, longitudinal, linked population datasets (including imaging results, treatments, and clinical and patient related outcomes) to fully define what constitutes normal in various contexts. A single conclusion might not apply in all circumstances. In the intensive care unit, for example, in each individual case we need to establish the value that the patient and family place on simply extending survival versus attempting to provide quality of remaining life. In obstetrics, what is the optimal outcome we are trying to achieve in terms of the decision to employ a caesarian section or not? Is it a minimal rate of surgery, a minimal rate of fetal mortality, or a minimal rate of maternal mortality? These are complex, non-trivial issues in terms of providing such decision support to the obstetrician. One of the indirect benefits of AI might be in forcing us to clearly define the major challenges in health care in a way that we have not been forced to do before. Anyone who has attempted to develop clinical software understands that coding an application requires so-called black and white solutions (eg, symptoms yes or no, disease present or absent) that are often difficult to obtain in medicine. We can only design AI for those issues in which we possess a precise, contextual, and optimally complete level of understanding of health and disease. And after we do develop AI solutions, we will still need to continue to provide supervision using human intelligence with all its quirks, inconsistencies, and potential for deficits in situational awareness. AI is not going to produce a perfect medical system, but if thoughtfully designed and implemented, it has the potential to produce a better one. We declare no competing interests. We thank Dr Mazyeh Ghassemi for inspiring this work during a debate at the Machine Learning for Health Unconference in Toronto in May 28, 2019.
BACKGROUND:In 2012, the Ontario government withdrew public insurance coverage of imaging tests for uncomplicated low back pain. We studied the impact of this restriction on test ordering by physicians.METHODS:We compared the numbers of lumbar spine radiography, computed tomography (CT) and single-segment magnetic resonance imaging (MRI) studies ordered by physicians in the 3 years before and after the policy change. We linked claims data from the Ontario Health Insurance Program with physician details to calculate rates per test-ordering physician. We compared changes in rates of monthly test ordering by family physicians and specialists before and after the policy change using segmented regression analysis of interrupted time series data.RESULTS:The number of lumbar spine radiography and spine CT studies ordered by family physicians decreased by 98 597 (28.7%) and 17 499 (28.7%), respectively, in the year after the policy change; there was little change in ordering by specialists. The number of lumbar spine radiography studies ordered per family physician by month decreased by 0.81 tests (p < 0.001) after the intervention, followed by a smaller rebound increase that remained below baseline. Monthly ordering of spine CT per family physician declined by 0.1 tests (p < 0.001), and that of limited spine MRI rose before the intervention, decreased by 0.18 tests (p < 0.001) after the intervention, then started to rise again. Monthly ordering of limited spine MRI by specialists, which had been stable before the policy change, decreased by 0.1 tests per specialist (p < 0.001) afterward, then rose to preintervention levels.INTERPRETATION:The restriction in coverage of imaging tests caused a larger decrease in test ordering by family physicians than by specialists and a larger, more sustained reduction in the use of lumbar spine radiography and spine CT than of spine MRI.