As artificial intelligence (AI), machine learning (ML), and other forms of advanced automation are increasingly considered for deployment in safety-critical industries, there is an urgent need for evaluation methods which reliably identify risks of deployment prior to people being harmed. In this narrative review, we discuss the benefits and drawbacks of 11 major methodological decisions underpinning evaluations of AI-infused technologies from the perspective of cognitive systems engineering (CSE) and naturalistic decision making (NDM). These methodological decisions are organized around four aspirations central to the perspective of CSE and NDM: evaluations of AI-infused technologies should be (1) integrated, (2) naturalistic, (3) grounded, and (4) pattern-centered. We use these aspirations to interpret common human-AI evaluation methods and discuss new evaluation challenges for emerging AI-infused technologies. This narrative review is meant to guide both current methods and future research toward safe and effective strategies for evaluating AI-infused technologies, especially in safety-critical settings.
BackgroundConsumer fitness tracker devices offer scalable opportunities to monitor real-world behavior and support health in cancer survivorship. However, adoption and sustained use outside structured research settings remain incompletely characterized, limiting their integration into survivorship care. ObjectiveThe primary objectives were to describe the prevalence of fitness tracker ownership and use patterns among cancer survivors and to identify sociodemographic, psychosocial, and usability-related correlates of device ownership and frequent use. MethodsWe conducted a cross-sectional survey of 893 cancer survivors enrolled in the Total Cancer Care protocol at a comprehensive cancer center. Participants completed an adapted online questionnaire assessing fitness tracker ownership, frequency of use, and perceived barriers and facilitators. Multivariable logistic regression models were used to identify sociodemographic, psychosocial, and usability-related correlates of device ownership and frequent use. ResultsMore than half of the participants (506/893, 56.7%) reported owning a fitness tracker, and among owners, 82.2% (416/506) reported frequent use (most days or every day), including 71.3% (361/506) who wore the device every day. The most commonly used devices were the Apple iWatch (272/506, 53.8%) and Fitbit (147/506, 29.1%). In multivariable analyses, fitness tracker ownership was independently associated with sex, household income, and cancer site. Male participants had lower odds of ownership (adjusted odds ratio [aOR] 0.61, 95% CI 0.42-0.87; P=.006), while higher household income was associated with greater ownership (US $50,000-$99,999: aOR 1.70, 95% CI 1.13-2.57; P=.01; ≥US $100,000: aOR 3.64, 95% CI 2.41-5.50; P<.001). Ownership also differed by cancer site. Among owners, frequent use was concurrently and inversely associated with self-reported device discomfort (P<.001), low motivation (P<.001), information overload (P=.04), and limited app integration (P=.007). ConclusionsIn this single-center, predominantly White, higher-income sample of cancer survivors, fitness tracker ownership was common and patterned by demographic and socioeconomic characteristics, while sustained engagement among device owners was associated with psychosocial and usability factors. These findings suggest that scalable, fitness tracker–enabled survivorship care will need to address both structural disparities in ownership as well as behavioral readiness and user experience to ensure clinically meaningful implementation.
Methods to evaluate the potential benefits and risks of human-AI systems in high-risk domains remain an acknowledged research gap. Joint Activity Testing (JAT) has shown considerable promise to reveal these benefits and risks. However, applications of JAT have been currently restricted to static scenarios, limiting the applicability to dynamic settings. In this work, we expand JAT techniques to incorporate a sequence of decisions in a temporally extended task: detecting active satellite maneuvers. We report the results of a pilot study to assess the feasibility of JAT for evaluating space situational awareness tasks. Our results show distinct performance–challenge relationships across scenarios, suggesting the value of utilizing JAT to assess dynamic joint human-AI event detection.
Joint activity describes when more than one agent (human or machine) contributes to the completion of a task or activity. Designing for joint activity focuses on explicitly supporting the interdependencies between agents necessary for effective coordination among agents engaged in the joint activity. This builds and expands upon designing for usability to further address how technologies can be designed to act as effective team players. Effective joint activity requires supporting, at minimum, five primary macrocognitive functions within teams: Event Detection, Sensemaking, Adaptability, Perspective-Shifting, and Coordination. Supporting these functions is equally as important as making technologies usable. We synthesized fourteen heuristics from relevant literature including display design, human factors, cognitive systems engineering, cognitive psychology, and computer science to aid the design, development, and evaluation of technologies that support joint human-machine activity.
Coordination underpins adaptive capacity in resilient systems. However, research on specific patterns of coordination as they relate to adaptive capacity is relatively scarce. The study focuses on the Sterile Processing Department (SPD) as an exemplar of a large complex system with significant coordination challenges. The study aimed to identify adaptive performance under countervailing pressures, and coordination patterns that enable adaptive capacity. Six observations and eighteen interviews of OR and SPD workers were conducted. Data were analyzed for themes and inherent semantic relationships. Schematic diagrams were developed to compactly represent the adaptive workflow patterns along with the associated interactions supporting coordination. Coordination patterns, such as concatenated loops, were identified as markers of system brittleness. Implications for supporting coordination and adaptive capacity under complexity are provided.
There is an urgent need for human factors methods which can match the pace and scale of AI development. In this study, we compared two data collection mechanisms to probe nurses’ understanding of an AI-infused patient data display. We collected responses to equivalent questions in two alternative formats: (1) unstructured free text responses and (2) structured multi-select responses. We found nurses were significantly more likely to report 11 of 15 response categories with the multi-select questions. Additionally, we found nurses exhibited significantly different behaviors interacting with the display; however, we found little evidence of differences in task performance. Our findings emphasize that structuring data collection mechanisms to support quick data processing can induce different responses and behaviors that are not directly comparable to unstructured responses.
Designing AI-infused technologies in healthcare involves multidisciplinary design considerations, like the usability of the display and explainability of the AI algorithm. We present three vignettes from a case study in which design decisions intended to improve usability inadvertently undermined explainability, and vice versa. Our results illustrate potential conflicts between disciplines and suggest the need for joint activity as a unifying design perspective. From these insights, we discuss revisions to our AI-infused technology.
Abstract Background Despite the growth of digital healthcare data, surveillance for healthcare associated infections (HAIs) often consists of reviewing cases using lists, charts, or tables. For infections, such as C.diff which can be transmitted patient-to-environment and patient-to-patient, it can be useful to visualize the location of these infections to detect spatial clustering. Here we discuss the development of a web application (GeoHAI) for infection preventionists (IPs) to visualize C.diff in rooms over time and quantify the burden of C.diff in these rooms.Figure 1:Screenshot of GeoHAI application (Development environment with simulated case data) This screenshot shows a view of a floor within one of the hospital buildings. Each circle in a room relates to a patient who was in that room on the chosen date. Clicking on a circle will give details on that patient (Patient ID, admit/discharge date, room number, date room was entered and exited, as well as whether the patient had C. diff (HO- or CO-CDI) and what was the date of the positive test). Orange circles indicate CO-CDI, red are HO-CDI. Shading corresponds to the cumulative number of hours patients with C. diff (HO- or CO-CDI) were in each room for the previous 30 days. The user can also look at the shading for 90 days or 365 days by choosing one of those options in the upper left-hand corner. The shading is intended to give a sense of overall burden of C. diff in each room. The list of floors on the left panel allows the user to switch between floors, and boxes in this area give a preview of high interest rooms on that floor. The building can be selected from the drop-down box in the upper right corner. Methods Development of GeoHAI consisted of (1) processing clinical and geographic data; (2) software development; and (3) deployment within the medical center (MC). (1) Floor plan data was converted to GeoJSON files which contain coordinates to draw polygons corresponding to rooms within the hospital. Rooms were given a spaceID, a unique number for each building-floor-room. Clinical and patient location data include patient ID, admission/discharge date, C.diff result, and entry/exit times for room transfers. We wrote R code to clean the data and identify hospital onset vs community onset C.diff (HO-CDI, CO-CDI), to link patient room to spaceID, and to calculate the historical burden of C.diff in each room, defined as the cumulative number of hours that patients with active C.diff spent in each room. (2) With IP input, we created design specifications which were used to develop the application. Node-RED was used to load clean, clinical data from CSV files into the application database. (3) IT worked with software developers to host the application on the MC’s cloud servers using Azure DevOps. Results GeoHAI is accessed by authorized users in development and production environments on the MC network. Users choose from 42 unique floors in 6 buildings, can select a patient within a room to see if they had C.diff during their stay, the date of C.diff, and room entry/exit date. The degree of room shading correlates to the cumulative hours patients with C.diff were in a room (Figure 1). New data can be loaded via an intuitive web interface. Conclusion GeoHAI displays C.diff cases in a large academic hospital. We next plan to study how GeoHAI impacts HAI investigation, to expand GeoHAI features, and to add additional HAIs. Disclosures All Authors: No reported disclosures
Processes to assure the safe, effective, and responsible deployment of artificial intelligence (AI) in safety-critical settings are urgently needed. Here we show a procedure to empirically evaluate the impacts of AI augmentation as a basis for responsible deployment. We evaluated three augmentative AI technologies nurses used to recognize imminent patient emergencies, including combinations of AI recommendations and explanations. The evaluation involved 450 nursing students and 12 licensed nurses assessing 10 historical patient cases. With each technology, nurses' performance was both improved and degraded when the AI algorithm was most correct and misleading, respectively. Our findings caution that AI capabilities alone do not guarantee a safe and effective joint human-AI system. We propose two minimum requirements for evaluating AI in safety-critical settings: (1) empirically measure the performance of people and AI together and (2) examine a range of challenging cases which produce a range of strong, mediocre, and poor AI performance.
Using ethnographic methods in community-based participatory research (CBPR) fosters mutual trust and collaboration between researchers and community members. However, scaling these methods is challenged by limited researcher resources and community member involvement. This paper introduces computational ethnography (CE), which uses computational assists to deliver rich insights. We argue that traditional ethnography, even with computational techniques, is insufficient for scaling effectively. Our Machine–Readable Co-design (MaRC) toolkit, a CE method, leverages human–machine teaming to address the collection and analysis of community stories at scale in CBPR. MaRC integrates AI/ML technologies to support effective community co-planning. In this study, two researchers used MaRC to capture and analyze over 100 community stories. While researchers spent over 120 hr across 2 months manually measuring and recording data, MaRC computed each story within 8 to 10 minutes, achieving an 84% concordance rate with our manual, baseline analysis. MaRC demonstrates significant potential for scaling future CBPR efforts.
A common pattern among intervention implementations is that, in many cases, initial adoption and early signals of success are followed by dwindling participation, signals that the intervention is not achieving its goals, and abandonment of the intervention as it was originally conceived. The IMPActS Framework facilitates anticipation of misalignment among stakeholders during the design, pitch, implementation, and sustainability of a candidate intervention’s implementation. We adapted the IMPActS Framework to design an IMPActS Workshop to occur within the context of an ongoing research with a partner organization to develop their proactive performance monitoring and sustained adaptability management program. The IMPActS Workshop prospectively tests the implementation sustainability of candidate interventions and sharpened intervention design to improve the likelihood of sustainment across the organization. We report the motivation, methods, results, and lessons learned from this effort.
Sterile Processing Departments (SPDs) must clean, maintain, store, and organize surgical instruments which are then delivered to Operating Rooms (ORs) using a Courier Network, with regular coordination occurring across departmental boundaries. To represent these relationships, we utilized the Systems Engineering Initiative for Patient Safety (SEIPS) 101 Toolkit, which helps model how health-related outcomes are affected by healthcare work systems. Through observations and interviews which built on prior work system analyses, we developed a SEIPS 101 journey map, PETT scan, and tasks matrices to represent the instrument reprocessing work system, revealing complex interdependencies between the people, tools, and tasks occurring within it. The SPD, OR and Courier teams are found to have overlapping responsibilities and a clear co-dependence, with critical implications for the successful functioning of the whole hospital system.
Explainable AI must simultaneously help people understand the world, the AI, and when the AI is misaligned to the world. We propose situated interpretation and data (SID) as a design technique to satisfy these requirements. We trained two machine learning algorithms, one transparent and one opaque, to predict future patient events that would require an emergency response team (ERT) mobilization. An SID display combined the outputs of the two algorithms with patient data and custom annotations to implicitly convey the alignment of the transparent algorithm to the underlying data. SID displays were shown to 30 nurses with 10 actual patient cases. Nurses reported their concern level (1–10) and intended response (1–4) for each patient. For all cases where the algorithms predicted no ERT (correctly or incorrectly), nurses correctly differentiated ERT from non-ERT in both concern and response. For all cases where the algorithms predicted an ERT, nurses differentiated ERT from non-ERT in response, but not concern. Results also suggest that nurses’ reported urgency was unduly influenced by misleading algorithm guidance in cases where the algorithm overpredicted and underpredicted the future ERT. However, nurses reported concern that was as or more appropriate than the predictions in 8 of 10 cases and differentiated ERT from non-ERT cases better than both algorithms, even the more accurate opaque algorithm, when the two predictions conflicted. Therefore, SID appears a promising design technique to reduce, but not eliminate, the negative impacts of misleading opaque and transparent algorithms.
Deployments of artificial intelligence (AI) and machine learning (ML) in healthcare can both help and harm patient outcomes, amplifying calls for a human-centered approach to AI/ML development. This paper details one approach guided by three principles: (1) pursue human-machine team (HMT) performance, not algorithm performance, (2) build interpretability throughout, and (3) constrain development to deconstrain interactions. We describe how these principles influenced our development of two algorithms predicting patient decompensation events five minutes into the future. These algorithms showed comparable performance to other similar models with enhanced interpretability that greatly expanded HMT interaction possibilities. Our early investments in the potential for teaming appeared to pay dividends for the resultant HMT.
Background: The Flexible Adaptive Algorithmic Surveillance Testing (FAAST) program represents an innovative approach for improving the detection of new cases of infectious disease; it is deployed here to screen and diagnose SARS-CoV-2. With the advent of treatment for COVID-19, finding individuals infected with SARS-CoV-2 is an urgent clinical and public health priority. While these kinds of Bayesian search algorithms are used widely in other settings (eg, to find downed aircraft, in submarine recovery, and to aid in oil exploration), this is the first time that Bayesian adaptive approaches have been used for active disease surveillance in the field.Objective: This study's objective was to evaluate a Bayesian search algorithm to target hotspots of SARS-CoV-2 transmission in the community with the goal of detecting the most cases over time across multiple locations in Columbus, Ohio, from August to October 2021.Methods: The algorithm used to direct pop-up SARS-CoV-2 testing for this project is based on Thompson sampling, in which the aim is to maximize the average number of new cases of SARS-CoV-2 diagnosed among a set of testing locations based on sampling from prior probability distributions for each testing site. An academic-governmental partnership between Yale University, The Ohio State University, Wake Forest University, the Ohio Department of Health, the Ohio National Guard, and the Columbus Metropolitan Libraries conducted a study of bandit algorithms to maximize the detection of new cases of SARS-CoV-2 in this Ohio city in 2021. The initiative established pop-up COVID-19 testing sites at 13 Columbus locations, including library branches, recreational and community centers, movie theaters, homeless shelters, family services centers, and community event sites. Our team conducted between 0 and 56 tests at the 16 testing events, with an overall average of 25.3 tests conducted per event and a moving average that increased over time. Small incentives-including gift cards and take-home rapid antigen tests-were offered to those who approached the pop-up sites to encourage their participation.Results: Over time, as expected, the Bayesian search algorithm directed testing efforts to locations with higher yields of new diagnoses. Surprisingly, the use of the algorithm also maximized the identification of cases among minority residents of underserved communities, particularly African Americans, with the pool of participants overrepresenting these people relative to the demographic profile of the local zip code in which testing sites were located.Conclusions: This study demonstrated that a pop-up testing strategy using a bandit algorithm can be feasibly deployed in an urban setting during a pandemic. It is the first real-world use of these kinds of algorithms for disease surveillance and represents a key step in evaluating the effectiveness of their use in maximizing the detection of undiagnosed cases of SARS-CoV-2 and other infections, such as HIV.
OBJECTIVE:Observe patient-clinician communication to gain insight about the reasons underlying the choice of patients with unilateral breast cancer to undergo contralateral prophylactic mastectomy (CPM), despite lack of survival benefit, risk of harms, and cautions expressed by surgical guidelines and clinicians. METHODS & MEASURES:WORDS is a prospective study that explored patient-clinician communication and patient decision making. Participants recorded clinical visits through a downloadable mobile application. We analyzed 44 recordings from 22 patients: 9 who chose CPM, 8 who considered CPM but decided against it, and 5 who never considered CPM. We used abductive analysis combined with constructivist grounded theory methods. RESULTS:Decisions to undergo CPM are patient-driven and motivated by perceptions that CPM is the most aggressive, and therefore safest, treatment option available. These decisions are shaped not primarily by the content of conversations with clinicians, but by the history of cancer in patients' families, their own first-hand experiences with cancers among loved ones, fear for their children, and anxiety about cancer recurrence. CONCLUSION:The perception that CPM is the safest, most aggressive option strongly influences patients, despite scientific evidence to the contrary. Future efforts to address high CPM rates should focus on patient-driven decision making and cancer-related fears.
Objective Reduce nurse response time for emergency and high-priority alarms by increasing discriminability between emergency and all other alarms and suppressing redundant and likely false high-priority alarms in a secondary alarm notification system (SANS). Background Emergency alarms are the most urgent, requiring immediate action to address a dangerous situation. They are clinician-triggered and have higher positive predictive value (PPV). High-priority alarms are automatically triggered and have lower PPV. Method We performed a retrospective pre-post study, analyzing data 15 months before and 25 months after a SANS redesign was implemented in four hospitals. For emergency alarms, we incorporated digitized human speech to distinguish them from automatically triggered alarms, leaving their onset and escalation pathways unchanged. For automatically triggered alarms, we suppressed some by delaying initial onset and escalation by 20 s. We used linear mixed models to assess the change in response time, Fisher's exact test for the proportion of response times longer than 120 s, and control charts for process stability. Results Response time for emergency alarms decreased at all hospitals (main, from 26.91 s to 22.32 s, p < .001; cardiac, from 127.10 s to 52.43 s, p < .001; cancer, from 18.03 s to 15.39 s, p < .001). Improvements were sustained. Automatically triggered alarms decreased 25.0%. Response time for the three automatically triggered cardiac alarms increased at the four hospitals. Conclusion Auditory sound disambiguation was associated with a sustained reduced nurse response time for emergency alarms, but suppressing some high-priority automatically triggered alarms was not. Application Distinguishing and escalating urgent, actionable alarms with higher PPV improves response time.
In this panel, we present perspectives on how to incorporate fundamental concepts in resilience engineering into health care human-in-the-loop simulations. Our panelists have successfully implemented concepts, but also continue to experience challenges with convincing colleagues of the importance and value of doing so for simulations that have other training, technology evaluation, or quality improvement objectives.