OBJECTIVE:Accurately characterizing research trends is critical for identifying cutting-edge scientific breakthroughs in their infancy and informing strategic priorities. This research contributes a pipeline that utilizes generative AI technologies to develop research topic taxonomies from publication keywords and analyze keyword evolution within topics, methodological and domain trends, and topic co-occurrences. We demonstrated the pipeline by conducting a retrospective analysis of biomedical informatics research trends in the Journal of Biomedical Informatics (JBI). METHODS:We identified the JBI publications with keywords available on PubMed, spanning 2011-2025. We downloaded all the keywords and categorized them into methodological innovations and health domains, identified topics, assigned topic names, and constructed their hierarchies, all using large-language models (LLMs). We introduced an automated method for evaluating topics, leveraging MeSH terminology as the underlying knowledge base. RESULTS:Using 6,930 unique keywords from 2,427 publications, we derived 1,028 distinct topics related to methodological innovations, with each topic associated with medians of four keywords (Q1: 2, Q3: 13) and six publications (Q1: 2, Q3: 19). We identified 904 topics related to health domains, with each topic associated with three keywords (Q1: 1, Q3: 11) and four publications (Q1: 1, Q3: 15). Based on the topics, we analyzed the prominent research areas, trends in publication volume, evolution of keyword distributions within each topic, and patterns of co-occurring topics. Among the 2,379 eligible publications, 2,009 (84.4%) exhibited overlap between the keyword-derived MeSH terms and the MeSH terms assigned to the publication by the National Library of Medicine. CONCLUSION:This study presents a method that leverages modern generative AI technologies for retrospective analysis of a scientific field to identify emerging topics and to detect shifts in scholarly focus. Illustrated by data for JBI and correlated with historical background events and policy changes, our findings demonstrate the effectiveness and utility of the methods while providing a powerful lens to understand the evolution of biomedical informatics research priorities in JBI.
Artificial intelligence (AI) is increasingly embedded in clinical environments, raising questions of trust, fairness, empathy, and governance. The ethical terrain surrounding AI in medicine remains unstable despite its rapid adoption. We introduce the “Seven Deadly Sins of AI in Medicine”, a conceptual framework of recurring systemic failure modes: (i) Blind Trust, (ii) Overregulation, (iii) Dehumanization, (iv) Misaligned Optimization, (v) Overinforming and False Forecasting, (vi) Misapplied Statistics, and (vii) Self-Referential Evaluation. The framework was developed through systematic synthesis of scientific literature, clinical guidelines, and regulatory frameworks prior to any empirical data collection. To validate this pre-established framework, we conducted a global, cross-professional opinion poll of 914 stakeholders from 143 countries between July 2024 and March 2025. Results confirmed broad agreement with each pre-identified risk, revealing cross-cultural convergence in ethical concern alongside persistent divides in attitudes toward regulation—particularly between technologically advanced nations and emerging economies. We further propose an inversion of the framework into seven cardinal virtues for AI in medicine, offering actionable principles to guide responsible development and governance. The goal is to move beyond scattered ethical guidelines toward a unified diagnostic tool for trustworthy, human-centered medical AI.
Metrics and instruments can provide guidance for clinical researchers to assess their potential research projects at an early stage before significant investment. Furthermore, metrics can also provide structured criteria for peer reviewers to assess others’ clinical research manuscripts or grant proposals. This study aimed to develop, test, validate, and use evaluation metrics and instruments to accurately, consistently, systematically, and conveniently assess the quality of scientific hypotheses for clinical research projects. Metrics development went through iterative stages, including literature review, metrics and instrument development, internal and external testing and validation, and continuous revisions in each stage based on feedback. Furthermore, two experiments were conducted to determine brief and comprehensive versions of the instrument. The brief version of the instrument contained three dimensions: validity, significance, and feasibility. The comprehensive version of metrics included novelty, clinical relevance, potential benefits and risks, ethicality, testability, clarity, interestingness, and the three dimensions of the brief version. Each evaluation dimension included 2 to 5 subitems to evaluate the specific aspects of each dimension. For example, validity included clinical validity and scientific validity. The brief and comprehensive versions of the instruments included 12 and 39 subitems, respectively. Each subitem used a 5-point Likert scale. The validated brief and comprehensive versions of metrics can provide standardized, consistent, systematic, and generic measurements for clinical research hypotheses, allow clinical researchers to prioritize their research ideas systematically, objectively, and consistently, and can be used as a tool for quality assessment during the peer review process.
We conducted a data-driven hypothesis generation study with clinical researchers using VIADS (a visual interactive analysis tool for filtering and summarizing large data sets coded with hierarchical terminologies) or other analytical tools (as control, e.g., SPSS, SAS, R). The participants analyzed the same datasets and developed hypotheses using a think-aloud verbal protocol. Their screen activities and audio were recorded, transcribed, and coded for cognitive events. We analyzed the recordings to identify the cognitive events (e.g., "Analyze data") during hypothesis generation. The VIADS group exhibited the lowest mean number of cognitive events per hypothesis with the smallest standard deviation. The highest percentages of cognitive events in hypothesis generation were "Using analysis results" (30%) and "Seeking connections" (23%). The results suggest that VIADS may guide participants better than the control group. Our framework for scientific hypothesis generation in clinical research contexts guides the elaboration of the underlying cognitive mechanism of the process.
As evident from the discussions throughout this book, workflow plays a central role in ensuring smooth functioning of all clinical activities—from patient encounter to medication administration to population health management. Any disruption to workflow can result in severe, adverse consequences such as decreased time efficiency and greater patient safety risks. In the recent two decades, the most systemic disruption to clinical workflow across the globe is associated with the widespread implementation of health IT systems, electronic health records (EHR) in particular.
Objectives: To obtain insights about inexperienced clinical researchers’ hypothesis quality and associated factors. The findings inform the development of informatics tools to aid the hypothesis generation process. Methods: We analyzed an existing dataset collected through a randomized controlled study, focusing on individual hypotheses and participants. We invited clinical researchers to analyze datasets and develop hypotheses using the think-aloud method. Participants’ screen activity and audio were recorded, transcribed, coded, and analyzed to measure the time and cognitive events (a granular unit of thought processes used by the participants while generating hypotheses). Hypotheses were rated by an expert panel. Here we analyzed (1) the top 5-rated hypotheses, (2) the bottom 5-rated hypotheses, and (3) the participants who generated them. Results: Participants who generated the top 5-rated hypotheses utilized fewer cognitive events and a shorter range of time per hypothesis; their hypotheses presented a higher valid rate, and they were more experienced. Conclusion: Having more experience is positively associated with higher quality and valid rates of the generated hypotheses. The higher-rated hypotheses seem to be positively associated with slightly fewer cognitive events and shorter time. The effect may not be linear. These analyses provide evidence for customized study designs or tool development based on these associated factors.
The ability to track and detect the activities and process that constitute clinical workflow for performance analysis and error detection has been enhanced with the inclusion of modern technological interventions in clinical environments. One such important intervention is automated location tracking which is a system that detects the movement of clinically relevant entities (physicians, nurses, patients, and equipment). In this chapter, we elucidate the technologies associated with automated location tracking focusing on the two most widely used: Radio-Frequency Identification (RFID) and Bluetooth. We describe specific systems of each type to give readers a general model of the technological requirements for similar setups in clinical environments. Our goal in writing this chapter is the provide readers with an overview of the state-of-the-art technologies and analytic methods. This can hopefully serve as a guide for similar setups in other medical organizations and clinical sites. The process of using a location tracking system to perform novel workflow data analytics is achieved using computational methods whose efficacy has been enhanced by the continuous collection of tracking data that can be potentially collected. Case studies from our own research in the emergency department at the Mayo Clinic, are used as illustration. Finally, we use visualization techniques that can be used to convey workflow related information as well as a proof-of-concept visualization dashboard, which can be used to provide continuous and consistent feedback to clinical target users. This will facilitate self-assessment of workflow and related behaviors and potentially detect bottlenecks and sources of error.
Objectives:To compare how clinical researchers generate data-driven hypotheses with a visual interactive analytic tool (VIADS, a visual interactive analysis tool for filtering and summarizing large datasets coded with hierarchical terminologies) or other tools.Methods:We recruited clinical researchers and separated them into "experienced" and "inexperienced" groups. Participants were randomly assigned to a VIADS or control group within the groups. Each participant conducted a remote 2-hour study session for hypothesis generation with the same study facilitator on the same datasets by following a think-aloud protocol. Screen activities and audio were recorded, transcribed, coded, and analyzed. Hypotheses were evaluated by seven experts on their validity, significance, and feasibility. We conducted multilevel random effect modeling for statistical tests.Results:Eighteen participants generated 227 hypotheses, of which 147 (65%) were valid. The VIADS and control groups generated a similar number of hypotheses. The VIADS group took a significantly shorter time to generate one hypothesis (e.g., among inexperienced clinical researchers, 258 s versus 379 s, p = 0.046, power = 0.437, ICC = 0.15). The VIADS group received significantly lower ratings than the control group on feasibility and the combination rating of validity, significance, and feasibility.Conclusion:The role of VIADS in hypothesis generation seems inconclusive. The VIADS group took a significantly shorter time to generate each hypothesis. However, the combined validity, significance, and feasibility ratings of their hypotheses were significantly lower. Further characterization of hypotheses, including specifics on how they might be improved, could guide future tool development.
Health information technologies have become vital tools for the practice of clinical medicine. However, numerous challenges remain for the fuller realization of its potential as instruments that advance clinical care and enhance patient safety. Human-computer interaction (HCI) is a discipline rooted in computer science as well as the social and behavioral sciences. It is focally concerned with evaluating and improving user experience, usability, and usefulness of technology. HCI in medicine and healthcare, the subject matter of this volume, extends across clinical and consumer health informatics, addressing a range of user populations including providers, biomedical scientists and patients. The breadth of HCI in biomedicine and healthcare is rather broad including thousands of journal articles across medical disciplines and consumer health domains. Although the 14 chapters in this volume are rather varied in subject matter and scope, there is greater focus on clinical informatics with a couple of chapters addressing consumer health informatics issues. This introductory chapter provides a brief overview of the other chapters in this volume.
Objectives:We invited inexperienced clinical researchers to analyze coded health datasets and develop hypotheses. We recorded and analyzed their hypothesis generation process. All the hypotheses generated in the process were rated by the same group of seven experts by using the same metrics. This case study examines the higher quality (i.e., higher ratings) and lower quality of hypotheses and participants who generated them. We characterized the contextual factors associated with the quality of hypotheses. Methods:All participants (i.e., clinical researchers) completed a 2-hour study session to analyze data and generate scientific hypotheses using the think-aloud method. Participants' screen activity and audio were recorded and transcribed. These transcriptions were used to measure the time used to generate each hypothesis and to code cognitive events (i.e., cognitive activities used when generating hypotheses, for example, "Seeking for Connection" describes an attempt to draw connections between data points). The hypothesis ratings by the expert panel were used as the quality of the hypotheses during the analysis. We analyzed the factors associated with (1) the five highest and (2) five lowest rated hypotheses and (3) the participants who generated them, including the number of hypotheses per participant, the validity of those hypotheses, the number of cognitive events used for each hypothesis, as well as the participant's research experience and basic demographics. Results:Participants who generated the five highest-rated hypotheses used similar lengths of time (difference 3:03), whereas those who generated the five lowest-rated hypotheses used more varying lengths of time (difference 7:13). Participants who generated the five highest-rated hypotheses also utilized slightly fewer cognitive events on average compared to the five lowest-rated hypotheses (4 per hypothesis vs. 4.8 per hypothesis). When we examine the participants (who generated the five highest and five lowest hypotheses) and their total hypotheses generated during the 2-hour study sessions, the participants with the five highest-rated hypotheses again had a shorter range of time per hypothesis on average (0:03:34 vs. 0:07:17). They (with the five highest ratings) used fewer cognitive events per hypothesis (3.498 vs. 4.626). They (with the five highest ratings) also had a higher percentage of valid rate (75.51% vs. 63.63%) and generally had more experience with clinical research. Conclusion:The quality of the hypotheses was shown to be associated with the time taken to generate them, where too long or too short time to generate hypotheses appears to be negatively associated with the hypotheses' quality ratings. Also, having more experience seems to positively correlate with higher ratings of hypotheses and higher valid rates. Validity is a quality dimension used by the expert panel during rating. However, we acknowledge that our results are anecdotal. The effect may not be simply linear, and future research is necessary. These results underscore the multi-factor nature of hypothesis generation.
Hypothesis generation is an early and critical step in any hypothesis-driven clinical research project. Because it is not yet a well-understood cognitive process, the need to improve the process goes unrecognized. Without an impactful hypothesis, the significance of any research project can be questionable, regardless of the rigor or diligence applied in other steps of the study, e.g., study design, data collection, and result analysis. In this perspective article, the authors provide a literature review on the following topics first: scientific thinking, reasoning, medical reasoning, literature-based discovery, and a field study to explore scientific thinking and discovery. Over the years, scientific thinking has shown excellent progress in cognitive science and its applied areas: education, medicine, and biomedical research. However, a review of the literature reveals the lack of original studies on hypothesis generation in clinical research. The authors then summarize their first human participant study exploring data-driven hypothesis generation by clinical researchers in a simulated setting. The results indicate that a secondary data analytical tool, VIADS—a visual interactive analytic tool for filtering, summarizing, and visualizing large health data sets coded with hierarchical terminologies, can shorten the time participants need, on average, to generate a hypothesis and also requires fewer cognitive events to generate each hypothesis. As a counterpoint, this exploration also indicates that the quality ratings of the hypotheses thus generated carry significantly lower ratings for feasibility when applying VIADS. Despite its small scale, the study confirmed the feasibility of conducting a human participant study directly to explore the hypothesis generation process in clinical research. This study provides supporting evidence to conduct a larger-scale study with a specifically designed tool to facilitate the hypothesis-generation process among inexperienced clinical researchers. A larger study could provide generalizable evidence, which in turn can potentially improve clinical research productivity and overall clinical research enterprise.
Background:The US Department of Veterans Affairs (VA) launched the VA Video Connect (VVC) video conferencing platform to connect veterans with VA clinicians in 2018. We assessed practices, concerns, and perceptions toward VVC encounters among physicians within the VA New Mexico Healthcare System (VANMHCS). Methods:Medicine Service Physicians of VANMHCS who had previously completed ≥ 1 VVC encounter were invited to semistructured interviews. Questions were constructed to assess the following domains: overarching views of video telehealth, perceptions of the VVC application, and barriers to the broad implementation of video telehealth. Interviews were assessed using a qualitative, open-coding approach. Themes were constructed both deductively, through direct responses to interview questions, and inductively, by identifying emerging patterns in the data. Results:Of the 64 physicians invited to participate, 13 (20%) were interviewed. Of those interviewed, 9 (69%) were female, 10 (77%) were specialists, 8 (62%) had been practicing for ≥ 10 years, and 7 (54%) completed ≥ 5 VVC visits. Interviews ranged from 10 to 25 minutes. Five themes were observed: (1) VVC software and internet connection issues affected implementation; (2) patient technological literacy affected both veteran and physician comfort with VVC; (3) integration of supportive measures is desired; (4) clinical video telehealth (CVT) services may increasingly enhance access to care; and (5) in-person encounters provided unique advantages over CVT. Conclusions:Physicians believe VVC could lead to improved access to care for veterans facing geographical challenges. Efforts should focus on improving VVC user interface and addressing technological issues, educating veterans/physicians on the use of CVT, and integrating supportive measures for successful VVC encounters.
BACKGROUND:We developed a prototype patient decision aid, EyeChoose, to assist college-aged students in selecting a refractive surgery. EyeChoose can educate patients on refractive errors and surgeries, generate evidence-based recommendations based on a user's medical history and personal preferences, and refer patients to local refractive surgeons.OBJECTIVES:We conducted an evaluative study on EyeChoose to assess the alignment of surgical modality recommendations with a user's medical history and personal preferences, and to examine the tool's usefulness and usability.METHODS:We designed a mixed methods study on EyeChoose through simulations of test cases to provide a quantitative measure of the customized recommendations, an online survey to evaluate the usefulness and usability, and a focus group interview to obtain an in-depth understanding of user experience and feedback.RESULTS:We used stratified random sampling to generate 245 test cases. Simulated execution indicated EyeChoose's recommendations aligned with the reference standard in 243 (99%). A survey of 55 participants with 16 questions on usefulness, usability, and general impression showed that 14 questions recorded more than 80% positive responses. A follow-up focus group with 10 participants confirmed EyeChoose's useful features of patient education, decision assistance, surgeon referral, as well as good usability with multimedia resources, visual comparison among the surgical modalities, and the overall aesthetically pleasing design. Potential areas for improvement included offering nuances in soliciting user preferences, providing additional details on pricing, effectiveness, and reversibility of surgeries, expanding the function of surgeon referral, and fixing specific usability issues.CONCLUSION:The initial evaluation of EyeChoose suggests that it could provide effective patient education, generate appropriate recommendations, connect to local refractive surgeons, and demonstrate good system usability in a test environment. Future research is required to enhance the system functions, fully implement and evaluate the tool in naturalistic settings, and examine the findings' generalizability to other populations.
Clinical cognition is central to a clinician's daily tasks, such as making diagnostic and therapeutic decisions. For example, doctors rely on their memory to recall relevant facts, concepts and experiences that can help them diagnose and treat their patients. Memory is needed for clinicians to
BACKGROUND:Ward rounds offer a rich environment for learning about team clinical reasoning. We aimed to assess how team clinical reasoning occurs on ward rounds to inform efforts to enhance the teaching of clinical reasoning.METHODS:We performed focused ethnography of ward rounds over a 6-week period, during which we observed five different teams. Each day team comprised one senior physician, one senior resident, one junior resident, two interns and one medical student. Twelve 'night-float' residents who discussed new patients with the day team were also included. Field notes were analysed using content analysis.FINDINGS:We analysed 41 new patient presentations and discussions on 23 different ward rounds. The median duration of case presentations and discussions was 13.0 minutes (IQR, 10.0-18.0 minutes). More time was devoted to information sharing (median 5.5 minutes; IQR, 4.0-7.0 minutes) than any other activity, followed by discussion of management plans (median 4.0 minutes; IQR, 3.0-7.8 minutes). Nineteen (46%) cases did not include discussion of a differential diagnosis for the chief concern. We identified two themes relevant to learning: (1) linear versus iterative approaches to team-based diagnosis and (2) the influence of hierarchy on participation in clinical reasoning discussions.CONCLUSION:The ward teams we observed spent far less time discussing differential diagnoses compared with information sharing. Junior learners such as medical students and interns contributed less frequently to team clinical reasoning discussions. In order to maximise student learning, strategies to engage junior learners in team clinical reasoning discussions on ward rounds may be needed.
Objectives:This study aims to identify the cognitive events related to information use (e.g., "Analyze data", "Seek connection") during hypothesis generation among clinical researchers. Specifically, we describe hypothesis generation using cognitive event counts and compare them between groups. Methods:The participants used the same datasets, followed the same scripts, used VIADS (a visual interactive analysis tool for filtering and summarizing large data sets coded with hierarchical terminologies) or other analytical tools (as control) to analyze the datasets, and came up with hypotheses while following the think-aloud protocol. Their screen activities and audio were recorded and then transcribed and coded for cognitive events. Results:The VIADS group exhibited the lowest mean number of cognitive events per hypothesis and the smallest standard deviation. The experienced clinical researchers had approximately 10% more valid hypotheses than the inexperienced group. The VIADS users among the inexperienced clinical researchers exhibit a similar trend as the experienced clinical researchers in terms of the number of cognitive events and their respective percentages out of all the cognitive events. The highest percentages of cognitive events in hypothesis generation were "Using analysis results" (30%) and "Seeking connections" (23%). Conclusion:VIADS helped inexperienced clinical researchers use fewer cognitive events to generate hypotheses than the control group. This suggests that VIADS may guide participants to be more structured during hypothesis generation compared with the control group. The results provide evidence to explain the shorter average time needed by the VIADS group in generating each hypothesis.
BACKGROUND:Visualization can be a powerful tool to comprehend data sets, especially when they can be represented via hierarchical structures. Enhanced comprehension can facilitate the development of scientific hypotheses. However, the inclusion of excessive data can make visualizations overwhelming. OBJECTIVE:We developed a visual interactive analytic tool for filtering and summarizing large health data sets coded with hierarchical terminologies (VIADS). In this study, we evaluated the usability of VIADS for visualizing data sets of patient diagnoses and procedures coded in the International Classification of Diseases, Ninth Revision, Clinical Modification (ICD-9-CM). METHODS:We used mixed methods in the study. A group of 12 clinical researchers participated in the generation of data-driven hypotheses using the same data sets and time frame (a 1-hour training session and a 2-hour study session) utilizing VIADS via the think-aloud protocol. The audio and screen activities were recorded remotely. A modified version of the System Usability Scale (SUS) survey and a brief survey with open-ended questions were administered after the study to assess the usability of VIADS and verify their intense usage experience with VIADS. RESULTS:The range of SUS scores was 37.5 to 87.5. The mean SUS score for VIADS was 71.88 (out of a possible 100, SD 14.62), and the median SUS was 75. The participants unanimously agreed that VIADS offers new perspectives on data sets (12/12, 100%), while 75% (8/12) agreed that VIADS facilitates understanding, presentation, and interpretation of underlying data sets. The comments on the utility of VIADS were positive and aligned well with the design objectives of VIADS. The answers to the open-ended questions in the modified SUS provided specific suggestions regarding potential improvements for VIADS, and the identified problems with usability were used to update the tool. CONCLUSIONS:This usability study demonstrates that VIADS is a usable tool for analyzing secondary data sets with good average usability, good SUS score, and favorable utility. Currently, VIADS accepts data sets with hierarchical codes and their corresponding frequencies. Consequently, only specific types of use cases are supported by the analytical results. Participants agreed, however, that VIADS provides new perspectives on data sets and is relatively easy to use. The VIADS functionalities most appreciated by participants were the ability to filter, summarize, compare, and visualize data. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID):RR2-10.2196/39414.