
Background:In the digital era, smartphone technology has significantly advanced ocular imaging, allowing for high-resolution anterior segment photography (ASP) to enhance referrals, decision-making, and disease monitoring within clinical workflows. Objective:This exploratory quality-improvement project aims to examine clinician experiences with 4 smartphone-based and slit-lamp-based ASP workflows used in eye casualty and corneal services. Methods:A mixed methods project was conducted at a tertiary eye hospital in England (United Kingdom) between October 21, 2024, and December 6, 2024. The full imaging setups of 4 ASP workflows were evaluated within an established clinical setting: the Magnifier app (a smartphone-based app; Apple Inc), QuikVue smartphone adapter (a smartphone-based camera lens attachment; VisuScience Meditech Co Ltd), NexYZ Slit-Lamp Adapter (a slit-lamp-based attachment; Celestron LLC), and Zeiss Slit Lamp Imaging Solution (ZSLIS; an integrated slit-lamp camera system; Carl Zeiss AG). Ophthalmology residents and corneal fellows completed nonvalidated usability surveys and participated in one-on-one interviews examining the role of ASP workflows in clinical practice. Survey data were analyzed using descriptive statistics. Interviews were transcribed and subjected to reflexive thematic analysis. Results:Among 25 survey respondents (17 residents and 8 fellows) and 12 interviewees, clinician-rated assessments of usability, cost-effectiveness, and clinical utility favored the Magnifier workflow, which received the highest ratings across these domains (19/25, 76%; 23/25, 92%; and 18/25, 72% strongly agreed, respectively). Five themes emerged from qualitative analysis: obstacles to clinical implementation, accessibility, quality and workflow, digital ethics, and innovation. Conclusions:No single ASP workflow proved superior in usability ratings and interview themes regarding clinical application. However, the use of the Magnifier workflow appeared to offer the most favorable balance of usability, cost-effectiveness, and clinical utility in clinician-rated assessments, though this reflects the full imaging workflow rather than the app alone. These findings are exploratory and require future work to examine workflow integration, data governance for smartphone imaging, and patient perspectives on ASP workflows.
Background:Commercially available wearable devices, capable of measuring cardiac rhythm, are gaining popularity for convenient heart health monitoring. Accurate measurement of key intervals, specifically RR intervals for rhythm and QT and corrected QT (QTc) intervals for arrhythmia risk, is crucial for assessing potential cardiac morbidity. Objective:This study aimed to evaluate the agreement of RR, QT, and QTc intervals measured by a consumer wearable device (Withings ScanWatch series 1) with those measured using a research-grade 3-lead electrocardiogram (ECG; Biopac) and a clinically approved single-lead patch (Cardea SOLO). Methods:This was a cross-sectional study of 28 healthy adults aged 21 to 65 years, with a mean age of 39.4 (SD 13.0) years. Participants underwent 4 concurrent ECG measurements per device: twice seated at rest, once supine, and once after exercise. All RR and QT intervals were manually measured using standardized caliper software (EP calipers) by trained reviewers, and QTc was calculated using the Bazett formula. Agreement was assessed via repeated-measures correlation (rrm) and linear mixed-effects models (LMMs) for repeated-measures Bland-Altman analysis. Results:Strong agreement was observed for RR intervals when comparing the Withings ScanWatch with both the Biopac (rrm=0.84; P<.001) and the SOLO (rrm=0.80; P<.001). Agreement was weaker for QT intervals (rrm=0.37; P=.001 compared to Biopac and rrm=0.37; P=.001 compared to SOLO) and QTc intervals (rrm=0.27; P=.01 compared to Biopac and rrm=0.27; P=.03 compared to SOLO). Bland-Altman LMM analysis demonstrated that the Withings ScanWatch systematically underestimated QT intervals, with a mean error (ME) of -42.9 (SD 30.1) milliseconds versus Biopac and -15.8 (SD 20.6) milliseconds versus SOLO. Similarly, QTc intervals were underestimated, with an ME of -51.4 (SD 32.3) milliseconds versus Biopac and -18.8 (SD 21.5) milliseconds versus SOLO. Conclusions:While the Withings ScanWatch showed strong agreement for RR interval measurement, its variability and systematic underestimation in QT and QTc measurements limited its clinical applicability for accurate assessment of arrhythmia risk. The device may serve as a valuable tool for tracking trends over time, but it is currently unreliable for precise interval screening.
Background:Remote patient management (RPM) is a strategy to track daily health data in patients with heart failure (HF) to enable early detection and treatment of decompensation to prevent hospital readmissions. The implementation of RPM alters standard care and redistributes provider responsibilities. Although randomized controlled trials demonstrate positive clinical outcomes, the impact of implementing RPM on patients' health care usage, as well as the care time and workload of health care providers in a real-world setting, remains unclear. Objective:The objective of this retrospective study is to explore the effects of implementing RPM for HF on patients' cardiac care consumption and caregiver workload within a regional care network through the use of real-world data. Methods:Two hospitals in the same regional care network with aligned care pathways for usual care (UC) were compared: one with RPM (hospital 1) and one without (hospital 2). Between June 2021 and July 2022, patients with HF were included from a clinical registry with a 1-year follow-up. Cardiac health care consumption was quantified by the number and duration of consultations (in-person, phone, and RPM) with nurse practitioners specialized in HF care and cardiologists. HF-related readmissions were also collected. Three groups were analyzed: hospital 1 RPM, hospital 1 UC, and hospital 2 UC. Results:A total of 131 patients were included (hospital 1 RPM: 34, hospital 1 UC: 64, hospital 2 UC: 33). Consultation frequency and duration were significantly higher in the hospital 1 RPM cohort compared to the hospital 1 UC and the hospital 2 UC cohorts (frequency: mean 18.5, SD 9.3 vs mean 9.3, SD 6.0 and mean 11.6, SD 6.3; P<.001; duration: mean 233.2, SD 95.2 min vs mean 121.2, SD 47.9 min and mean 113.0, SD 66.0 min; P<.001), mainly due to increased number of consultations of nurse practitioners specialized in HF care (hospital 1 RPM: mean 15.3, SD 8.4 vs hospital 1 UC: mean 5.3, SD 2.4 and hospital 2 UC: mean 8.6, SD 6.1; P<.001). No significant differences were observed regarding cardiologist care time (P=.27), cardiologist number of consultations (P=.43), or HF-related readmissions (P=.47). Conclusions:In this real-world study, the implementation of RPM for patients with HF was associated with a higher nursing workload, with no observed differences in cardiologists' consultations or patient outcomes. These findings highlight the need for optimization at organizational and technological levels and underline the importance of real-world evaluation of the implementation of innovations.
Abstract Background Basic life support (BLS) significantly improves survival and neurological outcomes after out-of-hospital cardiac arrest (OHCA). However, bystander cardiopulmonary resuscitation (CPR) rates vary widely worldwide, reaching only 40% in Geneva, Switzerland. The European Resuscitation Council’s “Kids Save Lives” statement advocates for integrating BLS education into mandatory school curricula to improve bystander CPR rates, and thus OHCA outcomes. Training medical students as BLS instructors could help address the shortage of qualified instructors needed to implement school-based BLS programs in primary schools. Objective The primary aim of this pilot study was to assess the feasibility of involving medical students as BLS instructors for primary school classes and to evaluate the impact of their teaching on primary school students’ immediate knowledge acquisition and self-confidence. Methods This pilot implementation study was conducted in Geneva, Switzerland. Second- and third-year medical students were presented the project during their emergency skills training and invited to complete an online questionnaire. Sixteen students were selected and underwent instructor training, enabling them to deliver a BLS course in a primary school class. Participating seventh-grade primary school students completed a questionnaire designed to assess immediate knowledge acquisition and confidence. The primary outcome was the feasibility of involving medical students as BLS instructors, including the proportion of medical students willing to become BLS instructors, the number trained and certified, and the number of BLS courses delivered. Secondary outcomes included children’s mean knowledge score and confidence in calling the emergency phone number. Predictors of willingness, and of knowledge and confidence, were also examined. Results Among 325 eligible students, 107 (32.9%) completed the questionnaire, and 87 (26.8% of all eligible students; 81.3% of respondents) expressed willingness to fully participate in this program. The 16 selected students successfully completed the instructor training and delivered courses in 16 primary school classes. No significant predictors of willingness were identified. A total of 295 (81.7%) of 361 children completed the postcourse questionnaires. Knowledge acquisition was high, with a mean score of 5.4 (SD 1.1) out of 7. High proportions of correct responses were observed for most BLS concepts, except for breathing assessment and recovery position indications. Overall, 78% (230/295) of children reported confidence in calling emergency services. No significant predictors of knowledge or confidence were found. Conclusions Medical students can effectively deliver BLS training to primary school students and may represent a useful workforce for school-based BLS programs. However, the low response rate introduces a selection bias, as many students did not respond and the observed willingness may not reflect that of the overall student population. Sustainable scale-up will require involving other instructors. Further studies should assess the feasibility of large-scale implementation.
Background:Generative AI can reduce the academic-writing burden on clinical health care professionals, but unsupervised use introduces citation hallucination (the confident fabrication or misattribution of references), which threatens research integrity. When a machine invents a source, it is termed "hallucination," and when a person does it, it is termed "fabrication," yet both are equally unacceptable. Existing health professions education writing workshops have rarely translated this concern into a concrete, reproducible source-verification procedure. Objective:This study aimed to describe the development and delivery of a 2-day workshop teaching the Atomic Sentence method, a source-anchored knowledge-modeling technique for AI-assisted literature synthesis, and evaluate its feasibility, acceptability, and short-term effect on research-idea development among clinical health care professionals. Methods:We conducted a single-cohort educational program evaluation of a 2-day workshop (≈16 contact hours) for 18 health care professionals across 4 campuses of a Buddhist medical network in Taiwan. Curriculum development used the ADDIE (Analysis, Design, Development, Implementation, and Evaluation) model; outcomes were framed with the Kirkpatrick model (levels 1-2). The workshop implemented a published 7-step AI-assisted research workflow (topic exploration, literature search, knowledge-base management, reading, synthesis, writing, and peer-review simulation). Citation integrity was protected by confining citation generation to a source-grounded tool (NotebookLM) combined with atomic-sentence extraction and mandatory cross-reference verification; other, nongrounded tools supported discovery, reading, and writing. Feasibility was assessed by the workshop completion rate and the questionnaire response rate. Outcomes were an acceptability questionnaire covering 5 domains (usefulness of instructional materials, curriculum planning, time allocation, personal-goal attainment, and administrative support; each rated 0-10) and a pre/post research-topic-transformation analysis (4 categories; 2 independent coders, Cohen κ); responses to an open-ended item on improvement suggestions were analyzed by qualitative content analysis. Reporting follows the GAMER (Guidance for AI Use in Medical Education Reporting) guidance for AI use in medical education. Results:All 18 (100%) participants completed the workshop and evaluation. Acceptability was high across all 5 domains (each median 10; overall median 10, IQR 9.3-10). The most variable domain was time allocation (range 5-10). Short-term improvement in the research topic (substantial transformation or refinement) occurred in 11 (61%; κ=0.77) participants; 2 (11%) had no structured question before or after. All 18 (100%) would recommend the workshop. Conclusions:A short workshop teaching the Atomic Sentence method was feasible and well accepted and was associated with short-term research-idea development in a multidisciplinary clinical cohort. Whether the method reduces citation hallucination or improves citation accuracy was not tested and requires controlled studies with objective outcomes, blinded raters, and longitudinal follow-up.
Background:The SimZones model is an organizational framework for simulation-based education that structures learning across 5 progressive zones, from self-directed preparatory activities (zone 0) to team-based clinical scenarios (zones 1-4). However, empirical evidence on the impact of zone 0 regarding procedural skill acquisition and retention remains limited. Objective:The present pilot study aimed to build on this evidence gap by examining whether the addition of a structured zone 0 self-directed preparatory phase, delivered through interactive video (IV), could enhance cardiopulmonary resuscitation (CPR) competence acquisition, skill retention, and CPR quality among nursing students when combined with conventional instructor-led zone 1 training, while also assessing the feasibility and acceptability of the intervention. Methods:A total of 52 nursing students from San Pablo CEU University (Madrid, Spain) were randomly assigned to an experimental group (n=33, SimZone 0 IV followed by SimZone 1 instructor-led seminar) or a control group (n=19, SimZone 1 instructor-led seminar only). CPR competence acquisition, skill retention, CPR quality, and feasibility were assessed immediately after zone 1 training and at 3 and 6 months. Results:In the self-directed zone 0 phase, 73% (24/33) of experimental group students achieved competence. After zone 1 training, 100% (28/28) of the experimental group achieved competence vs 84% (16/19) of the control group (P=.06, Fisher exact test). At 3 months, competence was 96% (23/24) vs 71% (10/14; P=.05), and at 6 months, 96% (22/23) vs 62% (8/13; P=.02). Mean CPR quality scores were consistently higher in the experimental group across all time points (zone 1: 94%, SD 11.8% vs 90%, SD 16.4%; 3 months: 92%, SD 10.9% vs 90%, SD 12.1%; and 6 months: 92%, SD 10.9% vs 85%, SD 17.8%). The intervention was feasible, well accepted, and free of adverse events or technical issues. Conclusions:Self-directed IV learning as a preparatory phase before instructor-led training is feasible, acceptable, and is associated with preliminary signals of enhanced CPR skill acquisition and retention among nursing students. These findings should be interpreted with caution given the pilot nature of the study and the small sample size.
Unlabelled:This study assesses the feasibility and costs of reaching, screening, and enrolling sexual and gender minority adolescents aged 13-17 years in the Southern United States in an online mixed method study, with a focus on state-level differences in recruitment and expense.
Background:Autism spectrum disorder is underdiagnosed in adults, with increasing demand on diagnostic services and prolonged waiting times. AI-powered tools may offer scalable solutions for early screening and triage. Objective:This proof-of-concept study aimed to evaluate the feasibility, acceptability, and user experience of ASIST (Autism Screening With Intelligent Supportive Technology), an AI-powered conversational platform designed to support structured history taking and referral preparation for adults seeking to explore autistic traits. Methods:A mixed methods feasibility study was conducted. Adults (N=12) interacted with a voice-based AI chatbot delivering validated screening tools (10-item Autism Spectrum Quotient and 2-Minute Autism Detection Scale). Quantitative acceptability and usability were assessed using items informed by the theoretical framework of acceptability alongside open-text qualitative feedback. Results:Eleven patient and public involvement and engagement contributors informed the development and refinement of the study, and 12 adults completed the pilot evaluation. Participants generally reported positive perceptions of the chatbot, including low effort, favorable confidence, and perceived fairness. Open-text feedback highlighted the perceived value of ASIST as a history taking and referral support tool while also identifying areas for refinement, including pacing, speech clarity, and response format. Conclusions:AI-powered conversational tools may offer scalable solutions for structured history taking, referral preparation, and early triage. Further large-scale validation, pathway integration, and equity-focused evaluation are required before wider implementation.
Background:Systematic literature reviews (SLRs) are essential for evidence synthesis in health research but remain labor-intensive, especially at the screening stage. Manual review of titles and abstracts requires substantial human effort, while existing automation tools still have limited adoption in health technology assessment. The EQ-5D questionnaire, a widely used patient-reported outcome measure for health-related quality of life, provides data that frequently underpin reimbursement and policy decisions. Objective:This pilot study evaluated whether recent large language models (LLMs) can support the identification of publications reporting EQ-5D data in PubMed records, using only publicly available metadata (title, abstract, and keywords). Methods:A total of 200 publications retrieved through the EuroQol PubMed filter were manually labeled by experts as reporting or not reporting EQ-5D data. The dataset was split into stratified training, validation, and test subsets. Several machine learning approaches were compared, including a Naïve Bayes baseline using bag-of-words features, a decision-tree model based on full-text keyword occurrence, and transformer-based LLMs (Bidirectional Encoder Representations from Transformers [BERT], Biomedical BERT [BioBERT], Scientific BERT [SciBERT], and Biomedical Language Understanding Evaluation BERT [BlueBERT]). Both classifier-only and fine-tuned configurations were tested across multiple learning rates. Model performance was assessed using accuracy, precision, recall, and F1-score. Results:Baseline approaches achieved near-random test performance (accuracy around 0.53). Classifier-only LLMs modestly improved results (accuracy up to 0.64 with SciBERT). Fine-tuned models substantially outperformed these baselines, with BERT and BioBERT achieving the best performance (accuracy=0.70; F1-score=0.68). In screening-oriented evaluation, this configuration achieved 90.0% sensitivity, 40.0% specificity, and 6 false negatives on the held-out test set. The models reproduced human screening tendencies despite the small dataset size, demonstrating the technical feasibility of LLM-assisted article selection. Conclusions:This study provides the first demonstration of LLM-assisted identification of EQ-5D data in biomedical literature. The findings support technical feasibility but do not establish a reliable stand-alone automated screening tool. Although limited by dataset size, the proposed workflow is reproducible and adaptable to other patient-reported outcome measures. Because validation was based on a single small train-validation-test split, the results should be interpreted as preliminary; future work will scale data collection, include statistical testing, and explore semisupervised learning to further reduce manual screening workload.
Background:The prospective prescription review system can improve prescription rationality, but its effectiveness in high-volume specialty care settings is not well established. Objective:We aimed to evaluate the effectiveness of the prospective prescription review system in reducing irrational prescriptions and to analyze factors associated with successful interception. Methods:This retrospective cohort study analyzed all outpatient and emergency prescriptions issued between January 1 and December 31, 2024, at an eye, ear, nose, and throat tertiary hospital in Shanghai, China. Among 2,559,342 prescriptions, 123,914 (4.84%) flagged as irrational by the prospective prescription review system were included. Results:The overall interception success rate was 32.99% (40,877/123,914). The prescription rationality rate increased from 95.23% (2,476,305/2,600,219) to 96.79% (2,477,144/2,559,342) after excluding intercepted prescriptions (P<.001). Higher review levels were strongly associated with higher interception success rates: 0.99% (517/52,146) for reminder, 54.2% (37,039/68,341) for warning, and 96.91% (3321/3427) for mandatory (P<.001). Revised prescriptions had a higher interception success rate than prescriptions without revision (35,008/50,418, 69.44% vs 5869/73,496, 7.99%; P<.001). Prescriptions without documented reasons showed a higher interception success rate than those with documented reasons (40,086/98,070, 40.87% vs 791/25,844, 3.06%; P<.001). Physician acceptance of pharmacist disapproval was associated with a successful interception rate of 85.63% (137/160), compared with 3.33% (2/60) when the disapproval was rejected (P<.001). Multivariable logistic regression showed that inappropriate dosage or frequency (odds ratio [OR] 3.11, 95% CI 2.99-3.24) and inappropriate prescription quantity (OR 2.09, 95% CI 2-2.18) were most likely to be intercepted, whereas inappropriate indications (OR 0.04, 95% CI 0.03-0.04) and radiation oncology (OR 0.01, 95% CI 0.01-0.02; P<.001) were more likely to be issued. Pharmacist-led rule revision for mometasone furoate nasal spray led to a 1000-fold increase in identified irrational prescriptions (from 2-4 per month to 3022 in May). The interrupted time series analysis showed that the April revision was associated with an immediate and statistically significant increase in the monthly overall interception success rate (intervention coefficient = 0.15; P =.005). Conclusions:The prospective prescription review system substantially reduced irrational prescriptions and improved prescription rationality. Higher review levels were associated with higher interception rates, and physicians' attitudes also played a critical role. Paired-organ dose standardization emerged as an important consideration in ophthalmology and otolaryngology practice. Collaboration between the prospective prescription review system and pharmacists is essential for optimizing prescription review outcomes.
Unlabelled:Despite widespread discussion of opportunities and risks about AI in medicine, few health care professionals routinely used AI during this study period. Physicians and men were more likely to report frequent AI use, and frequent AI users more often reported positive sentiments about AI's future impacts on pay, enjoyment, and productivity at work. The youngest professionals (≤29 years) were less engaged and more skeptical about AI's future benefits. These findings highlight a gap between AI's promise and medical practice, as adoption varies by role, experience, and demographics.
Background National digital health systems are increasingly recognized as a key tool for strengthening health care systems and improving patient outcomes. Global guidance from the World Health Organization and European Union initiatives emphasizes interoperability, robust data governance, and secure exchange of health information as prerequisites for system-wide benefits. However, translating these principles into operational national digital health infrastructures remains challenging, and empirical evidence on long-term, centralized implementation models is limited. Lithuania is among a small number of European countries that have implemented a fully centralized national digital health platform (DHP), providing a valuable case for examining the development, structure, and performance of such systems. Objective This study aimed to describe and analyze Lithuania’s national digital health system across its historical development, governance, technical architecture, interoperability and adoption, and to identify challenges and lessons learned. Methods We conducted a descriptive national case study of Lithuania’s digital health system using longitudinal document analysis. Publicly available policy documents, legislation, strategy reports, audits, and official statistics published between 2000 and 2025 were systematically reviewed following the readying, extracting, analyzing, and distilling (READ) framework. Thematic analysis was applied to examine governance, system architecture, interoperability, adoption, and implementation challenges. Peer-reviewed literature was used for contextualization. Results Lithuania’s digital health system evolved from fragmented initiatives in the early 2000s into a fully centralized national DHP that has been operational nationwide since 2015. Since 2018, health care providers have been required to record core clinical data electronically, and all reimbursed medicines are prescribed exclusively through the national electronic prescription system. The DHP integrates electronic health records, electronic prescription system, appointment booking, and other national subsystems with standardized interfaces. System use expanded rapidly: the annual number of electronic documents stored in the DHP increased from approximately 7 million in 2017 to over 100 million in 2024, reflecting widespread uptake of outpatient documentation, referrals, and prescriptions. Patient-facing services, including online access to records and a nationwide appointment booking system, also expanded, although patient-initiated use remains uneven across regions. The platform supports secure primary use for care delivery and secondary use of health data through a centralized access pathway. While core services have reached national scale, some subsystems, such as imaging exchange and telemedicine, remain less mature. Ongoing modernization focuses on improving data interoperability, enabling advanced analytics and future AI-based services, strengthening cross-border eHealth services, and extending coverage to sensitive and previously excluded domains such as mental health, maternal, and newborn care through dedicated subsystems. Conclusions Lithuania’s centralized digital health system illustrates both the potential and the limitations of national-scale integration driven by legal mandates and standardized interoperability. Its experience provides transferable insights for health systems seeking to consolidate fragmented infrastructures while maintaining data security, regulatory compliance, and long-term system adaptability.
BACKGROUND:In the United States and worldwide, chronic pain affects a vast number of people and is one of the leading reasons adults seek medical care. In the United States, 24.3% of adults reported experiencing chronic pain in the prior 3 months in 2023. Chronic pain is defined as pain that lasts ≥3 months, significantly disrupting one's daily functioning and quality of life. Chronic pain can also be accompanied by other conditions, including anxiety and depression. Health care providers should be aware of several limitations in their current treatment modalities. Although opioid analgesics are used for moderate to severe pain management, they cause many serious adverse effects, including sedation, respiratory depression, constipation, and a high risk of dependence and addiction. Other pain medications, such as nonopioid analgesics, including nonsteroidal anti-inflammatory drugs, adjuvant analgesics, and corticosteroids, also cause a range of side effects and organ toxicity. OBJECTIVE:We conducted a single-arm exploratory pilot study to explore the potential role of a noninvasive, transdermal, audio-based therapeutic system among 24 police professionals aged 21 to 65 years with chronic pain and/or associated conditions, including migraine, sleep disorders, stress, anxiety, and conventional headaches. METHODS:In this STROBE (Strengthening the Reporting of Observational Studies in Epidemiology)-guided pilot study, we explored the feasibility and perceived effectiveness of a self-administered intervention based on the Transdermal Acoustic Pain Suppression (TAPS) system among police personnel in West Palm Beach, Florida. Participants completed a flexible 30-day protocol integrated into their work schedules, with data collection spanning 5 months due to operational constraints. We used percentages as measures of effect and the chi-square test and Fisher exact test for significance testing. All P values were 2-sided, with a level of significance of .05. RESULTS:In this small pilot study, following TAPS therapy, more than one-third (n=9, 37.5%) of participants experienced improvement in their condition, and half (n=12, 50%) reported reduced symptoms, including notable pain reduction. CONCLUSIONS:Although the findings from this small pilot study are preliminary, we believe they support the rationale for larger analytic studies designed a priori to test the hypothesis.
Background:Health care professional-generated vignettes are commonly used to illustrate and analyze patient experiences, shaped through clinical reflection to support provider understanding, training, and practice improvement. However, the subjective interpretation of narratives and the time-consuming manual process limit this approach. As an alternative, patients could record video blogs (vlogs) of their experiences, which can then be transformed into vignettes using GenAI and carefully designed prompts. Objective:This study aims to compare the feasibility and utility of 2 different vignette creation methods, manual and ChatGPT-3.0, from user-generated vlogs about living with chronic pain. Methods:Short commentary videos were recorded by a person living with chronic pain of unknown origin. These were transcribed and used to create vignettes. From 64 videos, three 1-minute clips were selected, each representing a key life stream: university student, international traveler, and patient with chronic pain, based on prior thematic analysis. In total, 3 vignettes were created for each creation method (manual and AI-generated). Each vignette incorporated a profile description, 3 selected user-generated video clips representing the participant's 3 life streams, researcher-generated analytic commentary for each clip (produced by a researcher with 15 years of experience in manual vignette creation for clinical use), and a brief summary of each video. A comparison between manually constructed and AI-generated vignettes evaluated coherence, contextual accuracy, narrative depth, and utility in health communication and user experience research. Additionally, the person responsible for generating the raw videos (Participant X) provided commentary on this process as a cocreator of a new method, offering critical insights into the authenticity, emotional resonance, and perceived usefulness of each vignette format. Results:This study demonstrates that manual and ChatGPT vignette creation methods are feasible approaches for analyzing and representing user-generated vlogs about living with chronic pain. Each method offers distinct utility: manual vignettes provide rich, nuanced narratives capturing emotional and social complexities, while ChatGPT-generated vignettes offer efficient, concise summaries with some loss of detail. Participant X's reflections reinforced these findings. Conclusions:ChatGPT-generated vignettes, when combined with human review, offer the potential for an efficient and scalable approach to capturing the experiences of patients with chronic pain. These vignettes could support training, communication, and sharing of patient insights.
Background:AI is increasingly integrated into medical education, offering new ways for students to acquire knowledge and support clinical reasoning. However, the extent, patterns, and implications of AI use among medical students remain incompletely understood. Prior studies have relied on retrospective surveys that are susceptible to recall bias and have not quantified AI use as a proportion of total study time. Objective:This pilot study aimed to quantify real-time AI use among medical students, including the proportion of study time devoted to AI, and how use varies by training stage and engagement style (active vs passive). Active use was defined as iterative, bidirectional engagement; passive use was defined as unidirectional consultation with limited interrogation. Methods:This longitudinal observational cohort study recruited medical students from 2 osteopathic medical schools (April-May 2025) to complete a baseline survey and 7 digital diary entries over a 21-day period, delivered via automated SMS every 3 days. Students reported total study time, AI use time, tools used, and purposes of use. The data were analyzed using Stata 19. Multiple linear regression models examined associations between AI use (total minutes and percentage of study time) and key variables, and a mixed-effects model using diary-level data with a random intercept per student addressed within-person variability across entries. Results:A total of 71 of 1332 (response rate: 5.3%) eligible students completed the baseline survey (mean age 26.6, SD 2.8 y; n=39, 55% identified as men; n=32, 45% identified as women; n=55, 77% preclinical). On average, students reported using AI tools during 19% of their total study time (mean 35.8 of 185.6 min per diary, SD 35.8 min). The most used tool was ChatGPT (n=63, 89%), followed by Google Gemini (n=22, 31%). Clinical-phase students (MS3-MS4) used AI significantly more than preclinical students (MS1-MS2), with an adjusted increase of 19% (P=.003). Students classified as active users spent significantly more total time using AI than passive users (P=.002). Across groups, AI use was primarily passive, including simplifying complex concepts, answering practice questions, and generating summaries. In the multilevel models, preclinical students reported significantly lower AI use than clinical-phase students (P=.02). The intraclass correlation coefficient was 0.47 (95% CI 0.34-0.60). Conclusions:Although preliminary, these findings suggest that medical students are incorporating AI into a substantial proportion of their study time, with greater use among clinical trainees and active users. Despite this, most use remains passive. Given mixed evidence on AI's impact on deep learning, further research on learning outcomes is needed. Institutions may consider providing guidance on responsible AI use, including critical evaluation and verification of outputs. The digital diary methodology offers a practical approach for capturing real-time AI use and may inform future educational research and intervention design.
Background:Respiratory inductance plethysmography (RIP) belts are the standard for measuring thoracoabdominal movements in home sleep apnea testing (HSAT), but are often cumbersome and power-intensive. Objective:This study aimed to characterize abdominal movement signals measured by a wireless, single-point, abdomen-worn sensor to evaluate the sensor's capability to track respiratory dynamics. Methods:Overnight recordings were obtained from 37 participants using a wireless abdomen-worn sensor and a thoracic RIP belt during HSAT. The abdominal movement signal was analyzed for breath detection, respiratory rate (RR) estimation, and waveform similarity relative to the RIP signal. Factors influencing the agreement between the 2 signals were also investigated. Results:Data from 34 participants were analyzed. The abdominal movement signal showed moderate agreement with the thoracic RIP signal, achieving a sensitivity of 82.44%, a positive predictive value of 76.22%, and an F1-score of 78.99% for breath-cycle detection. RR estimation yielded a mean absolute percentage error of 5.42% and limits of agreement of ±3 breaths per minute (bpm). Morphological similarity was moderate, with an average distance correlation of 0.73 and a mean squared error of 0.64. The agreement between the 2 signals declined with increasing respiratory event severity. Conclusions:Compared with a standard thoracic RIP belt, the single-point, wireless, abdomen-worn sensor tracked basic respiratory metrics and waveform morphology with moderate agreement. These findings show promising baseline performance and suggest its viability as a simplified, low-profile data-acquisition platform for home-based respiratory monitoring.
Background:Popular discourse often frames prescription stimulants Ritalin and Adderall as drugs for teens and emerging adults with greater financial resources (eg, students or young professionals), while illicit stimulants such as crack or methamphetamine are often considered drugs for older adults with lower incomes. Crack and methamphetamine are also frequently used in combination with opioids to offset the sedating effects of adulterants such as fentanyl and xylazine in the unregulated drug supply. This combination of stimulant and opioid use creates additional overdose risks and may require new public health approaches. Our team posited that creating effective messages to protect against overdose in the context of stimulant use entailed first developing a better understanding of which stimulants people are using in combination with opioids. Objective:While conducting formative research to develop a new intervention, our team sought to examine prescription stimulant use among participants, including middle-aged, lower-income, and homeless people. This entailed first ascertaining the prevalence of Ritalin and Adderall use among participants, and then asking people why they used them. We also sought to implement a novel AI-assisted coding methodology and to write up the steps we took as a replicable model that other research teams could readily use to facilitate their own data analysis. Methods:We collected substance use screenings during 3 waves of data collection in 2025 (N=102 participants). In early 2026, we conducted 16 qualitative interviews with people who reported using Ritalin or Adderall. We then conducted mixed methods analyses to examine reported substance use, including the use of opioids and Ritalin or Adderall in combination, as well as potential relationships between using opioids with prescription stimulants and reporting a desire to stop using drugs. After conducting the interviews, we implemented a new hybrid methodology using AI tools to code interview transcripts and generate detailed reports accompanied by supporting quotes. Results:Participants described using Ritalin or Adderall to manage negative effects of increasingly powerful opioids, including oversedation and withdrawal symptoms. A subset of participants who reported using both Ritalin or Adderall and opioids were 3.5 times more likely to agree or strongly agree with the statement "I want to stop using drugs but I need help to do that" compared to those who did not report using these substances in combination (19/22, 86.4% vs 18/28, 64.3%; odds ratio 3.52, 95% CI 0.83-14.89; P=.09). Although not statistically significant, this difference may prove practically significant and merits further study. Conclusions:The use of prescription stimulants appears to be increasingly common among new populations of people who use opioids, including those who report very low incomes and even homelessness. Additional research is warranted to examine how this type of polysubstance use may require different types of prevention strategies.
Background:Large language models (LLMs) have shown promising performance on medical examinations across specialties. However, comparative evaluations of current-generation LLMs across multiple European anesthesiology examinations, alongside structured assessment of hallucinations vs question-related confusion, remain lacking. Objective:This study aimed to compare the performance of 4 state-of-the-art LLMs on anesthesiology and intensive medicine examination questions and assess their hallucination rates. Methods:This computational comparative study analyzed 437 multiple-choice questions (1748 queries) from 3 sources: nurse anesthetist school examinations (infirmier anesthésiste diplômé d'État [registered nurse anesthetist]; n=100, 22.9%), European Diploma in Anaesthesiology and Intensive Care (EDAIC; n=219, 50.1%), and EDAIC On-Line Assessment (n=118, 27.0%). Each question was submitted to 4 LLMs (Claude Sonnet 4.5, Gemini 2.5 Pro, GPT-5, and Grok 4) using standardized prompts via default web interface settings. Responses were evaluated through structured consensus review by 2 examiners for accuracy, hallucinations, and question-related confusion. Statistical analysis included Friedman and Wilcoxon signed-rank tests with Holm-Bonferroni correction, the Cochran Q test, and generalized estimating equations. Results:Average success rates ranged from 86% (SD 18%) to 94% (SD 10%) across LLMs and examination types, exceeding the EDAIC part I passing threshold, representing substantial improvement over previously reported GPT-3.5 performance. For the EDAIC, overall intermodel differences were significant (Friedman χ23=13.9; P=.003; W=0.02), with Gemini outperforming GPT-5 as the only pairwise difference. Hallucination rates ranged from 11% (11/100) to 20.1% (44/219) without significant intermodel differences. All models exceeded the EDAIC passing threshold. Conclusions:Current-generation LLMs demonstrated consistently high performance across multiple European anesthesiology examinations but continue to produce clinically relevant hallucinations, supporting their role as supervised educational tools rather than autonomous learning resources. These findings underscore the need for structured integration frameworks and systematic verification when deploying LLMs as learning tools in medical education.