The impact of incorporating artificial intelligence (AI) into a double-read breast-screening workflow, including arbitration, is unclear. This retrospective study included 50,000 representative women from two NHS breast-screening centers. All the women had long-term follow-up, allowing us to determine whether use of AI leads to earlier cancer detection. Cases requiring arbitration (8,732 cases) were read by 22 readers in a reader study, following their normal arbitration workflow. Overall, after arbitration, replacing the second reader with AI was noninferior (5% margin) to two human readers in terms of sensitivity and specificity (P < 0.001) while offering a workload benefit. Arbitration improved the specificity of the AI arm by overruling cases incorrectly recalled by the AI tool; however, it also overruled the AI tool recall decision for some interval and next-round cancers. Further development of the AI tool alongside improvement in its explainability could lead to the earlier detection of cancers.
Despite advances in deep learning and transformer architectures, prior reviews have focused narrowly on traditional clinical decision support systems (CDSS) or single medical domains, leaving significant gaps in understanding contemporary AI-driven predictive tools. This systematic review and meta-analysis evaluated the predictive performance of artificial intelligence-based CDSS (AI-CDSS) across multiple medical specialties. Following PRISMA guidelines, PubMed and Cochrane Library were searched through December 2024 for studies evaluating predictive AI-CDSS using real-world clinical data. Two reviewers independently screened 3,296 records (κ = 0.833), with study quality assessed via QUADAS-2 and performance measures pooled using random-effects meta-analysis. Fifty studies spanning 17 medical specialties were included. Meta-analysis demonstrated moderate discriminatory ability (pooled AUC: 0.652, 95% CI: 0.562-0.743), high specificity (0.819, 95% CI: 0.793-0.844), moderate accuracy (0.765, 95% CI: 0.734-0.796), and variable sensitivity (0.660, 95% CI: 0.535-0.785), with substantial heterogeneity across all measures (I² ≥ 98.9%). Only 24% of studies involved prospective deployment, and 64% reported exclusively technical metrics without clinical workflow data. Predictive AI-CDSS demonstrate moderate-to-good diagnostic performance with strong specificity; however, the predominance of retrospective study designs and limited implementation reporting reveal critical gaps between technical validation and real-world clinical utility. To address these shortcomings, we propose the ROADMAP framework, structured around seven domains: Representative development, Outcomes-focused evaluation, Assessment for deployment, Data harmonization, Monitoring for bias, Allocation via economic evaluations, and Priorities for standardized reporting and prospective validation. This framework provides a practical roadmap for bridging the gap between algorithmic performance and meaningful clinical integration.
Regulation plays a pivotal role in shaping the adoption of AI-based surgical tools to enhance patient care. However, differences in global regulatory frameworks can influence how quickly and safely these surgical AI devices reach clinical practice. This article compares the regulatory landscapes of the United States (US), European Union (EU), United Kingdom (UK), and China, and examines how their respective regulatory authorities approach the approval of surgical AI devices, including those used for preoperative risk stratification and planning, intraoperative navigation and computer vision, and postoperative monitoring. While all jurisdictions follow a risk-based paradigm, key differences centre on classification thresholds, pathways for approval, clinical evidence requirements, and the pace of market entry. The US Food and Drug Administration (FDA) relies heavily on the 510(k) process, facilitating moderate-risk device clearances at a rapid pace, though concerns persist regarding predicate creep for iteratively evolving algorithms. Manufacturers seeking approval in Europe must navigate the Medical Device Regulation (MDR) and the emerging requirements of the EU AI Act, which introduce additional obligations for high-risk AI systems. The UK continues to shape its framework post-Brexit under the Medicines and Healthcare products Regulatory Agency (MHRA), including through novel oversight models such as the AI Airlock regulatory sandbox. In China, the National Medical Products Administration (NMPA) designates most decision-support AI tools as high-risk, increasing approval scrutiny but potentially slowing market entry. This comparative analysis underscores a growing need for transparency and international collaboration to enhance patient safety, anticipate evolving technologies, and ensure that the potential of AI is realised in global surgical practice.
Connected Medical Devices (CMD) are redefining care within the NHS but exposing it to bi-directional cyber-physical threats that traverse physical, network and cloud layers. These vulnerabilities blur the boundary between technology and patient safety. This Comment argues that the MHRA should elevate cybersecurity to a clinical-safety mandate, enforcing a unified socio-technical framework with security-by-design, cross-layer risk assessment and continuous post-market vigilance.
Current antimicrobial resistance (AMR) surveillance relies on fragmented indicators that fail to capture institutional AMR burden complexity. To develop the AMR Burden Score, a three-round modified eDelphi study engaged interdisciplinary experts (Round 1: n = 17, Rounds 2 and 3: n = 7), including clinicians, microbiologists, pharmacists, health economists, and public health specialists. The AMR Burden Score comprises six weighted domains: Resistance (25%), Effectiveness (30%), Monitoring (30%), Adoption (5%), Processes (5%), and Systems (5%). Strong consensus emerged for core indicators, including incidence of resistant infections (unanimous Round 3 agreement, median 8.0), pathogen-specific resistance rates (median 7.0), and staff training programmes (median 8.0). The AMR Burden Score provides a structured framework for institutional AMR assessment, though implementation requires context-specific adaptation and further validation.
BackgroundEndoscopic submucosal dissection (ESD) enables en bloc resection of early gastrointestinal (GI) neoplasia, but its adoption in Western countries has been limited by training and service constraints. This study aimed to describe the current landscape of ESD in the UK.MethodsA nationwide, cross-sectional, anonymised online survey was conducted among UK endoscopists independently performing either upper and/or lower GI ESD.ResultsTwenty-eight responses were analysed. Most respondents were gastroenterologists (79%) and male (82%). Responses were received from across the UK, with the largest proportions from Greater London (39%) and the South East (18%). Overall, 68% had completed an advanced fellowship (commonly in the UK or Japan), and 82% had attended accredited ESD courses, often on multiple occasions. Across all ESD sites, only a minority perform more than 20 procedures annually, with high-volume practice largely confined to rectal ESD. Narrow Band Imaging (93%) and white light endoscopy (86%) were the most commonly used delineation methods; colloid-based solutions (75%), epinephrine (82%), and blue dye (100%) were widely used for submucosal injection. Most respondents routinely used tunnelling (93%) and traction-assisted techniques (82%). General anaesthesia was preferred for upper GI ESD (82%), and conscious sedation for lower GI ESD (46%). Reported barriers to training included heavy workload (37%), low caseload (26%), and limited institutional support (19%). Nearly all respondents (93%) believed robotics has a future role in ESD.ConclusionsIndependent ESD practice in the UK is delivered by a small workforce, with heterogeneous training backgrounds and procedural volumes often below recommended thresholds. National strategies for structured training, centralisation, and coordinated service planning are needed to support safe and sustainable expansion of ESD.
Background: Virtual consultations (VC) are now integral to primary care delivery, but their implementation has been hampered by technological limitations, user engagement issues, and the lack of validated implementation frameworks.Objective: This study sought to identify and validate key elements required for the successful adoption of high-quality VC and to develop recommendations for its implementation in primary care.Methods: An electronic Delphi (eDelphi) study was conducted between March and December 2024 with Primary Care Physicians (PCPs) from 24 countries who had experience using VC in routine practice. Participants assessed 25 proposed toolkit components across six domains using a 5-point Likert scale. Descriptive statistics were applied, and consensus was defined as ≥70% agreement.Results: A total of 135 PCPs participated across three Delphi rounds. Consensus was achieved for 24 of the 25 framework elements. High priority was placed on long-term strategic planning and financial sustainability, sustained investment in VC infrastructure, enhanced cybersecurity and system resilience, improved digital literacy, and clear communication strategies for end-users. The sole component failing to reach consensus related to religious factors affecting VC acceptability, with most participants favouring inclusion of religion within broader cultural considerations rather than as a separate element.Conclusion: This study validated key components of a flexible VC implementation framework for primary care. The lack of consensus on certain contextual factors highlights the importance of culturally sensitive adaptation when implementing VC. Future work should focus on real-world piloting of the framework to evaluate its effectiveness at practice, policy, and system levels.
Objective This scoping review aims to provide an overview of the methodology used when defining the learning curve (LC) in live laparoscopic surgery. Design This review was performed in line with the Preferred Reporting Items for Systematic Reviews and Meta Analyses extension for Scoping Reviews (PRISMA-ScR) guidelines. English-language articles were included systematically using Boolean operators for search string combinations. Selected studies were qualitatively analyzed for LC measurement methodology, including statistical approaches. Additional analysis was undertaken by grouping LC metrics into intrinsic, surrogate, and outcome-based. Setting Only original publications in the English language were included. No date restrictions were applied. Eligible studies were those that measured LC in laparoscopic cases (i.e., not simulation studies) regardless of their surgical specialty. Robotic-assisted laparoscopic operations were also included. Articles analyzing non-laparoscopic procedures were not included. Results A total of 87 articles were extracted for review. The majority (45/87) were conducted within general surgery alongside gynecology, vascular, pediatrics, and urology. A total of 50 studies analyzed the LC for a new technology or novel technique. Multiple metrics were used to quantify LC, with the most common being operative time (n = 75), followed by complications (n = 31). Only 3 studies exclusively used intrinsic performance measures for LC analysis. There was significant heterogeneity in the statistical analysis and description of LC. The majority of studies (51/87) grouped cases temporally. Nineteen studies used a Cumulative Sum (CUSUM) statistical analysis. Conclusions Measurement of the LC in live laparoscopic surgery has significant potential for training and appraisal of new techniques. Currently, the time-intensive nature of LC measurement limits its widespread adoption. However, newer techniques used alongside real-time measurement of metrics within the OR may overcome this. Future work should focus on identifying metrics that can be transferred between procedures to allow comparability and standardized statistical analysis.
Background:Social media has transformed the landscape of health communication. Video content can optimally activate our cognitive systems, enhance learning, and deliver accessible information. Evidence has suggested the positive impact of videos on health knowledge and health-related behaviors, yet the impact of social media videos on quantitative health outcomes is underresearched. Evaluating such outcomes poses unique challenges in measuring exposure and outcomes within internet-based populations. Objective:We aimed to evaluate the impact of social media videos on quantitative health outcomes, examine methodologies used to measure these effects, and describe the characteristics of video interventions and their delivery. Methods:In accordance with PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, MEDLINE, Embase, Web of Science, CINAHL, and Google Scholar were searched. Studies were eligible if they were original research evaluating long-form social media video interventions addressing any health-related condition, delivered via social media platforms, and reported quantitative health outcomes. The primary outcome was the effect of social media videos on quantitative health outcomes. Additional outcomes included participant characteristics, video features, delivery methods, and the use of theoretical frameworks. A narrative synthesis was conducted. A subgroup meta-analysis was performed to synthesize health outcomes mentioned in 2 or more studies with sufficient homogeneity. Risk of bias assessment was conducted using Cochrane Risk of Bias 2, ROBINS-I, or National Institutes of Health Quality Assessment Tool, depending on the study design. One reviewer screened titles and abstracts. Two reviewers independently conducted full-text screening, data extraction, and risk of bias assessment. Results:A systematic search was conducted on October 25, 2023, and was updated on June 12, 2025, yielding a total of 41,172 records after duplicate removal. Sixteen studies were included, involving 4158 participants. Mental health-related conditions were the most studied (10 studies). Most video interventions were delivered via YouTube (12 studies). Studies have reported that video interventions were associated with significant improvements in peri-procedural anxiety, mood, and physical activity levels, although most findings were limited to individual studies with variable methodological quality. Three studies that developed videos with user input and theoretical frameworks significantly impacted study-specific primary outcomes. A subgroup meta-analysis demonstrated a significant moderate impact of online video interventions in improving peri-procedural anxiety (standard mean difference=0.57, 95% CI 0.09-1.05). All but one study showed some concern or high risk of bias. Conclusions:We demonstrated a potential positive impact of social media videos on quantitative health outcomes, notably in improving peri-procedural anxiety. Videos developed with user input and theoretical frameworks significantly impacted study-specific primary outcomes. Nevertheless, there is the need to shift focus toward measuring physical health-related outcomes and to develop better designed, innovative methodologies to measure the impact that can better simulate the social media environment.
Frailty prevalence among older adults is rising globally, with significant implications for quality of life, healthcare costs, and economic productivity. While multicomponent lifestyle interventions can reverse frailty, existing programmes such as Vivifrail and the Otago Exercise Programme are limited by poor long-term adherence. Digital technologies combined with behavioural science principles offer a potential solution, yet few interventions have been developed with meaningful input from older adults themselves. To co-design a digitally-enabled intervention and to evaluate its feasibility, acceptability, and preliminary effects on health behaviours and frailty status in community-dwelling older adults. A mixed-methods feasibility study was conducted with 30 community-dwelling adults aged 65 and over (mean age 72.6 years; 45% men, 55% women) scoring 0-4 on the Fried Frailty Scale at enrolment. The "Healthy Habits" intervention was iteratively developed through four co-design cycles, incorporating wearable sensors (smartwatch, smart scale, sleep analyser), and a custom mobile application providing personalised feedback on seven habit domains. Participants engaged for a minimum of six weeks, with optional extended participation. Outcomes included retention, engagement metrics, changes in daily step count and sedentary time, grip strength, and frailty scores (Fried and Edmonton scales). Retention was 97% (29/30) at six weeks, with 83% of completers electing to continue participation up to a maximum of 54 weeks. Smartwatch adherence was high (80% wear time). Among participants, 79% (22/28) demonstrated greater than 10% improvement in either daily step count or sedentary time. In the subgroup receiving the digital application (n=7), 71% increased daily steps by more than 10% within six weeks. Mean grip strength improved by 8% (p=0.013) and Edmonton Frailty Scale scores decreased by 48% (p=0.015). A co-designed, digitally enabled intervention based on habit formation principles is feasible and acceptable to older adults, with high retention and preliminary evidence of improvements in physical activity, sedentary behaviour, and frailty markers. These findings support progression to a randomised controlled trial to evaluate efficacy. N/A
Background Artificial intelligence (AI) and machine learning (ML) have shown immense potential in cardiology, leveraging data-driven insights to enhance diagnosis, treatment planning and patient care. This study presents a comprehensive evaluation of US Food and Drug Administration (FDA)-approved AI/ML devices in cardiology, analysing trends in clinical applications, regulatory pathways and evidence transparency.Methods FDA clearance summaries from the AI/ML medical device database were reviewed to identify cardiology-specific applications. Devices were categorised using the descriptive, diagnostic, predictive and prescriptive framework. Regulatory pathways, AI technologies and validation data were critically assessed.Results Of 1016 FDA-approved AI/ML devices, 277 (27.3%) had cardiology applications, predominantly for imaging (65.3%) and diagnostics (64.3%). Predictive and prescriptive tools constituted only 5.4% and 0.7%, respectively. Most devices (97.1%) were cleared via the 510(k) pathway, with 58.0% at risk of predicate creep. Quality of clinical evidence was limited, with only 3.2% of devices supported by high-quality trials. The type of AI technology was often underreported (58.8%).Conclusion While AI/ML technologies are reshaping cardiology, regulatory challenges and reporting transparency impede their optimal use. Strengthened regulatory frameworks, improved trial design and robust post-market surveillance are essential to ensure safety, efficacy and equity in the deployment of AI tools in cardiology.
Abstract Background Shoulder dysfunction is common after axillary surgery for breast cancer, impairing range-of-motion (ROM), strength, and quality-of-life (QoL). NHS rehabilitation is inconsistent, often limited to exercise leaflets without structured follow-up. Wearable devices may enable personalised, feedback-driven rehabilitation. This study evaluated feasibility, acceptability, and preliminary clinical outcomes of wearable-driven rehabilitation versus standard care. Methods In this single-centre feasibility trial, 72 patients undergoing axillary surgery were randomised 2:1:1 to feedback-enabled wearable rehabilitation, or to one of two control arms: standard care with passive monitoring or standard care alone. Feasibility outcomes included recruitment, retention, adherence, fidelity, and acceptability. Exploratory outcomes, assessed at baseline and four weeks, included upper-limb activity, ROM, strength, lymphoedema, and patient-reported pain, disability, and QoL. Results Feasibility targets were met, with 82% recruitment, 97% retention and high adherence and acceptability. Daily activity returned to baseline by day 10 in the intervention group (100.2%, 95% c.i. 97.8–102.6) but remained below baseline at day 30 in controls (83.6%, 95% c.i. 78.9–88.3). At four weeks, median recovery of shoulder ROM and strength was greater in the intervention arm: flexion ROM 98.9% (IQR 90.1–101.1) versus 78.7% (IQR 68.6–87.3); flexion strength 98.4% (IQR 95.4–105.1) versus 91.2% (IQR 88.6–93.9). Patient-reported outcomes also favoured the intervention, with less pain, lower disability, and higher QoL (P < 0.001). No differences were seen in lymphoedema or complications. Conclusions Wearable-driven rehabilitation after axillary surgery is feasible, safe, and acceptable. Early signals suggest functional and QoL benefits, supporting progression to a multi-centre RCT to test clinical efficacy.
Background and objectives:Modern health systems require physicians to not only provide high-quality clinical care but also understand, navigate, and lead complex healthcare organizations. However, undergraduate medical education in Portugal remains predominantly focused on clinical skills, with minimal exposure to health policy and management (HPM). This study aimed to assess Portuguese medical students' exposure to, attitudes toward, and preferences for HPM education. Methods:We conducted a cross-sectional survey of 483 medical students across 10 Portuguese medical schools. The questionnaire assessed prior exposure to HPM education, self-perceived knowledge of the national health system, curricular preferences, and civic participation. Results:Only 29.2% of participants (n = 141; 95% CI 25.1-33.2) reported any previous HPM training, and those with exposure were more likely to demonstrate higher institutional literacy and greater confidence in understanding health system governance. Overall, 94.8% supported the inclusion of HPM in the medical curriculum, and 64.4% supported making it compulsory, with stronger support among civically engaged students. Conclusion:Portuguese medical students had limited formal exposure to HPM but expressed strong demand for structured training in this area. These findings highlight a misalignment between current curricula and students' perceived needs and support the introduction of a mandatory HPM course in Portuguese medical schools to better prepare future physicians for leadership and governance roles within the health system.
Endoscopic submucosal dissection (ESD) is technically demanding and associated with a steep learning curve and increased complication risk. Constraints of conventional endoscopes, together with the observed benefits of robotic assistance in selected surgical procedures, have driven development of robotic systems for advanced endoscopic applications. This review maps the landscape of robotic endoscopic systems in the context of ESD. A PRISMA-ScR-compliant scoping review was conducted. Six databases were searched in 2025. Studies evaluating robotic endoscopic platforms for ESD in preclinical or clinical settings were included. Two reviewers independently screened studies. Data were extracted on study characteristics, platform features, experimental design, and outcomes. Findings were synthesised descriptively. Twenty-seven studies were published between 2010 and 2025, with most from 2019 onwards (22; 82
Background:The growing reliance on virtual consultations in primary care has reshaped traditional general practitioner (GP)-patient communication dynamics, presenting new challenges that affect care quality and safety. Objective:This study explores communication challenges and gaps, particularly relevant to virtual consultations compared with face-to-face interactions, as well as identifying mitigation strategies from both GPs' and patients' perspectives. Methods:This qualitative study employed 4 online focus group discussions with a purposive sample of UK-based GPs and patients. Data were analyzed using a deductive-inductive thematic approach with NVivo software. The extended Shannon-Weaver communication model and the Capability, Opportunity, Motivation and Behavior model guided the analysis of communication challenges and mitigation strategies, respectively. The Consolidated Criteria for Reporting Qualitative Research were followed to ensure rigorous reporting. Results:A total of 21 participants (12 patients and 9 GPs) took part in 4 online focus group discussions, 2 for patients and 2 for GPs. Six key themes on communication challenges emerged: 5 aligned with the extended Shannon-Weaver communication model (related to the sender-encoder, message, channel, receiver-decoder-feedback, and context), and a new one was inductively identified (patient autonomy and inclusivity). GPs, as senders, highlighted missing visual cues, affecting message clarity in remote communication channels. Patients, as receivers, reported difficulties explaining symptoms remotely, reduced emotional connection, and perceived empathy, linked to contextual challenges and the need for inclusive communication. Mitigation strategies were mapped to the Capability, Opportunity, Motivation and Behavior model: capability (training/resources), opportunity (triage/tools), and motivation (patient engagement/system adaptability), with participants emphasizing tailored training, standardized approaches, and flexible models to support effective and inclusive virtual communication. Conclusions:This study highlights communication gaps in virtual consultations and proposes actionable mitigation strategies. Tailored use of virtual modalities, supported by structured training and policy efforts, is essential to ensure effective and safe remote communication.